Training evaluation method, system and equipment based on multi-modal data fusion and storage medium

Through multimodal data fusion technology, accurate evaluation and real-time error correction of physiological signals and operation trajectory data are achieved, solving the problems of delayed evaluation and insufficient correction in traditional medical training, and improving the targetedness and efficiency of training.

CN120809112APending Publication Date: 2025-10-17CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510901391.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional medical training methods rely on single-modality data acquisition and are unable to combine operation trajectories and physiological signals, resulting in delayed assessments and the inability to correct errors in real time, making it difficult to meet the needs of accurate assessment and real-time intervention.

Method used

By collecting physiological signal data and operation trajectory data, multimodal fusion is performed, three-dimensional space coordinate projection and time interpolation algorithm are used for spatiotemporal alignment, a correlation matrix is ​​constructed, feature vectors are extracted and fused, and a hierarchical evaluation mechanism is used to output comprehensive ability assessment results and trigger dynamic intervention measures.

Benefits of technology

It achieves accurate fusion evaluation of physiological signals and operation trajectory data, provides real-time feedback and error correction, improves the pertinence and efficiency of training, and meets the needs of medical training for accurate evaluation and real-time intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809112A_ABST
    Figure CN120809112A_ABST
Patent Text Reader

Abstract

The invention provides a training evaluation method, system and device based on multi-modal data fusion and a storage medium, which can be applied to operation practical operation training in the field of medical health, and the method comprises the following steps: collecting physiological signal data and operation track data generated in a training process; carrying out space-time alignment on the physiological signal data and the operation track data with different sampling rates through a three-dimensional space coordinate projection and time interpolation algorithm, and constructing an incidence matrix; performing feature extraction and fusion on the physiological signal data and the operation track data based on the incidence matrix to generate a fusion feature vector; processing the fusion feature vector by adopting a hierarchical evaluation mechanism including basic hierarchical evaluation and optimized hierarchical evaluation, and outputting a comprehensive capability evaluation result; triggering a corresponding dynamic intervention measure according to a comprehensive capability evaluation result; therefore, multi-modal fusion is carried out on the physiological signal data and the operation track data, a comprehensive capability evaluation result is accurately output, and intervention and error correction are carried out in time in the practical operation process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of training evaluation, and particularly relates to a training evaluation method, system and device based on multi-modal data fusion and a storage medium. BACKGROUND

[0002] In the medical health field, especially in the surgical skill training for professionals, accurately evaluating the training effect and providing real-time feedback are crucial to improving the training quality. Traditional training methods mostly rely on manual recording of operation logs, paper examination and other discrete information, which leads to fragmented data collection and limits the training effect evaluation to single surface indicators, failing to reflect the real ability of the trainee.

[0003] In related technologies, an electronic system is introduced to digitally store the theoretical test scores and operation duration of the trainee, but this method is also limited by single-modal data collection and evaluation, failing to combine the operation trajectory of the trainee in the operation space, failing to establish the correlation between the operation trajectory and the physiological signals of the trainee, and the feedback is seriously delayed, failing to correct the wrong operation in the operation process in time, and failing to meet the demand of medical training for precise evaluation and real-time intervention. SUMMARY

[0004] The application provides a training evaluation method, system, device and storage medium based on multi-modal data fusion, which can perform multi-modal fusion on physiological signal data and operation trajectory data, accurately output comprehensive ability evaluation results, and timely intervene and correct errors in the operation process.

[0005] To solve the above technical problems, in a first aspect, the application provides a training evaluation method based on multi-modal data fusion, comprising the following steps:

[0006] Collecting physiological signal data and operation trajectory data generated in the training process;

[0007] Performing space-time alignment on the physiological signal data and the operation trajectory data with different sampling rates through three-dimensional space coordinate projection and time interpolation algorithm, and constructing a correlation matrix;

[0008] Extracting and fusing features of the physiological signal data and the operation trajectory data based on the correlation matrix to generate a fusion feature vector;

[0009] Processing the fusion feature vector by using a hierarchical evaluation mechanism including basic level evaluation and optimization level evaluation, and outputting a comprehensive ability evaluation result;

[0010] Triggering corresponding dynamic intervention measures according to the comprehensive ability evaluation result.

[0011] As a further improvement of the present application, the physiological signal data at least includes physiological sign data and operation force data during the training process of the trainer.

[0012] The operation trajectory data at least includes motion trajectory data of the operation instrument during the training process.

[0013] As a further improvement of the present application, the space-time alignment of the physiological signal data and the operation trajectory data with different sampling rates is performed by a three-dimensional space coordinate projection and a time interpolation algorithm, and a correlation matrix is constructed, including:

[0014] A three-dimensional space coordinate system is established.

[0015] Spatial mapping features are extracted from the physiological signal data and the operation trajectory data.

[0016] The spatial mapping features are projected into the three-dimensional space coordinate system.

[0017] The projected physiological signal data and operation trajectory data are time-synchronized based on a time interpolation algorithm, so that the time deviation between the synchronized data does not exceed a preset deviation.

[0018] The space-time aligned physiological signal data and operation trajectory data are correlated and mapped to construct a correlation matrix reflecting the relationship between the physiological state of the trainer and the operation instrument.

[0019] As a further improvement of the present application, the physiological signal data and the operation trajectory data are extracted and fused based on the correlation matrix to generate a fused feature vector, including:

[0020] Based on the space-time aligned physiological signal data and operation trajectory data, the following operations are performed:

[0021] Physiological sign features and operation force features in the physiological signal data are extracted.

[0022] Motion trajectory features in the operation trajectory data are extracted.

[0023] Spatial position deviation features in the correlation matrix are extracted.

[0024] The physiological sign features, operation force features, motion trajectory features, and spatial position deviation features are integrated to generate an initial feature vector.

[0025] The initial feature vector is subjected to dimension compression processing to generate the fused feature vector.

[0026] As a further improvement of the present application, the fusion feature vector at least includes a motion trajectory feature vector, an operation force feature vector, a physiological sign feature vector, and a spatial position deviation feature vector.

[0027] As a further improvement of the present application, the fusion feature vector is processed by using a hierarchical evaluation mechanism including a basic level evaluation and an optimization level evaluation, and an integrated capability evaluation result is output, including:

[0028] Based on the motion trajectory feature vector and the spatial position deviation feature vector, a matching degree and a deviation degree between a training operation trajectory and a standard trajectory are calculated;

[0029] The matching degree and the deviation degree are weighted and fused to generate a basic level evaluation result;

[0030] Based on the operation force feature vector, a stability index of an operation instrument holding is calculated, and a risk evaluation value is predicted based on the physiological sign feature vector;

[0031] The stability index and the risk evaluation value are weighted and fused to generate an optimization level evaluation result;

[0032] The integrated capability evaluation result is output in combination with the basic level evaluation result and the optimization level evaluation result.

[0033] As a further improvement of the present application, the corresponding dynamic intervention measures are triggered according to the integrated capability evaluation result, including:

[0034] When the deviation degree in the basic level evaluation result exceeds a first preset threshold, a warning intervention is triggered;

[0035] When the stability index in the optimization level evaluation result is lower than a second preset threshold, or when the risk evaluation value exceeds a third preset threshold, a correction intervention is triggered;

[0036] When the integrated capability evaluation result is lower than a fourth preset threshold, a targeted reinforcement intervention is triggered.

[0037] In a second aspect, the present application provides a training evaluation system based on multi-modal data fusion, which includes:

[0038] An acquisition unit is configured to acquire physiological signal data and operation trajectory data generated during a training process;

[0039] A construction unit is configured to perform space-time alignment on the physiological signal data and the operation trajectory data with different sampling rates by using a three-dimensional space coordinate projection and a time interpolation algorithm, and construct a correlation matrix;

[0040] a fusion unit configured to perform feature extraction and fusion on the physiological signal data and the operation trajectory data based on the correlation matrix to generate a fusion feature vector;

[0041] an evaluation unit configured to process the fusion feature vector using a hierarchical evaluation mechanism including a basic level evaluation and an optimization level evaluation to output a comprehensive ability evaluation result;

[0042] a triggering unit configured to trigger a corresponding dynamic intervention measure according to the comprehensive ability evaluation result.

[0043] In a third aspect, the present application provides a computer device, which comprises a processor and a memory coupled to the processor, and the memory stores a computer program, and the computer program is executed by the processor to cause the processor to perform the steps of the training evaluation method based on multi-modal data fusion in any one of the above aspects.

[0044] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the training evaluation method based on multi-modal data fusion in any one of the above aspects.

[0045] Compared with the prior art, the training evaluation method, system, device and storage medium based on multi-modal data fusion provided by the present application avoid data fragmentation by collecting physiological signal data and operation trajectory data in real time; the spatio-temporal alignment of the physiological signal data and the operation trajectory data is performed by a three-dimensional space coordinate projection and a time interpolation algorithm, a correlation matrix is constructed to ensure that heterogeneous data with different sampling rates can be correlated under the same spatio-temporal reference; multi-dimensional features are extracted and fused based on the correlation matrix to generate a fusion feature vector, and a hierarchical evaluation mechanism is used to perform comprehensive evaluation from the basic operation and optimization operation levels, thereby outputting a comprehensive evaluation result reflecting the ability of the trainer, ensuring the comprehensiveness and accuracy of the evaluation; finally, corresponding dynamic intervention measures are triggered according to the evaluation result, and real-time feedback and correction of operation deviation are performed for different levels of problems, thereby improving the pertinence and efficiency of the training, and providing a precise and real-time evaluation and intervention scheme for the training in the medical and health field. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0047] Figure 1A flowchart of the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 1.

[0048] Figure 2 A flowchart of constructing the correlation matrix in the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 2.

[0049] Figure 3 A flowchart of generating the fusion feature vector in the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 3.

[0050] Figure 4 A flowchart of outputting the comprehensive capability evaluation result in the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 4.

[0051] Figure 5 A structural schematic diagram of outputting the comprehensive capability evaluation result in the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 5.

[0052] Figure 6 A flowchart of triggering the dynamic intervention measure in the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 6.

[0053] Figure 7 A structural schematic diagram of the training evaluation system based on multi-modal data fusion provided by the embodiment of the present application is shown in FIG. 7.

[0054] Figure 8 A structural schematic diagram of the computer device provided by the embodiment of the present application is shown in FIG. 8.

[0055] Figure 9 A structural schematic diagram of the storage medium provided by the embodiment of the present application is shown in FIG. 9. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, and are not used to limit the embodiments of the present application.

[0057] In the description of the embodiments of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise explicitly and specifically limited. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative positional relationship, movement condition, etc. between the components in a certain posture (as shown in the drawings), if the certain posture changes, the directional indications also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.

[0058] To provide a more detailed and complete description of the present disclosure, the following provides illustrative descriptions of implementation methods and specific embodiments of the present disclosure; however, these descriptions are not intended to be the only way to implement or use the embodiments of the present disclosure. The implementation methods cover features of various specific embodiments, as well as the method steps and sequences for constructing and operating these specific embodiments. However, other specific embodiments may also be used to achieve the same or equivalent functionality and step sequences.

[0059] In the embodiments of this application, the terms "exemplary," "in some embodiments," and "in another embodiment" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" should not be construed as preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner.

[0060] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0061] In the healthcare sector, especially in surgical skills training for professionals, accurately evaluating training effectiveness and providing real-time feedback are crucial to improving training quality. Traditional training methods rely on discrete information such as manual operation logs and paper-based assessments, resulting in fragmented data collection. Training effectiveness evaluation is limited to a single, superficial metric that fails to reflect the trainee's true capabilities.

[0062] Related technologies have introduced electronic systems to digitally store trainees' theoretical test scores and operation time. However, this method is also limited to single-modality data collection and evaluation. It cannot be combined with the trainees' actual operation trajectory in the surgical space, and cannot establish a correlation between the operation trajectory and data such as the trainees' physiological signals. In addition, the feedback delay is serious, and incorrect operations cannot be corrected in time during the actual operation. It is difficult to meet the medical training needs for accurate evaluation and real-time intervention.

[0063] In view of this, please refer to Figures 1-9 The embodiments of the present application provide a training evaluation method, system, device and storage medium based on multimodal data fusion, which can perform multimodal fusion of physiological signal data and operation trajectory data, accurately output comprehensive ability evaluation results, and intervene and correct errors in a timely manner during actual operation.

[0064] It can be understood that in the surgical skill training in the medical health field, accurate assessment of the operation level and real-time feedback are particularly important to improve the operation ability of the trainer. The existing evaluation methods are mostly single mode information such as recording theoretical test scores and operation time, so that the training evaluation effect only stays at the surface index statistics level such as operation time and step completion degree, and is limited to a single surface index. It is particularly important to evaluate the operation ability of the trainer from a comprehensive perspective and improve the operation ability of the trainer. Therefore, the following will continue to take the surgical skill training in the medical health field as an example to describe the training evaluation method based on multi-modal data fusion provided by the present application in detail.

[0065] Please refer to Figure 1 The flowchart of the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application, the detection method comprises the following steps:

[0066] Step S1: Collecting physiological signal data and operation trajectory data generated in the training process;

[0067] As an optional implementation, the physiological signal data at least includes physiological sign data of the trainer in the training process, and operation force data;

[0068] The operation trajectory data at least includes motion trajectory data of the operating instrument in the training process.

[0069] In the embodiment of the present application, the physiological signal data and operation trajectory data generated by the trainer in the training process need to be collected in real time.

[0070] Among them, the physiological signal data includes physiological sign data that can reflect the body state of the trainer in the training process, and operation force data generated when the trainer interacts with the surgical instrument. The operation trajectory data at least includes the motion trajectory coordinates of the operating instrument, i.e. the surgical instrument, in the three-dimensional space.

[0071] Since the physiological signal data can reflect the cognitive load, emotional state and other characteristics of the trainer, the operation trajectory data can further reflect the operation trajectory of the surgical instrument in the training process and the action accuracy, so as to convert the surface recorded data such as operation time and step completion degree into three-dimensional data by combining the physiological signal data and the operation trajectory data, so as to provide a multi-dimensional data basis for subsequent accurate evaluation.

[0072] In an optional embodiment, the trainer can collect PPG (Photoplethysmography) signals and EDA (Electrodermal Activity) signals in real time by wearing a biological sign monitoring bracelet, and the PPG signals reflect the heart rate variability of the trainer during the training process, and the EDA signals reflect the tension of the trainer during the training process.

[0073] At the same time, the trainer also needs to wear a force feedback glove to collect the change value of the force of the training personnel holding the surgical instrument during the operation, and when the training personnel wearing the force feedback glove contacts the surgical instrument, the pressure sensor arranged on the force feedback glove will synchronously collect the actual force direction and the specific force value.

[0074] Further, the trainer can also wear an AR head-mounted device to capture the gaze point coordinates of the trainer in real time through the eye tracking module of the AR head-mounted device, or extract the pupil diameter change rate of the trainer, and it is also feasible to obtain other types of physiological signal data through other means, which is not limited by the present application.

[0075] In the embodiment of the present application, the operation trajectory data collected above should at least include the motion trajectory data of the operating instrument during the training process, for example, when the trainer applies a medical simulator for simulation training, the three-dimensional motion trajectory of the surgical instrument, preferably the tip of the surgical instrument, can be collected in real time through an API (Application Programming Interface) interface, and the trigger time point of the action of the trainer through the surgical machine for clamping, cutting, and large amplitude shaking is recorded synchronously, so as to establish the association between the physiological signal data and the operation trajectory data in the subsequent process, and provide data support for the subsequent hierarchical evaluation mechanism.

[0076] It can be understood that the physiological signal data provided by the present application is not limited to the physiological sign data and the operation force data of the trainer during the training process provided above, and the operation trajectory data is not limited to the motion trajectory data of the operating instrument during the training process, and other types of physiological signal data and operation trajectory data such as body temperature data are also feasible, and those skilled in the art can make corresponding adjustments according to the specific training scene.

[0077] In a specific embodiment, Zhang needs to perform a resection operation training in a medical simulator, and he needs to wear at least a biological sign monitoring bracelet, a force feedback surgical glove, and connect the simulated operation table to the training system before the training starts, so that the training system can collect three-dimensional motion data of the tip of the surgical instrument in real time.

[0078] Optionally, Zhang can also wear an AR head-mounted device, and the AR head-mounted device collects the coordinates of the fixation points of Zhang during the training process and the pupil diameter change rate. Through the above preparation work, the physiological signal data and operation trajectory data of Zhang in the simulation surgery are comprehensively obtained.

[0079] Step S2: The physiological signal data and the operation trajectory data with different sampling rates are spatio-temporally aligned through a three-dimensional space coordinate projection and a time interpolation algorithm, and a correlation matrix is constructed.

[0080] As an optional implementation, please refer to Figure 2 The flowchart for constructing the correlation matrix in the training evaluation method based on multi-modal data fusion provided by the embodiment of the application is as follows: the physiological signal data and the operation trajectory data with different sampling rates are spatio-temporally aligned through a three-dimensional space coordinate projection and a time interpolation algorithm, and a correlation matrix is constructed, which includes:

[0081] Step S20: A three-dimensional space coordinate system is established.

[0082] Step S21: Spatial mapping features are extracted from the physiological signal data and the operation trajectory data.

[0083] Step S22: The spatial mapping features are projected into the three-dimensional space coordinate system.

[0084] Step S23: The physiological signal data and the operation trajectory data after projection are time-synchronized based on a time interpolation algorithm, so that the time deviation between the synchronized data does not exceed a preset deviation.

[0085] Step S24: The physiological signal data and the operation trajectory data after spatio-temporal alignment are correlated and mapped, and a correlation matrix reflecting the relationship between the physiological state of the trainer and the operating instrument is constructed.

[0086] In the embodiment of the application, the three-dimensional space coordinate system needs to be established based on the operation platform of the medical simulator. The three-dimensional space coordinate system can be established with the center position of the operation platform as the coordinate origin, or with one of the end angles of the operation platform as the coordinate origin. The above-provided establishment methods of the three-dimensional space coordinate system are all feasible, as long as a unified space reference system can be provided for the subsequent physiological signal data and operation trajectory data, and the projection into the unified space reference system is ensured. Those skilled in the art should know this.

[0087] Further, the spatial mapping features are extracted from the physiological signal data and the operation trajectory data, that is, the spatial mapping features related to the spatial position are extracted from the physiological signal data and the operation trajectory data.

[0088] For example, the data related to the spatial position in the physiological signal data at least includes the coordinates of the visual focus points in the training process, and the data related to the spatial position in the operation trajectory data at least includes the coordinates of the motion trajectory of the surgical instrument in the training process.

[0089] After the spatial mapping features are extracted from the physiological signal data and the operation trajectory data, the spatial mapping features need to be projected into a three-dimensional coordinate system, that is, the coordinates of the visual focus points and the coordinates of the motion trajectory of the surgical instrument are projected into a three-dimensional coordinate system established in advance, to ensure that the spatial mapping features projected into the three-dimensional coordinate system are only related to the operation scene and can reflect the features related to the spatial position, avoiding the introduction of other invalid data affecting the subsequent processing efficiency.

[0090] Further, since the collection rate of the AR head-mounted device and the collection rate of the motion trajectory of the surgical instrument can not be the same, time interpolation algorithm is needed to synchronize the projected physiological signal data and operation trajectory data in time, so that the time deviation between the synchronized data does not exceed the preset deviation.

[0091] For example, the collection rate of a certain AR head-mounted device is 30Hz, and the collection rate of the motion trajectory of the surgical instrument is 100HZ, so based on the size relationship between the collection rate of the AR head-mounted device and the collection rate of the motion trajectory of the surgical instrument, the relatively higher collection rate of the motion trajectory of the surgical instrument 100HZ is used as a reference, and the missing time points of the visual focus point data with a relatively lower collection rate of 30Hz are completed by using a cubic spline interpolation algorithm.

[0092] For example, the instrument trajectory has data at time points t=10ms, 20ms…, while the visual focus point data is only recorded at t=0ms, 33ms, 66ms…, so a cubic spline interpolation algorithm is used to synchronize the two in time, to ensure that the visual focus point data generates an equal interval of 10ms after interpolation. Synchronized time series, realize the time synchronization between the motion trajectory data of the surgical instrument and the visual focus point data.

[0093] Further, NTP (Network Time Protocol) can also be used to ensure that the time deviation between the motion trajectory data of the surgical instrument and the visual focus point data after synchronization does not exceed the preset deviation, such as not exceeding 1ms, so as to ensure that the visual focus point coordinates and the motion trajectory coordinates of the surgical instrument under each timestamp are in a one-to-one correspondence, avoiding the situation that the visual focus point coordinates and the motion trajectory coordinates of the surgical instrument cannot be corresponded due to time errors, affecting the subsequent evaluation effect.

[0094] In the embodiment of the present application, after the time synchronization between the physiological signal data and the operation trajectory data, specifically the movement trajectory data of the surgical instrument and the visual focus point data, is completed, the movement trajectory data of the surgical instrument and the visual focus point data need to be associated and mapped to construct an association matrix reflecting the relationship between the physiological state of the trainer and the operation instrument.

[0095] Specifically, the synchronized visual focus point coordinates and the movement trajectory coordinates of the surgical instrument can be associated and mapped with the timestamp as the index to form an association matrix, thereby constructing an association matrix containing the relationship between the visual focus point of the trainer and the movement trajectory of the surgical instrument.

[0096] Step S3: performing feature extraction and fusion on the physiological signal data and the operation trajectory data based on the association matrix to generate a fusion feature vector;

[0097] As an optional implementation, please refer to Figure 3 The flowchart for generating a fusion feature vector in the training evaluation method based on multi-modal data fusion provided by the embodiment of the present application is shown in the above, and the feature extraction and fusion on the physiological signal data and the operation trajectory data based on the association matrix to generate a fusion feature vector includes:

[0098] Based on the spatio-temporal aligned physiological signal data and operation trajectory data, the following operations are performed:

[0099] Step S30: extracting physiological sign features and operation force features in the physiological signal data;

[0100] Step S31: extracting movement trajectory features in the operation trajectory data;

[0101] Step S32: extracting spatial position deviation features in the association matrix;

[0102] Step S33: integrating the physiological sign features, operation force features, movement trajectory features and spatial position deviation features to generate an initial feature vector;

[0103] Step S34: performing dimension compression processing on the initial feature vector to generate the fusion feature vector.

[0104] In the embodiment of the present application, the spatio-temporal aligned physiological signal data and operation trajectory data need to be subjected to corresponding operations. Since the physiological signal data includes physiological sign data and operation force data, the corresponding physiological sign features need to be extracted from the physiological sign data, and the corresponding operation force features need to be extracted from the operation force data.

[0105] Specifically, since the physiological sign data is collected in real time by the trainees wearing a biological sign monitoring bracelet, the EDA (Electrodermal Activity) signal can be subjected to wavelet transformation and feature extraction to extract the energy proportion of the corresponding autonomic nervous activity in the 0.2-1.5Hz frequency band. An increase in the energy in this frequency band indicates an increase in the trainee's tension, which is extracted as a physiological sign feature.

[0106] Similarly, the pressure gradient change rate, that is, the change value of pressure per unit time, is calculated from the operation force data to reflect the stability of the trainee's force during the operation. The spatial change entropy value of the pressure value can also be calculated from the operation force data. The higher the entropy value, the more balanced the force applied by the trainee during the operation. The degree of pressure dispersion between different fingers is quantified and extracted as the operation force feature.

[0107] In this way, the required physiological sign characteristics and operation force characteristics can be extracted from the physiological signal data.

[0108] Optionally, the trainee's visual focus hotspot areas can be identified based on the visual focus point data so as to subsequently evaluate the trainee's attention. This application will not go into details here.

[0109] As an optional implementation, it is also necessary to extract motion trajectory features from the operation trajectory data, specifically including time domain analysis of the operation trajectory data of the surgical instrument, and calculating the motion trajectory features such as the movement speed, acceleration, trajectory curvature, etc. of the surgical instrument tip during the operation.

[0110] Furthermore, it is necessary to extract the spatial position deviation features in the association matrix. Since the motion trajectory data and visual focus point data of the surgical instrument have been time-synchronized in step S2, this application calculates the spatial deviation between the motion trajectory of the surgical instrument tip and the visual focus point based on the established three-dimensional spatial coordinate system, and uses the spatial deviation as the spatial position deviation feature to quantify the degree of matching between the trainee's attention and the surgical instrument, so as to subsequently accurately evaluate the trainee's hand-eye coordination ability; as for how to calculate the spatial deviation between the two based on the motion trajectory of the surgical instrument tip and the visual focus point, it can be calculated based on the common three-dimensional spatial distance formula. Since this calculation method has been widely used in the field of vision, this application will not elaborate on it.

[0111] In an embodiment of the present application, the extracted physiological sign features, operation force features, motion trajectory features, and spatial position deviation features are integrated to generate an initial feature vector, forming an initial feature vector of multimodal data including physiological, motion, and spatial data.

[0112] On this basis, since the initial feature vector is fused with multi-dimensional feature vectors, the application adopts a dimension reduction algorithm such as PCA (Principal Component Analysis) or an autoencoder to perform dimension compression on the initial feature vector, thereby removing redundant features in the initial feature vector, reducing the subsequent calculation complexity, while retaining the associated features between different modalities, so that the subsequent hierarchical evaluation mechanism can perform rapid evaluation processing, thereby timely intervention.

[0113] It can be understood that the fused feature vector after dimension compression processing at least includes a motion trajectory feature vector, an operation force feature vector, a physiological sign feature vector, and a spatial position deviation feature vector; wherein the motion trajectory feature vector corresponds to the motion trajectory feature, the operation force feature vector corresponds to the operation force feature, the physiological sign feature vector corresponds to the physiological sign feature, and the spatial position deviation feature vector corresponds to the spatial position deviation feature; of course, other types of feature vectors required according to actual needs can also be added for subsequent evaluation, which is not limited by the application.

[0114] Step S4: processing the fused feature vector by using a hierarchical evaluation mechanism including a basic level evaluation and an optimization level evaluation, and outputting a comprehensive ability evaluation result;

[0115] As an optional implementation, please refer to Figure 4 The flowchart for outputting a comprehensive ability evaluation result in the training evaluation method based on multi-modal data fusion provided by the embodiment of the application, the above processing of the fused feature vector by using a hierarchical evaluation mechanism including a basic level evaluation and an optimization level evaluation, and outputting a comprehensive ability evaluation result, includes:

[0116] Step S40: calculating the matching degree and the deviation degree between the training operation trajectory and the standard trajectory based on the motion trajectory feature vector and the spatial position deviation feature vector;

[0117] Step S41: weighting and fusing the matching degree and the deviation degree to generate a basic level evaluation result;

[0118] Step S42: calculating a stability index of the operation instrument holding based on the operation force feature vector; and pre-judging a risk evaluation value based on the physiological sign feature vector;

[0119] Step S43: weighting and fusing the stability index and the risk evaluation value to generate an optimization level evaluation result;

[0120] Step S44: combining the basic level evaluation result and the optimization level evaluation result to output the comprehensive ability evaluation result.

[0121] In the embodiments of the present application, please refer to Figure 5 The structure diagram for outputting the comprehensive ability evaluation result in the training evaluation method based on multi-modal data fusion provided in the embodiments of the present application can be based on the motion trajectory feature vector, adopt the DTW (Dynamic Time Warping) algorithm, compare the actual motion trajectory of the trainer with the standard motion trajectory after time sequence alignment, and thus obtain the matching degree between the training operation trajectory and the standard trajectory.

[0122] In an optional embodiment, the deviation degree between the visual focus point of the trainer and the instrument trajectory can be calculated based on the spatial position deviation feature vector, and thus the deviation degree between the training operation trajectory and the standard trajectory is obtained.

[0123] Further, the matching degree and the deviation degree obtained above are weighted and fused to generate the basic level evaluation result. The specific calculation method and calculation formula of the matching degree and the deviation degree are not described in detail.

[0124] In a specific embodiment provided in the present application, the basic reward function R1 = Σ (X x 0.6 + Y x 0.3) is adopted for calculation; wherein X is the matching degree, 0.6 is the matching degree weight, Y is the deviation degree, and 0.3 is the deviation degree weight. The respective calculation is performed for each time node within the entire training time, so as to integrate the trajectory matching degree and the spatial deviation degree into a unified index, and avoid one-sided evaluation by a single index.

[0125] As a further improvement of the present application, the operation instrument holding stability index is calculated by using the operation force feature vector mentioned above, so as to directly feed back the fine degree of instrument operation, and the risk evaluation value is obtained by using the physiological sign feature and the model to predict the possible operation risk, so as to evaluate the potential operation failure.

[0126] It can be understood that the stability index can be calculated by analyzing the pressure gradient change rate and the pressure distribution entropy value in the force feedback data, and the mapping relationship between the physiological state and the operation risk is established by using the machine learning models such as logistic regression and random forest, so as to realize the risk prediction to obtain the required risk evaluation value. The above calculation and prediction methods all belong to the conventional technical means in the field of signal processing and machine learning, and the existing technology can completely realize it.

[0127] Therefore, the present application does not involve the improvement of the specific calculation method or the prediction model, but directly uses the existing calculation method and prediction model to obtain the required stability index and risk evaluation value. Therefore, the specific implementation details in the above existing technology are not described in detail.

[0128] Further, the stability index and the risk assessment value obtained above are weighted and fused to generate an optimized hierarchical assessment result.

[0129] In one specific embodiment, a high-level reward function R2 = å(M x 0.6 + N x 0.3) is used for calculation; wherein M is the stability index, 0.6 is the stability index weight, Y is the risk assessment value, 0.3 is the risk assessment value weight, and the same calculation is performed for each time node in the entire training time, and then the stability index and the risk assessment value of each time node are weighted and fused to assess the comprehensive ability of the trainer in operation accuracy and risk control in real time, so as to identify potential safety risks.

[0130] On this basis, the application needs to combine the basic hierarchical assessment result and the optimized hierarchical assessment result to output a comprehensive ability assessment result.

[0131] For example, the calculation can be performed by the formula Q = p x R1 + (1-p) R2, wherein Q is the comprehensive ability assessment value, p is the dynamic weight, R1 is the basic hierarchical assessment value, R2 is the optimized hierarchical assessment value, and the specific value of p is not limited further herein.

[0132] It can be understood that the hierarchical assessment method combining the basic hierarchical assessment and the optimized hierarchical assessment can effectively avoid the situation that only the operation process is focused on and the potential risk is ignored, or the potential risk is excessively emphasized and the basic operation process is ignored.

[0133] Step S5: triggering a corresponding dynamic intervention measure according to the comprehensive ability assessment result.

[0134] As an optional implementation, please refer to Figure 6 The flowchart of triggering a dynamic intervention measure in the training assessment method based on multi-modal data fusion provided by the embodiment of the application, the above-mentioned triggering a corresponding dynamic intervention measure according to the comprehensive ability assessment result, comprises:

[0135] Step S50: triggering a warning intervention when the deviation degree in the basic hierarchical assessment result exceeds a first preset threshold;

[0136] Step S51: triggering a correction intervention when the stability index in the optimized hierarchical assessment result is lower than a second preset threshold, or when the risk assessment value exceeds a third preset threshold;

[0137] Step S52: triggering a targeted reinforcement intervention when the comprehensive ability assessment result is lower than a fourth preset threshold.

[0138] In the embodiments of the present application, the deviation degree in the basic level evaluation result is monitored in real time, when the deviation degree exceeds the first preset threshold, the current simulator scene is frozen, the operation track of the trainer is compared with the standard track in real time through holographic projection, and the operation track of the trainer is set to be different from the standard track in color, so as to provide the visual standard track for the trainer to improve.

[0139] Further, when the stability index in the optimization level evaluation is lower than the second preset threshold, it indicates that the trainer may currently have the situation of too large fluctuation of instrument holding force or unbalanced force, or when the risk evaluation value exceeds the third preset threshold, it indicates that the trainer currently has potential risks and may have the situation of prediction error, so that the correction intervention such as physical vibration and visual prompt can be triggered to remind the trainer in the early stage of the problem before causing serious errors, and the repeated reinforcement of the wrong operation is avoided.

[0140] Further, when the comprehensive ability evaluation value of the comprehensive ability evaluation result is lower than the fourth preset threshold, targeted reinforcement intervention can be triggered, the reinforcement scene is dynamically generated based on the scene of the error frequency, and the training frequency of the scene where the error occurs is improved subsequently.

[0141] It should be noted that the specific values of the first preset threshold, the second preset threshold, the third preset threshold and the fourth preset threshold are not limited in the present application; the setting of the thresholds can be flexibly adjusted according to the actual training scene or training demand in principle, for example, the first preset threshold and the second preset threshold can be set more strictly for fine operation scenes, and the triggering conditions of the thresholds can be appropriately relaxed for new trainee training; since the setting of the thresholds belongs to the routine selection of the person skilled in the art based on the specific application scene, the present application will not be described in detail here.

[0142] It can be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0143] The training evaluation method based on multi-modal data fusion provided in the application avoids fragmented data collection by collecting physiological signal data and operation trajectory data in real time during the training process; the physiological signal data and the operation trajectory data are spatio-temporally aligned by three-dimensional space coordinate projection and time interpolation algorithm, an association matrix is constructed, and it is ensured that heterogeneous data with different sampling rates can be associated under the same spatio-temporal reference; multi-dimensional features are extracted and fused based on the association matrix to generate a fusion feature vector, a hierarchical evaluation mechanism is used to comprehensively evaluate from the basic operation and optimization operation levels, and then a comprehensive evaluation result reflecting the ability of the trainer is output, thereby ensuring the comprehensiveness and accuracy of the evaluation; finally, corresponding dynamic intervention measures are triggered according to the evaluation result, real-time feedback and correction of operation deviation are carried out for different levels of problems, thereby improving the pertinence and efficiency of the training, and providing a precise and real-time evaluation and intervention scheme for the training in the medical and health field.

[0144] Based on the above training evaluation method based on multi-modal data fusion, the application provides a training evaluation system based on multi-modal data fusion, please refer to Figure 7 The training evaluation system based on multi-modal data fusion provided in the embodiments of the application has the structure shown in the figure, and the training evaluation system comprises an acquisition unit, a construction unit, a fusion unit, an evaluation unit and a triggering unit;

[0145] The acquisition unit is configured to acquire physiological signal data and operation trajectory data generated during the training process.

[0146] The construction unit is configured to perform spatio-temporal alignment on the physiological signal data and the operation trajectory data with different sampling rates by three-dimensional space coordinate projection and time interpolation algorithm, and construct an association matrix.

[0147] The fusion unit is configured to extract and fuse features of the physiological signal data and the operation trajectory data based on the association matrix, and generate a fusion feature vector.

[0148] The evaluation unit is configured to process the fusion feature vector by using a hierarchical evaluation mechanism comprising basic level evaluation and optimization level evaluation, and output a comprehensive ability evaluation result.

[0149] The triggering unit is configured to trigger corresponding dynamic intervention measures according to the comprehensive ability evaluation result.

[0150] As an optional implementation, the physiological signal data at least includes physiological sign data and operation force data of the trainer during the training process.

[0151] The operation trajectory data at least includes motion trajectory data of the operation instrument during the training process.

[0152] As an optional implementation, the time-space alignment of the physiological signal data and the operation trajectory data with different sampling rates by the three-dimensional space coordinate projection and the time interpolation algorithm, and the construction of the correlation matrix, comprise:

[0153] A three-dimensional space coordinate system is established.

[0154] Spatial mapping features are extracted from the physiological signal data and the operation trajectory data.

[0155] The spatial mapping features are projected into the three-dimensional space coordinate system.

[0156] The physiological signal data and the operation trajectory data after projection are time-synchronized based on the time interpolation algorithm, so that the time deviation between the synchronized data does not exceed a preset deviation.

[0157] The time-space aligned physiological signal data and operation trajectory data are correlated and mapped to construct a correlation matrix reflecting the relationship between the physiological state of the trainer and the operation instrument.

[0158] As an optional implementation, the feature extraction and fusion of the physiological signal data and the operation trajectory data based on the correlation matrix to generate a fusion feature vector, comprise:

[0159] Based on the time-space aligned physiological signal data and operation trajectory data, the following operations are performed:

[0160] Physiological sign features and operation force features in the physiological signal data are extracted.

[0161] Motion trajectory features in the operation trajectory data are extracted.

[0162] Spatial position deviation features in the correlation matrix are extracted.

[0163] The physiological sign features, operation force features, motion trajectory features, and spatial position deviation features are integrated to generate an initial feature vector.

[0164] The initial feature vector is subjected to dimension compression processing to generate the fusion feature vector.

[0165] As an optional implementation, the fusion feature vector at least includes a motion trajectory feature vector, an operation force feature vector, a physiological sign feature vector, and a spatial position deviation feature vector.

[0166] As an optional implementation, the fusion feature vector is processed by using a hierarchical evaluation mechanism comprising a basic level evaluation and an optimization level evaluation, and an integrated ability evaluation result is output, comprising:

[0167] Calculate a matching degree and a deviation degree between the training operation trajectory and the standard trajectory based on the motion trajectory feature vector and the spatial position deviation feature vector;

[0168] Weightedly fuse the matching degree and the deviation degree to generate a basic level evaluation result;

[0169] Calculate a stability index of the operation instrument holding based on the operation force feature vector, and predict a risk evaluation value based on the physiological sign feature vector;

[0170] Weightedly fuse the stability index and the risk evaluation value to generate an optimized level evaluation result;

[0171] Output the comprehensive capability evaluation result in combination with the basic level evaluation result and the optimized level evaluation result.

[0172] As an optional implementation, triggering the corresponding dynamic intervention measures according to the comprehensive capability evaluation result comprises:

[0173] Triggering a warning intervention when the deviation degree in the basic level evaluation result exceeds a first preset threshold;

[0174] Triggering a correction intervention when the stability index in the optimized level evaluation result is lower than a second preset threshold, or when the risk evaluation value exceeds a third preset threshold;

[0175] Triggering a targeted reinforcement intervention when the comprehensive capability evaluation result is lower than a fourth preset threshold.

[0176] For other details of the technical solutions of the units in the training evaluation system based on multi-modal data fusion provided in the above embodiments, refer to the description in the training evaluation method based on multi-modal data fusion in the above embodiments, which will not be repeated here.

[0177] It should be noted that each of the embodiments in the specification adopts a progressive manner for description, and each embodiment focuses on the different places from other embodiments, and the same and similar parts of each embodiment can be referred to each other. For system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0178] Please refer to Figure 8 The computer device 80 provided in the embodiments of the present application includes a processor 81 and a memory 82 coupled with the processor 81.

[0179] The memory 82 stores a computer program, and the computer program is executed by the processor 81 to enable the processor 81 to perform the steps of the training evaluation method based on multi-modal data fusion in the above embodiments.

[0180] The processor 81 can also be referred to as a CPU (Central Processing Unit). The processor 81 can be an integrated circuit chip having a processing capability of signals. The processor 81 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0181] Please refer to Figure 9 The computer readable storage medium of the embodiments of the present application stores a computer program 90, and the computer program 90 is executed by the processor to implement the above-mentioned artificial intelligence-based actuarial method. The computer program 90 can be stored in the above-mentioned storage medium in the form of a software product, including a plurality of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes, or a computer, a server, a mobile phone, a tablet computer, etc. The server can be a standalone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms, etc. Basic cloud computing services.

[0182] In several embodiments provided in the present application, it should be understood that the disclosed terminal, device and method can be implemented in other manners. For example, the embodiments of the device described above are merely schematic; for example, the division of the units is only a logical function division; there can be another division manner for the actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0183] In addition, each function unit in the embodiments of the present application can be integrated into a processing unit, or each unit can exist alone physically, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit. The above is merely an implementation manner of the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent flow transformation based on the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

[0184] The above implementation manners are merely exemplary implementation manners adopted for illustrating the principles of the embodiments of the present application, and the embodiments of the present application are not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the embodiments of the present application, and these modifications and improvements are also considered to be within the protection scope of the embodiments of the present application.

Claims

1. A training evaluation method based on multimodal data fusion, characterized in that: The following steps are involved: Collect physiological signal data and operation trajectory data generated during training; Performing spatiotemporal alignment on the physiological signal data and the operation trajectory data at different sampling rates through three-dimensional space coordinate projection and time interpolation algorithms to construct a correlation matrix; Performing feature extraction and fusion on the physiological signal data and the operation trajectory data based on the correlation matrix to generate a fusion feature vector; Processing the fused feature vector using a hierarchical evaluation mechanism including a basic level evaluation and an optimized level evaluation to output a comprehensive capability evaluation result; Corresponding dynamic intervention measures are triggered according to the comprehensive capability assessment results.

2. The training evaluation method based on multimodal data fusion according to claim 1, characterized in that: The physiological signal data at least includes physiological sign data and operation force data of the trainee during training; The operation trajectory data at least includes motion trajectory data of the operating instrument during training.

3. The training evaluation method based on multimodal data fusion according to claim 1, characterized in that: The method of performing spatiotemporal alignment on the physiological signal data and the operation trajectory data of different sampling rates by using a three-dimensional space coordinate projection and a time interpolation algorithm to construct a correlation matrix includes: Establish a three-dimensional space coordinate system; extracting spatial mapping features from the physiological signal data and the operation trajectory data; Projecting the spatial mapping feature into the three-dimensional spatial coordinate system; Performing time synchronization on the projected physiological signal data and the operation trajectory data based on a time interpolation algorithm so that the time deviation between the synchronized data does not exceed a preset deviation; The physiological signal data and the operation trajectory data that have been aligned in time and space are correlated and mapped to construct a correlation matrix reflecting the relationship between the trainee's physiological state and the operation equipment.

4. The training evaluation method based on multimodal data fusion according to claim 1, wherein: The extracting and fusing the physiological signal data and the operation trajectory data based on the correlation matrix to generate a fused feature vector includes: Based on the spatiotemporally aligned physiological signal data and operation trajectory data, the following operations are performed: Extracting physiological sign features and operation force features from the physiological signal data; extracting motion trajectory features from the operation trajectory data; Extracting spatial position deviation features in the correlation matrix; Integrating the physiological sign features, operation force features, motion trajectory features, and spatial position deviation features to generate an initial feature vector; Performing dimension compression processing on the initial feature vector to generate the fused feature vector.

5. The training evaluation method based on multimodal data fusion according to claim 4, characterized in that: The fused feature vector at least includes a motion trajectory feature vector, an operation force feature vector, a physiological sign feature vector, and a spatial position deviation feature vector.

6. The training evaluation method based on multimodal data fusion according to claim 5, characterized in that: The fused feature vector is processed using a hierarchical evaluation mechanism including a basic level evaluation and an optimized level evaluation, and a comprehensive capability evaluation result is output, including: Calculating the matching degree and the deviation degree between the training operation trajectory and the standard trajectory based on the motion trajectory feature vector and the spatial position deviation feature vector; Performing weighted fusion on the matching degree and the deviation degree to generate a basic level evaluation result; Calculating a stability index of the grip of the operating instrument based on the operating force characteristic vector; and predicting a risk assessment value based on the physiological sign characteristic vector; Performing weighted fusion on the stability index and the risk assessment value to generate an optimized level assessment result; The comprehensive capability evaluation result is output by combining the basic level evaluation result and the optimized level evaluation result.

7. The training evaluation method based on multimodal data fusion according to claim 1, wherein: The dynamic intervention measures triggered according to the comprehensive capacity assessment results include: When the deviation in the basic level evaluation result exceeds a first preset threshold, triggering early warning intervention; triggering a corrective intervention when the stability index in the optimization level evaluation result is lower than a second preset threshold, or when the risk assessment value exceeds a third preset threshold; When the comprehensive ability assessment result is lower than a fourth preset threshold, targeted intensive intervention is triggered.

8. A training evaluation system based on multimodal data fusion, characterized in that: include: An acquisition unit, used to collect physiological signal data and operation trajectory data generated during the training process; a construction unit, configured to perform spatiotemporal alignment on the physiological signal data and the operation trajectory data of different sampling rates through three-dimensional space coordinate projection and time interpolation algorithms to construct a correlation matrix; a fusion unit, configured to extract and fuse features of the physiological signal data and the operation trajectory data based on the correlation matrix to generate a fusion feature vector; an evaluation unit, configured to process the fused feature vector using a hierarchical evaluation mechanism including a basic level evaluation and an optimized level evaluation, and output a comprehensive capability evaluation result; A triggering unit is used to trigger corresponding dynamic intervention measures according to the comprehensive capability assessment result.

9. A computer device, characterized in that: The computer device includes a processor and a memory coupled to the processor, wherein a computing program is stored in the memory. When the computer program is executed by the processor, the processor performs the steps of the training and evaluation method based on multimodal data fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the training and evaluation method based on multimodal data fusion according to any one of claims 1 to 7.