Multi-modal inquiry assessment method based on dynamic weight network
Through the multimodal consultation evaluation method of dynamic weight network, combined with multimodal data collection, dynamic weight generation, cross-modal comparative learning and knowledge graph verification, the problem that the existing technology cannot comprehensively evaluate the consultation process is solved, and multi-stage dynamic scoring and accurate process evaluation of the consultation process are achieved.
Patent Information
- Application Number
- CN202510580412.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-09-26
AI Technical Summary
The existing medical consultation skills assessment system mainly focuses on procedural dialogue assessment of a preset case library. It cannot comprehensively evaluate the entire medical consultation process of medical students or interns and lacks dynamic evaluation of the process.
A multimodal medical consultation evaluation method based on dynamic weight network is adopted to achieve a comprehensive evaluation of the medical consultation process through multimodal data collection, dynamic weight generation, cross-modal comparative learning and knowledge graph verification.
It realizes multi-stage dynamic scoring of the consultation process, makes the scoring results more reasonable, and can assess each stage of the consultation skills in real time, thus improving the accuracy of teachers' evaluation of students' consultation process and the effectiveness of their guidance.
Smart Images

Figure CN120706918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and medical education technology, and in particular to a multimodal medical interview evaluation method based on a dynamic weight network. Background Art
[0002] In current medical teaching, it is an important means to improve the medical students' or interns' mastery of medical questioning skills to reasonably evaluate their mastery of medical students' or interns' skills and make up for their deficiencies. The current medical questioning skills evaluation system, such as the patent CN2020106399716, is mainly a pre-examination preset case library for process-based dialogue assessment; the results of the assessment are also mainly based on the results of the answers; however, in actual consultations, the doctor's professionalism is reflected in the entire process, and the focus should be on assessing the entire process of students or interns rather than the results; with the development of large models in recent years, the assessment of processes has been supported, so it is necessary to develop a method that can dynamically evaluate the students' medical consultation process and comprehensively evaluate the students' medical consultation process. Summary of the Invention
[0003] The present invention aims to provide a multimodal consultation evaluation method based on a dynamic weight network, which completes a comprehensive evaluation of the student consultation process through a dynamic weight generation module, a cross-modal comparative learning module and a knowledge graph verification module.
[0004] The present invention is implemented through the following technical solutions: a multimodal diagnostic evaluation method based on a dynamic weight network, comprising:
[0005] S1. Multimodal Data Acquisition Module: This module is used to collect voice data, video data, and electronic medical record data from the consultation process in real time; it also processes the data, converting voice data into speech-to-text data, video data into human posture feature data, and electronic medical record data into diagnostic text data.
[0006] S2. Dynamic Weight Generation Module: The teacher pre-sets the interview skill items and their weights. Based on the voice data, video data, and electronic medical record data from the multimodal data acquisition module, the dynamic weight generation module determines the interview stage. The module then automatically adjusts the interview skill items and their scores based on the identified interview stage.
[0007] S3. Cross-modal comparative learning module: This module performs deep semantic alignment on the speech and text data, human posture feature data, and diagnostic text data from the multimodal data acquisition module with the questioning skill items and their scores from the dynamic weight generation module to generate questioning data to be scored.
[0008] S4. Knowledge Graph Verification Module: Pre-set knowledge graphs for various cases; teachers set cases for students to be assessed, and generate diagnostic reference answers for the knowledge graphs based on the content of the knowledge graphs for the cases to be assessed; calculate the correlation between the diagnostic reference answers in the knowledge graphs and the interview data to be scored in the cross-modal comparative learning module; and assign scores to each interview skill item in the cross-modal comparative learning module;
[0009] S5. Visual feedback module: Arrange the scoring results of the knowledge graph verification module according to the timeline to generate a continuous visual feedback report.
[0010] Optionally, the multimodal data acquisition module steps are as follows:
[0011] S11. Real-time voice data collected during the consultation process is collected through a microphone, and the voice data is de-noised and converted into speech-to-text data;
[0012] S12. Real-time video data of the medical consultation process is collected through the camera, and the human posture of the video data is annotated through the large model to generate human posture feature data;
[0013] S13. Convert the electronic medical record data entered by students during the consultation process into diagnostic text data.
[0014] Optionally, the steps of the dynamic weight generation module are as follows:
[0015] S21. The teacher pre-determines the scoring of interviewing skills and divides the interviewing phase into the inquiry phase, the examination phase, and the diagnosis phase; and sets the weights of the scores of each interviewing skill in the three phases;
[0016] S22 identifies the multimodal data acquisition module voice and text data, human posture feature data, diagnostic text data;
[0017] S23. If the voice data is recognized as voice text data, the human posture feature data in the video data is not checked, and the diagnosis result is not entered in the diagnostic text data of the electronic medical record data; it is determined to be in the inquiry stage; the inquiry stage weight is generated according to the questioning skill item and the score value of each questioning skill item;
[0018] S24. If the human body posture feature data in the video data is recognized for inspection action, it is determined to be the inspection stage; the interview skill items and the score values of each interview skill item are generated according to the inspection stage weight;
[0019] S25. If the diagnosis result input in the diagnostic text data of the electronic medical record data is completed, it is determined to be in the diagnosis stage; generate the questioning skill items and the score value of each questioning skill item according to the diagnosis stage;
[0020] Optionally, the cross-modal contrastive learning module has the following steps:
[0021] S31. The multimodal data acquisition module's voice and text data, human posture feature data, diagnostic text data, and the dynamic weight generation module's rebirth questioning skills items are deeply semantically aligned, so that the voice and text data, human posture feature data, and diagnostic text data can be classified into the corresponding questioning skills items;
[0022] S32. Combine the medical questioning skill items containing voice and text data, human body posture feature data, and diagnostic text data with the regenerated scoring values of each medical questioning skill item to generate medical questioning data to be scored.
[0023] Optionally, the knowledge graph verification module steps are as follows:
[0024] S41. Pre-set knowledge graphs for various cases;
[0025] S42. The teacher presets the case that needs to be assessed; according to the case assessment needs, the voice and text data, human posture feature data, and diagnostic text data content of the knowledge map library of the pre-set case are retrieved;
[0026] S43 retrieves the voice and text data, human posture feature data, and diagnostic text data from the knowledge graph library and regenerates the dynamic weight generation module to semantically align the various diagnostic skills; generate diagnostic reference answer data;
[0027] S44. Perform correlation diagnosis on the question data to be scored in S32 and the diagnosis reference answer data, and assign points to each questioning skill item of the question data to be scored.
[0028] Optionally, the method of constructing the knowledge graph for various pre-set cases is as follows:
[0029] First, the hospital's diagnostic videos are used to build a training set through the large model inference module, and basic diagnostic reference answers are generated based on the training set; then, videos of senior doctors are selected to build a verification set to verify the diagnostic reference answers; finally, videos of doctors from other hospitals are used to build a verification test set to verify the diagnostic reference answers.
[0030] Optionally, the relevance diagnosis of the knowledge graph verification module is performed as follows: the relevance diagnosis is performed using a BiLSTM-CRF model, and the relevance of the inquiry content is evaluated by a graph traversal algorithm and scored. The formula is: Among them, E is the inquiry entity set and D is the diagnosis entity set.
[0031] Optionally, the text data, human posture features, and text data scores of the visual feedback module are displayed in the form of a 3D radar chart.
[0032] Optionally, the dynamic weight generation module: sets the score weight of each stage and combines it with the stage embedding vector to perform dynamic weight adjustment. The formula of the stage embedding vector is w i =σ(W s ·s+b s ), where s is the stage embedding vector, and Ws and bs are learnable parameters.
[0033] Optionally, the visual feedback module further includes dynamic learning path planning, which analyzes weak skills with low scores in the visual feedback report and recommends learning resources of different strengths according to the degree of skill weakness.
[0034] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0035] Through the multimodal data acquisition module and the dynamic weight generation module, the cross-modal comparative learning module is combined with the knowledge graph verification module to complete the assessment of students' medical interview skills from multiple stages of interview, examination, and diagnosis. The dynamic weight generation module can adjust the scoring weight of each stage, and the scoring results will be more reasonable; the knowledge graph verification module relies on the large model to summarize the preset case library, and the reference answer conclusions are more accurate; in addition, the medical interview skills at different stages can be assessed in real time, not just for the answers, which is more conducive to teachers' evaluation and guidance of the entire medical interview process for students, and more conducive to the comprehensive growth of students. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 This is a framework flow chart of a multimodal diagnostic evaluation method based on a dynamic weight network according to the present invention;
[0038] Figure 2 This is a flow chart of a multimodal data acquisition module of a multimodal diagnostic evaluation method based on a dynamic weight network according to the present invention;
[0039] Figure 3 This is a flow chart of a dynamic weight generation module of a multimodal diagnostic evaluation method based on a dynamic weight network according to the present invention;
[0040] Figure 4 This is a flowchart of the knowledge graph verification module of a multimodal interview evaluation method based on a dynamic weight network of the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0042] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0043] The present invention is further described below with reference to specific embodiments.
[0044] Figure 1-4 The present invention is a multimodal interview evaluation method based on a dynamic weight network, and the steps include:
[0045] S1. Multimodal Data Acquisition Module: This module is used to collect voice data, video data, and electronic medical record data from the consultation process in real time; it also processes the data, converting voice data into speech-to-text data, video data into human posture feature data, and electronic medical record data into diagnostic text data.
[0046] S2. Dynamic Weight Generation Module: The teacher pre-sets the interview skill items and their weights. Based on the voice data, video data, and electronic medical record data from the multimodal data acquisition module, the dynamic weight generation module determines the interview stage. The module then automatically adjusts the interview skill items and their scores based on the identified interview stage.
[0047] S3. Cross-modal comparative learning module: This module performs deep semantic alignment on the speech and text data, human posture feature data, and diagnostic text data from the multimodal data acquisition module with the questioning skill items and their scores from the dynamic weight generation module to generate questioning data to be scored.
[0048] S4. Knowledge Graph Verification Module: Pre-set knowledge graphs for various cases; teachers set cases for students to be assessed, and generate diagnostic reference answers for the knowledge graphs based on the content of the knowledge graphs for the cases to be assessed; calculate the correlation between the diagnostic reference answers in the knowledge graphs and the interview data to be scored in the cross-modal comparative learning module; and assign scores to each interview skill item in the cross-modal comparative learning module;
[0049] S5. Visual feedback module: Arrange the scoring results of the knowledge graph verification module according to the timeline to generate a continuous visual feedback report.
[0050] It should be noted that the scoring content of voice data, video data, and electronic medical record data in the dynamic weight generation module is different at different consultation stages; for example, in the consultation stage, the main assessment is the consultation skills, which can be scored according to the 10 language skills of West China Hospital. At this time, the weight of the language skills score will increase by 25%, and the weight of the non-verbal skills score will decrease; in the examination stage, the main assessment is the consultation skills, which can be scored according to the 5 language skills and 10 non-verbal skills of West China Hospital. At this time, the weight of the language skills score will increase by 25%; the weight of the non-verbal skills score will decrease; in addition, the video data uses OpenPose to extract human posture features, and infers the content of the action through human posture features, such as recognizing the human posture features as chest compressions.
[0051] Optionally, the multimodal data acquisition module steps are as follows:
[0052] S11. Real-time voice data collected during the consultation process is collected through a microphone, and the voice data is de-noised and converted into speech-to-text data;
[0053] S12. Real-time video data of the medical consultation process is collected through the camera, and the human posture of the video data is annotated through the large model to generate human posture feature data;
[0054] S13. Convert the electronic medical record data entered by students during the consultation process into diagnostic text data.
[0055] Optionally, the steps of the dynamic weight generation module are as follows:
[0056] S21. The teacher pre-determines the scoring of interviewing skills and divides the interviewing phase into the inquiry phase, the examination phase, and the diagnosis phase; and sets the weights of the scores of each interviewing skill in the three phases;
[0057] S22 identifies the multimodal data acquisition module voice and text data, human posture feature data, diagnostic text data;
[0058] S23. If the voice data is recognized as voice text data, the human posture feature data in the video data is not checked, and the diagnosis result is not entered in the diagnostic text data of the electronic medical record data; it is determined to be in the inquiry stage; the inquiry stage weight is generated according to the questioning skill item and the score value of each questioning skill item;
[0059] S24. If the human body posture feature data in the video data is recognized for inspection action, it is determined to be the inspection stage; the interview skill items and the score values of each interview skill item are generated according to the inspection stage weight;
[0060] S25. If it is recognized that the diagnosis result input in the diagnosis text data of the electronic medical record data is completed, it is determined that it belongs to the diagnosis stage; and the questioning skill items and the scoring values of each questioning skill item are generated according to the diagnosis stage.
[0061] It should be noted that the inquiry stage, examination stage, and diagnosis stage can be divided more finely according to different cases. For example, in the chest pain examination process, the examination stage can be divided into chest compression stage, blood pressure measurement stage, and inquiry interaction stage according to the steps; and different scoring standards can be generated for each stage.
[0062] Optionally, the cross-modal contrastive learning module has the following steps:
[0063] S31. The multimodal data acquisition module's voice and text data, human posture feature data, diagnostic text data, and the dynamic weight generation module's rebirth questioning skills items are deeply semantically aligned, so that the voice and text data, human posture feature data, and diagnostic text data can be classified into the corresponding questioning skills items;
[0064] S32. Combine the medical questioning skill items containing voice and text data, human body posture feature data, and diagnostic text data with the regenerated scoring values of each medical questioning skill item to generate medical questioning data to be scored.
[0065] Optionally, the knowledge graph verification module steps are as follows:
[0066] S41. Pre-set knowledge graphs for various cases;
[0067] S42. The teacher presets the case that needs to be assessed; according to the case assessment needs, the voice and text data, human posture feature data, and diagnostic text data content of the knowledge map library of the pre-set case are retrieved;
[0068] S43 retrieves the voice and text data, human posture feature data, and diagnostic text data from the knowledge graph library and regenerates the dynamic weight generation module to semantically align the various diagnostic skills; generate diagnostic reference answer data;
[0069] S44. Perform correlation diagnosis on the question data to be scored in S32 and the diagnosis reference answer data, and assign points to each questioning skill item of the question data to be scored.
[0070] Optionally, the knowledge graph library for various pre-set cases is constructed as follows:
[0071] First, the hospital's diagnostic videos are used to build a training set through the large model inference module, and basic diagnostic reference answers are generated based on the training set; then, videos of senior doctors are selected to build a verification set to verify the diagnostic reference answers; finally, videos of doctors from other hospitals are used to build a verification test set to verify the diagnostic reference answers.
[0072] Optionally, the relevance diagnosis of the knowledge graph verification module is performed as follows: the relevance diagnosis is performed using a BiLSTM-CRF model, and the relevance of the inquiry content is evaluated by a graph traversal algorithm and scored. The formula is: Among them, E is the inquiry entity set and D is the diagnosis entity set.
[0073] Optionally, the text data, human posture features, and text data scores of the visual feedback module are displayed in the form of a 3D radar chart.
[0074] Optionally, the dynamic weight generation module: sets the score weight of each stage and combines it with the stage embedding vector to perform dynamic weight adjustment. The formula of the stage embedding vector is w i =σ(W s ·s+b s ), where s is the stage embedding vector, and Ws and bs are learnable parameters.
[0075] Optionally, the visual feedback module also includes dynamic learning path planning, which analyzes the weak skills with low scores in the visual feedback report and recommends learning resources of different intensities according to the degree of skill weakness; for example, if the "chest compression skill" is weak, a 15-minute video teaching and 3 simulation exercises are recommended.
[0076] Specific implementation scenarios
[0077] For example, there are the following scenarios: training on interviewing patients with acute chest pain;
[0078] The teacher sets the assessment scenario as coronary heart disease and sets the weights for each stage;
[0079] The multimodal data acquisition module collected voice messages such as "chest pain," "how long has the chest pain lasted? Does the shoulder hurt? Does the patient have high blood pressure?", and "chest pain for two hours, shoulder pain, high blood pressure." It also collected video data of "the student leaning forward and nodding." It also collected records from the electronic medical record that "the patient has had chest pain for two hours, shoulder pain, and a history of high blood pressure."
[0080] If the dynamic weight generation module determines that the patient is in the consultation stage, the language skill score weight will be increased by 30%, with a total score of 60, divided into 5 language skill sub-items; the non-verbal skill weight will be reduced by 20%, with a total score of 10, divided into 2 non-verbal skill sub-items; the record content score weight will be reduced by 10%, with a total score of 20, divided into 3 record content sub-items; the total score will be 90;
[0081] The cross-modal comparative learning module divides the speech and text data collected by the multimodal data acquisition module into five language skill categories, divides the human posture features of the video data into two non-language skill categories, and divides the diagnostic text data into three recording skill categories, performing deep semantic alignment.
[0082] The knowledge graph verification module uses a pre-set knowledge graph of various cases, including video clips, audio clips, and electronic medical records of 100 cases of coronary heart disease consultations at West China Hospital. It generates diagnostic reference answers for 5 language skills sub-items, 2 non-verbal skills sub-items, and 3 recording skills sub-items during the consultation phase for coronary heart disease. The knowledge graph verification module calculates the correlation between the consultation data from the cross-modal comparative learning module and the diagnostic reference answers from the knowledge graph verification module, and the calculated correlation is 0.8 for the language skills sub-item, 0.9 for the non-verbal skills sub-item, and 0.7 for the recording skills sub-item. Finally, the students' language skills scores for the consultation phase were 48, 9 for the non-verbal skills sub-item, and 14 for the recording skills sub-item.
[0083] The human posture features collected from the video data are converted into chest compressions. The dynamic weight generation module determines to jump to the inspection stage and generates new scoring criteria. The cross-modal comparative learning module and the knowledge graph verification module regenerate new scores.
[0084] The final score of the knowledge graph verification module is visualized in the feedback report of the visual feedback module for teacher evaluation.
[0085] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0086] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A multimodal interview assessment method based on a dynamic weight network, characterized in that: include: S1. Multimodal Data Acquisition Module: This module is used to collect voice data, video data, and electronic medical record data from the consultation process in real time; it also processes the data, converting voice data into speech-to-text data, video data into human posture feature data, and electronic medical record data into diagnostic text data. S2. Dynamic Weight Generation Module: The teacher pre-sets the interview skill items and their weights. Based on the voice data, video data, and electronic medical record data from the multimodal data acquisition module, the dynamic weight generation module determines the interview stage. The module then automatically adjusts the interview skill items and their scores based on the identified interview stage. S3. Cross-modal comparative learning module: This module performs deep semantic alignment on the speech and text data, human posture feature data, and diagnostic text data from the multimodal data acquisition module with the questioning skill items and their scores from the dynamic weight generation module to generate questioning data to be scored. S4. Knowledge Graph Verification Module: Pre-set knowledge graphs for various cases; teachers set cases for students to be assessed, and generate diagnostic reference answers for the knowledge graphs based on the content of the knowledge graphs for the cases to be assessed; calculate the correlation between the diagnostic reference answers in the knowledge graphs and the interview data to be scored in the cross-modal comparative learning module; and assign scores to each interview skill item in the cross-modal comparative learning module; S5. Visual feedback module: Arrange the scoring results of the knowledge graph verification module according to the timeline to generate a continuous visual feedback report.
2. The multimodal interview and evaluation method based on a dynamic weight network according to claim 1, characterized in that: The multimodal data acquisition module steps are as follows: S11. Real-time voice data collected during the consultation process is collected through a microphone, and the voice data is de-noised and converted into speech-to-text data; S12. Real-time video data of the medical consultation process is collected through the camera, and the human posture of the video data is annotated through the large model to generate human posture feature data; S13. Convert the electronic medical record data entered by students during the consultation process into diagnostic text data.
3. The multimodal interview and evaluation method based on a dynamic weight network according to claim 1, characterized in that: The steps of the dynamic weight generation module are as follows: S21. The teacher pre-determines the scoring of interviewing skills and divides the interviewing phase into the inquiry phase, the examination phase, and the diagnosis phase; and sets the weights of the scores of each interviewing skill in the three phases; S22 identifies the multimodal data acquisition module voice and text data, human posture feature data, diagnostic text data; S23. If the voice data is recognized as voice text data, the human posture feature data in the video data is not checked, and the diagnosis result is not entered in the diagnostic text data of the electronic medical record data; it is determined to be in the inquiry stage; the inquiry stage weight is generated according to the questioning skill item and the score value of each questioning skill item; S24. If the human body posture feature data in the video data is recognized for inspection action, it is determined to be the inspection stage; the interview skill items and the score values of each interview skill item are generated according to the inspection stage weight; S25. If it is recognized that the diagnosis result input in the diagnosis text data of the electronic medical record data is completed, it is determined that it belongs to the diagnosis stage; and the questioning skill items and the scoring values of each questioning skill item are generated according to the diagnosis stage.
4. The multimodal interview and evaluation method based on a dynamic weight network according to claim 3, characterized in that: The steps of the cross-modal contrastive learning module are as follows: S31. The multimodal data acquisition module's voice and text data, human posture feature data, diagnostic text data, and the dynamic weight generation module's rebirth questioning skills items are deeply semantically aligned, so that the voice and text data, human posture feature data, and diagnostic text data can be classified into the corresponding questioning skills items; S32. Combine the medical questioning skill items containing voice and text data, human body posture feature data, and diagnostic text data with the regenerated scoring values of each medical questioning skill item to generate medical questioning data to be scored.
5. The multimodal interview and evaluation method based on a dynamic weight network according to claim 4, characterized in that: The steps of the knowledge graph verification module are as follows: S41. Pre-set knowledge graphs for various cases; S42. The teacher presets the case that needs to be assessed; according to the case assessment needs, the voice and text data, human posture feature data, and diagnostic text data content of the knowledge map library of the pre-set case are retrieved; S43 retrieves the voice and text data, human posture feature data, and diagnostic text data from the knowledge graph library and regenerates the dynamic weight generation module to semantically align the various diagnostic skills; generate diagnostic reference answer data; S44. Perform correlation diagnosis on the question data to be scored in S32 and the diagnosis reference answer data, and assign points to each questioning skill item of the question data to be scored.
6. A multimodal interview and evaluation method based on a dynamic weight network according to claim 5, characterized in that: The method of constructing the knowledge graph for presetting various cases is as follows: First, the hospital's diagnostic videos are used to build a training set through the large model inference module, and basic diagnostic reference answers are generated based on the training set; then, videos of senior doctors are selected to build a verification set to verify the diagnostic reference answers; finally, videos of doctors from other hospitals are used to build a verification test set to verify the diagnostic reference answers.
7. The multimodal interview and evaluation method based on a dynamic weight network according to claim 5, characterized in that: The relevance diagnosis of the knowledge graph verification module adopts the following method: the relevance diagnosis adopts BiLSTM-CRF model for diagnosis, and the relevance of the inquiry content is evaluated by graph traversal algorithm and scored. The formula is: Among them, E is the inquiry entity set and D is the diagnosis entity set.
8. The multimodal interview and evaluation method based on a dynamic weight network according to claim 1, characterized in that: The text data, human posture features, and text data scores of the visual feedback module are displayed in the form of a 3D radar chart.
9. The multimodal interview and evaluation method based on a dynamic weight network according to claim 3, characterized in that: The dynamic weight generation module: sets the score weight of each stage and combines the stage embedding vector to perform dynamic weight adjustment. The formula of the stage embedding vector is w i =σ(W s ·s+b s ), where s is the stage embedding vector, and Ws and bs are learnable parameters.
10. The multimodal interview and evaluation method based on a dynamic weight network according to claim 1, characterized in that: The visual feedback module also includes dynamic learning path planning, which analyzes weak skills with low scores in the visual feedback report and recommends learning resources of different strengths according to the degree of skill weakness.
Citation Information
Cited By
Multi-scene self-adaption-based pharmacist clinical ability assessment method and system
CN122089170A