A multi-dimensional medical education evaluation method and system based on speech recognition
Through speech recognition technology, multi-dimensional evaluation of student answers has been solved, and the cumbersome process and subjective deviation of traditional medical education evaluation has been solved, accurate evaluation of students' language expression and clinical thinking has been achieved, and the intelligent transformation of medical education has been promoted.
Patent Information
- Application Number
- CN202510563981.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The traditional medical education evaluation method has cumbersome processes, long time, single dimensions, and subjective deviations in the scores, which are difficult to truly reflect the learners' language expression ability, communication skills and clinical thinking in the diagnosis and treatment scenarios.
A multi-dimensional medical education evaluation method based on speech recognition is adopted, and the evaluation score is determined through speech signal slicing, text conversion, character grouping, similarity determination and multiple rounds of iterative keyword extraction, combined with preset keyword list and part-of-speech sorting.
It improves the reliability and accuracy of evaluation, can effectively deal with students' intermittent pronunciation and spoken flaws, significantly improves the fault tolerance of inflexible answers, and promotes the digital and intelligent development of medical education.
Smart Images

Figure CN120087804B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of educational evaluation, and particularly to a multi-dimensional medical education evaluation method and system based on speech recognition. Background Art
[0002] Speech recognition is a technology that converts audio signals into readable text. It captures sounds through a microphone or other audio input devices, and then analyzes and matches the characteristics of the sounds using acoustic models and language models, and finally outputs the corresponding text content.
[0003] Medical education refers to the cultivation and assessment of the theoretical knowledge, clinical skills, and professional qualities of medical students and medical staff through systematic curriculum teaching, clinical internships, and continuing education in the field of medical and health. The existing evaluation methods mainly rely on written tests and interviews, which are not only cumbersome and time-consuming, but also have a single dimension, and there are subjective biases in scoring, making it difficult to truly reflect the language expression ability, communication skills, and clinical thinking of learners in the diagnosis and treatment scenario.
[0004] : With the maturity of deep learning and speech recognition technologies, intelligent analysis and real-time feedback based on sounds have become possible. It can not only automatically identify pronunciation accuracy, term usage, and speech rate and intonation, but also evaluate response logic and humanistic care literacy in combination with natural language processing, injecting new vitality and opportunities into medical education evaluation. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-dimensional medical education evaluation method and system based on speech recognition, and solve the following technical problems:
[0006] The traditional evaluation method is not only cumbersome and time-consuming, but also has a single dimension, and there are subjective biases in scoring, making it difficult to truly reflect the language expression ability, communication skills, and clinical thinking of learners in the diagnosis and treatment scenario.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] A multi-dimensional medical education evaluation method based on speech recognition includes the following steps:
[0009] Obtain the speech signal of a student during the process of answering a single medical education evaluation question, slice the speech signal to obtain speech segments, and convert the speech segments into text form, denoted as text segments;
[0010] Sort the characters in a single text segment, group the characters based on the sorting, determine the pending words according to the similarity degree between adjacent two groups, obtain the target words based on the pending words, and determine the first keywords according to the positions of the target words in the sorting;
[0011] Extract keywords from the text segment based on a preset keyword list, denoted as the second keywords, sort the second keywords to obtain a keyword sorting; obtain the part-of-speech of the second keywords, where the part-of-speech includes symptoms, examinations, diagnoses, and treatments.
[0012] Sort the part-of-speech in the order of the keyword sorting to obtain a part-of-speech sorting, and determine an evaluation score based on the part-of-speech sorting and the first keywords.
[0013] As a further solution of the present invention: the speech segment satisfies the following constraints:
[0014] Number the speech segment, C KS (a+1) -C JS a ≥Cys, C KS (a+1) represents the time stamp corresponding to the start point of the speech segment numbered a + 1, C JS a ≥ represents the time stamp corresponding to the end point of the speech segment numbered a, Cys represents a preset time stamp difference, a ∈ [1, n - 1], and n represents the total number of the speech segments.
[0015] As a further solution of the present invention: the process of determining the first keywords includes:
[0016] Step 1: Sort the characters in a single text segment in the order of their appearance in the text segment to obtain a character sorting. Starting from the first character in the character sorting, divide adjacent m characters into the same group, where m is a preset number of characters.
[0017] Step 2: When the similarity degree of characters in any two groups is greater than a preset similarity degree threshold, combine the characters in these two groups to obtain a pending word, and obtain the similarity degree between the pending word and the keywords in the keyword list.
[0018] Mark the keyword corresponding to the maximum similarity degree as the target word.
[0019] Step 3: Obtain a new number of characters m' = m - 1, repeat Step 1 and Step 2, obtain a new target word, repeat the above steps until the new number of characters is greater than a preset number of character threshold.
[0020] Obtain two groups corresponding to a single target word, denoted as group i and group i - 1. The sorting position W of the first character in group i - 1 in the character sorting i-1 is on the left side of the sorting position W i of, and use the sorting position W i-1 as the target position.
[0021] Step 4: Obtain all the target positions, count the number of occurrences of the target words corresponding to the target positions within the sorting position interval [WMB - W, WMB + W], and obtain the target word corresponding to the maximum number of occurrences as the first keyword, where WMB represents the sorting position of the target position in the character sorting, and W represents the preset number of sorting positions.
[0022] As a further solution of the present invention: The process of determining the evaluation score based on the part-of-speech sorting and the first keyword includes:
[0023] Obtain the part of speech F1 at the first position in the part-of-speech sorting, and determine whether the part of speech F1 is the symptom;
[0024] If so, let the evaluation score P1 = 1, and start from the part of speech at the second position in the part-of-speech sorting to determine the first part of speech F2 that is not the symptom, and determine whether the part of speech F2 is an examination;
[0025] If so, obtain a new evaluation score P2 = P1 + 1, and start from the part of speech F2 to determine the first part of speech F3 that is not the examination, and determine whether the part of speech F3 is a diagnosis;
[0026] If so, obtain a new evaluation score P3 = P2 + 1, and start from the part of speech F3 to determine the first part of speech F4 that is not the diagnosis, and determine whether the part of speech F4 is a treatment;
[0027] If so, obtain a new evaluation score P4 = P3 + 1.
[0028] As a further solution of the present invention: The process of determining the evaluation score based on the part-of-speech sorting and the first keyword further includes a judgment step:
[0029] If the part of speech F1 is not the symptom, obtain the timestamp CF1 corresponding to the part of speech F1, obtain the first keyword whose corresponding timestamp is within the timestamp interval [CF1 - C, CF1 + C], denoted as the to-be-determined word, and obtain the part of speech of the to-be-determined word, where C is a preset value;
[0030] When there is a part of speech of the to-be-determined word that is a symptom, let the evaluation score P1 = 0.5, and execute the subsequent steps; when there is no part of speech of the to-be-determined word that is an examination, let the evaluation score P1 = 0, and execute the subsequent steps.
[0031] As a further solution of the present invention: The process of determining the evaluation score based on the part-of-speech sorting and the first keyword further includes:
[0032] When the part of speech F2 is not an examination, execute the judgment step;
[0033] When the part-of-speech F2 is not "diagnosis", execute the judgment step;
[0034] When the part-of-speech F2 is not "treatment", execute the judgment step.
[0035] As a further solution of the present invention: the part-of-speech of the second keyword and the word to be determined is based on manual annotation.
[0036] A multi-dimensional medical education evaluation system based on speech recognition, comprising:
[0037] Acquisition module: acquire the speech signal of a student during answering a single medical education evaluation question, slice the speech signal to obtain speech segments, and convert the speech segments into text form, denoted as text segments;
[0038] First keyword module: sort the characters in a single text segment, group the characters based on the sorting, determine the word to be determined according to the similarity between adjacent two groups, obtain the target word based on the word to be determined, and determine the first keyword according to the position of the target word in the sorting;
[0039] Second keyword module: extract the keywords in the text segment based on a preset keyword list, denoted as the second keyword, sort the second keyword to obtain a keyword sorting; obtain the part-of-speech of the second keyword, and the part-of-speech includes symptoms, examinations, diagnoses, and treatments;
[0040] Evaluation module: sort the part-of-speech in the order of the keyword sorting to obtain a part-of-speech sorting, and determine the evaluation score based on the part-of-speech sorting and the first keyword.
[0041] Advantages of the present invention: Compared with the prior art:
[0042] 1) Through the speech slicing method based on the preset timestamp difference, the present invention can effectively cope with the intermittent speech phenomenon caused by pauses or stutters during the student's answering process. This method maintains a sufficient time interval between each speech segment, ensures that the sliced text segments have high coherence and integrity, and significantly reduces the errors caused by overlapping or missing of speech signals, ensuring that the subsequent speech-to-text can more accurately complete the recognition task and reducing the error rate;
[0043] 2) The multi-round iterative keyword extraction method based on sliding window grouping and similarity determination proposed by the present invention can effectively compensate for the oral defects that occur when students have incoherent pronunciation, repeated characters, or stuttering. By presetting the character number sliding window and calculating the similarity between adjacent groups, the discontinuous characters are combined into words to be determined, and multi-round similarity matching is performed in a predefined term library, which can restore the most appropriate professional terms from the noisy text segments;
[0044] 3) In view of the repeated pronunciations and stutters caused by students' lack of proficiency in oral answers, the present invention innovatively introduces a first keyword extraction mechanism based on character sliding window grouping and similarity matching. When there are repeated or interrupted middle characters in a certain term, it is difficult for traditional direct matching methods based on a preset keyword list to accurately identify. Through multiple rounds of sliding window grouping, adjacent group similarity determination, and iterative refinement, the present invention can restore the most appropriate medical term from the noisy text and determine the first keyword by counting the positions of high-frequency target words. During the thinking chain evaluation process, when the second keyword extraction fails or is missing, this first keyword is used to assist in judging whether the student unfolds according to the logic of "symptom - examination - diagnosis - treatment", significantly improving the system's fault tolerance for unsmooth answers and evaluation accuracy, and contributing to promoting the digital transformation and intelligent development of medical education. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The present invention will be further described below with reference to the accompanying drawings.
[0046] Figure 1 is a schematic flowchart of a multi-dimensional medical education evaluation method based on speech recognition according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Please refer to Figure 1 As shown, the present invention is a multi-dimensional medical education evaluation method based on speech recognition, including the following steps:
[0049] Obtain the speech signal of a student during the process of answering a single medical education evaluation question, slice the speech signal to obtain speech segments, and convert the speech segments into text form, denoted as text segments;
[0050] In a preferred embodiment of the present invention, the speech segments satisfy the following constraints:
[0051] Number the speech segments, C KS (a+1) -C JS a ≥Cys, C KS (a+1) represents the time stamp corresponding to the start point of the speech segment numbered a + 1, C JS a≥ represents the timestamp corresponding to the end point of the speech segment numbered a, Cys represents a preset timestamp difference, a ∈ [1, n - 1], and n represents the total number of the speech segments;
[0052] Exemplarily, when a student answers the test question "What is heart failure", the system starts recording from 0.00 seconds when he starts speaking and continues recording until 18.70 seconds. At this time, he pauses for 2.30 seconds due to thinking (i.e., 21.00 seconds, the start point of the second segment minus 18.70 seconds, the end point of the first segment), satisfying 2.30s ≥ Cys. Then, 0.00 - 18.70 seconds is defined as the first segment; the second segment starts recording from 21.00 seconds to 35.20 seconds, and pauses again for 2.50 seconds (37.70 seconds, the start point of the third segment minus 35.20 seconds, the end point of the second segment), satisfying 2.50s ≥ Cys. Then, 21.00 - 35.20 seconds is defined as the second segment; the third segment is recorded from 37.70 seconds to 52.10 seconds. After the recording is completed, the system sends these three segments of speech into the recognition engine respectively, and obtains the corresponding text segments: the first segment is "Heart failure refers to the systemic circulatory disorder caused by insufficient cardiac pumping function", the second segment is "Patients often present with fatigue, shortness of breath and lower extremity edema", and the third segment is "Routine examinations include electrocardiogram, echocardiogram and blood biochemistry, etc.". In this way, the silent interval between each speech segment is not less than Cys, making the sliced text segments both coherent and complete, providing a clear and reliable written basis for subsequent medical education evaluation;
[0053] Sort the characters in a single said text segment, group the characters based on the sorting, determine the pending words according to the similarity degree between adjacent two groups, obtain the target words based on the pending words, and determine the first keyword according to the position of the target word in the sorting;
[0054] In another preferred embodiment of the present invention, the process of determining the first keyword includes:
[0055] Step 1: Sort the characters in a single said text segment in the order of appearance in the text segment to obtain a character sorting. Starting from the first character in the character sorting, divide adjacent m characters into the same group, where m is a preset number of characters;
[0056] Step 2: When the similarity degree of the characters in any two said groups is greater than a preset similarity degree threshold, combine the characters in these two groups to obtain a pending word, and obtain the similarity degree between the pending word and the keywords in the keyword list;
[0057] Mark the keyword corresponding to the maximum similarity degree as the target word;
[0058] Step 3: Obtain a new number of characters m' = m - 1, repeat Step 1 and Step 2 above to obtain a new target word, and repeat the above steps until the new number of characters is greater than a preset character number threshold;
[0059] Obtain two groups corresponding to a single target word, denoted as group i and group i - 1. The sorting position W of the first character in group i - 1 in the character sorting i-1 is at the sorting position W i on the left side, and use the sorting position W i-1 as the target position;
[0060] Step 4: Obtain all the target positions, count the number of times the target words corresponding to the target positions within the sorting position range [WMB - W, WMB + W] appear, and obtain the target word corresponding to the maximum number of times as the first keyword. WMB represents the sorting position of the target position in the character sorting, and W represents the preset number of sorting positions.
[0061] It should be noted that in actual spoken Q&A, students often have discontinuous, repeated, or jumping pronunciations due to unfamiliarity with complex medical terms. For example, when trying to express "coronary atherosclerotic heart disease", it may be pronounced as "coronary - artery - athero - athero - atherosclerotic - heart - heart - disease", making it difficult for the direct matching method based on the preset keyword list to locate the core concept. To solve this problem, the present invention robustly restores the true intention of students in complex noise by grouping characters in text fragments in multiple rounds and combining any characters;
[0062] Exemplarily, for the text fragment "Coronary athero - atherosclerotic heart heart disease is a chronic ischemic heart disease caused by coronary atherosclerosis", the length is more than 30 characters. First, arrange all characters in the order of appearance as:
[0063] [coronary, artery, athero, athero, atherosclerotic, heart, heart, disease, is, coronary, artery, athero, atherosclerotic, caused, by, chronic, ischemic, heart, disease];
[0064] Assuming that the initial sliding window length m = 8, the 1st to 8th characters "coronary atherosclerosis" are taken as group A, and the 9th to 16th characters "sclerotic heart disease is" are taken as group B. The system randomly selects 1 to 8 characters from these 16 characters in accordance with the mathematical permutation and combination method to construct several undetermined words, and calculates the similarity of each undetermined word with the terms "coronary atherosclerotic heart disease", "coronary atherosclerosis", "chronic ischemic heart disease" and so on in the preset term library one by one. Finally, "coronary atherosclerotic heart disease" is selected as the target word with the highest similarity. At the same time, its starting position 1 in the original sequence is recorded, and then m is reduced to 7, and the groups are re-divided (the 1st -7 "coronary atherosclerosis", 8-14 "heart sclerosis"), any combination of these 14 characters is matched again, and "coronary atherosclerosis" is identified again, and recorded at position 1; when m=6, the group "coronary atherosclerosis" and "heart disease is crown" are matched to obtain "coronary atherosclerosis", and recorded at position 1; when m=5 and 4, the term is located multiple times in different 8-10 character groups, and recorded at positions 3 and 5; until m=3, the group "coronary atherosclerosis" and "atherosclerosis" are also successfully extracted as core terms, and recorded at position 4. Finally, all target positions are summarized as {1,1,1,3,5,4};
[0065] For any target position, a corresponding target position interval is set, and the first keyword corresponding to the target position interval is determined. After the determination, new first keywords are determined for other target positions, and finally all first keywords are obtained. Through this multi-round iterative sliding window and arbitrary character combination strategy, the present invention can effectively filter out repeated and jumping character noise in students' stumbling expressions, accurately extract the most representative long terms, and ensure the robustness and accuracy of medical concept recognition.
[0066] Extracting keywords from the text segment based on a preset keyword list, recording them as second keywords, sorting the second keywords to obtain a keyword ranking; obtaining the part of speech of the second keyword, the part of speech including symptom, examination, diagnosis and treatment;
[0067] Sorting the parts of speech according to the order of the keyword sorting to obtain a part-of-speech sorting, and determining an evaluation score based on the part-of-speech sorting and the first keyword;
[0068] In a preferred embodiment of the present invention, the process of determining the evaluation score based on the part-of-speech ranking and the first keyword includes:
[0069] Obtaining the first part of speech F1 in the part of speech ranking, and determining whether the part of speech F1 is the symptom;
[0070] If so, set the evaluation score P1 = 1, and starting from the second part of speech in the part-of-speech ranking, determine the first part of speech F2 that is not the symptom, and determine whether the part of speech F2 is an examination;
[0071] If so, obtain a new evaluation score P2 = P1 + 1, and starting from the part of speech F2, determine the first part of speech F3 that is not the examination, and determine whether the part of speech F3 is a diagnosis;
[0072] If so, obtain a new evaluation score P3 = P2 + 1, and starting from the part of speech F3, determine the first part of speech F4 that is not the diagnosis, and determine whether the part of speech F4 is a treatment;
[0073] If so, obtain a new evaluation score P4 = P3 + 1;
[0074] Exemplarily, assume that in a certain text segment, the system has extracted the first keyword "acute myocardial infarction" through the foregoing method, and then matches six terms "chest pain", "shortness of breath", "electrocardiogram", "echocardiogram", "acute myocardial infarction", and "thrombolytic therapy" from the preset keyword list, and forms a keyword ranking according to their order of appearance in the segment. Then, according to the part of speech of each term ("chest pain" and "shortness of breath" are symptoms; "electrocardiogram" and "echocardiogram" are examinations; "acute myocardial infarction" is a diagnosis; "thrombolytic therapy" is a treatment), obtain the part-of-speech ranking [symptom, symptom, examination, examination, diagnosis, treatment]. The system first takes the first part of speech in the part-of-speech ranking F1 = symptom. After determining it as a symptom, set P1 = 1; then skip the second symptom, and starting from the third one, locate the first non-symptom F2 = examination. After determining it as an examination, set P2 = P1 + 1 = 2; then locate the first non-examination F3 = diagnosis after F2. After determining it as a diagnosis, set P3 = P2 + 1 = 3; finally, locate the first non-diagnosis F4 = treatment after F3. After determining it as a treatment, set P4 = P3 + 1 = 4. Thus, the entire scoring process is completed, and finally the evaluation score of this answer segment is obtained as 4 points;
[0075] In a preferred embodiment of the present invention, the process of determining the evaluation score based on the part-of-speech ranking and the first keyword further includes a judgment step:
[0076] If the part of speech F1 is not a symptom, obtain the time stamp CF1 corresponding to the part of speech F1, obtain the first keyword whose corresponding time stamp is within the time stamp interval [CF1 - C, CF1 + C], denoted as the pending word, and obtain the part of speech of the pending word, where C is a preset value;
[0077] When there is a part of speech of the pending word that is a symptom, set the evaluation score P1 = 0.5, and execute the subsequent steps; when there is no part of speech of the pending word that is an examination, set the evaluation score P1 = 0, and execute the subsequent steps;
[0078] It is understandable that the process of determining the evaluation score based on the part-of-speech sorting and the first keyword further includes:
[0079] When the part of speech F2 is not "examination", execute the judgment step;
[0080] When the part of speech F3 is not "diagnosis", execute the judgment step;
[0081] When the part of speech F4 is not "treatment", execute the judgment step;
[0082] It should be noted that in an oral answer, the system extracts a second keyword sequence from the student's answer. The first F1 is identified as "electrocardiogram examination", and its corresponding timestamp CF1 is 12.5 seconds. Since F1 does not belong to "symptom", the system enters the judgment step. When the preset constant C = 3 seconds (which can be preset according to experience), a time interval [CF1 - C, CF1 + C] = [9.5s, 15.5s] is constructed. The first keyword obtained by the previous sliding window combination method is retrieved within this interval. Suppose "chest pain" is matched at 13.2 seconds and its part of speech is "symptom", then it is determined that there is a pending word with the part of speech "symptom", and the evaluation score P1 = 0.5 is set, and the subsequent scoring continues; if no first keyword with the part of speech "symptom" is retrieved within [9.5s, 15.5s], it is determined that there is no qualified pending word, and the evaluation score P1 = 0 is set. Subsequently, if the same situation as F1 appears for F2, F3, and F4 (that is, F2 is not "examination", F3 is not "diagnosis", and F4 is not "treatment"), the subsequent F2, F2, and F4 are scored according to the same logic; by using the core terms and their time information extracted by the previous sliding window combination, partial scores can be given when the student does not explicitly mention the symptom but the corresponding concept already exists in the context, so as to compensate for the direct matching failure caused by the student's oral inarticulateness or omission of symptoms;
[0083] It is understandable that the part of speech of the second keyword and the pending word is based on manual annotation;
[0084] A multi-dimensional medical education evaluation system based on speech recognition, including:
[0085] Acquisition module: Obtain the voice signal of the student during the process of answering a single medical education evaluation question, slice the voice signal to obtain voice segments, and convert the voice segments into text form, denoted as text segments;
[0086] First keyword module: Sort the characters in a single text segment, group the characters based on the sorting, determine the pending word according to the similarity degree between adjacent two groups, obtain the target word based on the pending word, and determine the first keyword according to the position of the target word in the sorting;
[0087] Second keyword module: Extract keywords from the text segment based on a preset keyword list, denoted as second keywords, sort the second keywords to obtain a keyword sorting; obtain the part-of-speech of the second keywords, where the part-of-speech includes symptoms, examinations, diagnoses, and treatments.
[0088] Evaluation module: Sort the part-of-speech in the order of the keyword sorting to obtain a part-of-speech sorting, and determine an evaluation score based on the part-of-speech sorting and the first keyword.
[0089] The above has described an embodiment of the present invention in detail, but the content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equal changes and improvements made according to the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. A multi-dimensional medical education evaluation method based on speech recognition, characterized in that, It includes the following steps: Obtain the voice signal of a student during the process of answering a single medical education evaluation question, slice the voice signal to obtain voice segments, convert the voice segments into text form, and record them as text segments; Sort the characters in a single text segment, group the characters based on the sorting, determine the pending words according to the similarity between adjacent two groups, obtain the target words based on the pending words, and determine the first keyword according to the position of the target word in the sorting; Extract the keywords in the text segment based on a preset keyword list, record them as the second keywords, sort the second keywords to obtain a keyword sorting; obtain the part-of-speech of the second keywords, and the part-of-speech includes symptoms, examinations, diagnoses, and treatments; Sort the part-of-speech in the order of the keyword sorting to obtain a part-of-speech sorting, and determine the evaluation score based on the part-of-speech sorting and the first keyword; The process of determining the first keyword includes: Step 1: Sort the characters in a single text segment in the order of their appearance in the text segment to obtain a character sorting. Starting from the first character in the character sorting, divide adjacent m characters into the same group, where m is a preset number of characters; Step 2: When the similarity between the characters in any two groups is greater than a preset similarity threshold, combine the characters in these two groups to obtain a pending word, and obtain the similarity between the pending word and the keywords in the keyword list; Mark the keyword corresponding to the maximum similarity as the target word; Step 3: Obtain a new number of characters m' = m - 1, repeat Step 1 and Step 2, obtain a new target word, repeat the above steps until the new number of characters is greater than a preset number-of-characters threshold; Obtain two groups corresponding to a single said target word, denoted as group i and group i-1, where the sorting position W of the first character in the said group i-1 in the character sorting i-1 is at the sorting position W i on the left side, and use the said sorting position W i-1 as the target position; Step 4: Obtain all the target positions, count the number of times the target words corresponding to the target positions within the sorting position range [WMB - W, WMB + W] appear, and obtain the target word corresponding to the maximum number of times as the first keyword, where WMB represents the sorting position of the target position in the character sorting, and W represents the preset number of sorting positions.
2. The multi-dimensional medical education evaluation method based on speech recognition according to claim 1, characterized in that, The voice segments satisfy the following constraints: Number the speech segments, C KS (a+1) -C JS a ≥ Cys, C KS (a+1) represents the timestamp corresponding to the start point of the speech segment numbered a + 1, C JS a ≥ represents the timestamp corresponding to the end point of the speech segment numbered a, Cys represents a preset timestamp difference, a ∈ [1, n - 1], and n represents the total number of the speech segments.
3. A multi-dimensional medical education evaluation method based on speech recognition according to claim 1, characterized in that, The process of determining the evaluation score based on the part-of-speech sorting and the first keyword includes: Obtain the first part-of-speech F1 in the part-of-speech sorting, and determine whether the part-of-speech F1 is the symptom; If so, let the evaluation score P1 = 1, and start from the second part-of-speech in the part-of-speech sorting to determine the first part-of-speech F2 that is not the symptom, and determine whether the part-of-speech F2 is an examination; If so, obtain a new evaluation score P2 = P1 + 1, and start from the part-of-speech F2 to determine the first part-of-speech F3 that is not the examination, and determine whether the part-of-speech F3 is a diagnosis; If so, obtain a new evaluation score P3 = P2 + 1, and start from the part-of-speech F3 to determine the first part-of-speech F4 that is not the diagnosis, and determine whether the part-of-speech F4 is a treatment; If so, obtain a new evaluation score P4 = P3 + 1.
4. A multi-dimensional medical education evaluation method based on speech recognition according to claim 3, characterized in that, The process of determining the evaluation score based on the part-of-speech sorting and the first keyword also includes a judgment step: If the part of speech F1 is not a symptom, obtain the timestamp CF1 corresponding to the part of speech F1, obtain the first keywords whose corresponding timestamps are within the timestamp interval [CF1 - C, CF1 + C], denoted as the to-be-determined words, and obtain the part of speech of the to-be-determined words, where C is a preset value; When there exists a to-be-determined word whose part of speech is a symptom, set the evaluation score P1 = 0.5, and perform the subsequent steps; when there does not exist a to-be-determined word whose part of speech is an examination, set the evaluation score P1 = 0, and perform the subsequent steps.
5. A multi-dimensional medical education evaluation method based on speech recognition according to claim 4, characterized in that, The process of determining the evaluation score based on the part-of-speech sorting and the first keywords further includes: When the part of speech F2 is not an examination, perform the judgment step; When the part of speech F2 is not a diagnosis, perform the judgment step; When the part of speech F2 is not a treatment, perform the judgment step.
6. A multi-dimensional medical education evaluation method based on speech recognition according to claim 4, characterized in that, The part of speech of the second keyword and the to-be-determined word is based on manual annotation.
7. A multi-dimensional medical education evaluation system based on speech recognition, characterized in that Includes: Collection module: Obtain the voice signal of a student during the process of answering a single medical education evaluation question, slice the voice signal to obtain voice segments, and convert the voice segments into text form, denoted as text segments; First keyword module: Sort the characters in a single text segment, group the characters based on the sorting, determine the to-be-determined words according to the similarity degree between adjacent two groups, obtain the target words based on the to-be-determined words, and determine the first keywords according to the positions of the target words in the sorting; Second keyword module: Extract the keywords in the text segment based on a preset keyword list, denoted as the second keywords, sort the second keywords to obtain a keyword sorting; obtain the part of speech of the second keywords, where the part of speech includes symptoms, examinations, diagnoses, and treatments; Evaluation module: Sort the part of speech in the order of the keyword sorting to obtain a part-of-speech sorting, and determine the evaluation score based on the part-of-speech sorting and the first keywords; The process of determining the first keywords includes: Step 1: Sort the characters in a single text segment in the order of their appearance in the text segment to obtain a character sorting. Starting from the first character in the character sorting, divide adjacent m characters into the same group, where m is a preset number of characters; Step 2: When the similarity degree between the characters in any two groups is greater than a preset similarity degree threshold, combine the characters in these two groups to obtain the to-be-determined words, and obtain the similarity degree between the to-be-determined words and the keywords in the keyword list; Mark the keyword corresponding to the maximum similarity degree as the target word; Step 3: Obtain a new number of characters m' = m - 1, repeat Step 1 and Step 2, obtain a new target word, repeat the above steps until the new number of characters is greater than a preset number-of-characters threshold; Obtain two groups corresponding to a single said target word, denoted as group i and group i-1, and the sorting position W of the first character in the group i-1 in the character sorting i-1 Is at the sorting position W i On the left side of, use the sorting position W i-1 As the target position; Step 4: Obtain all the target positions, count the number of times the target words corresponding to the target positions within the sorting position interval [WMB - W, WMB + W] appear, and obtain the target word corresponding to the maximum number of times as the first keyword, where WMB represents the sorting position of the target position in the character sorting, and W represents the preset number of sorting positions.
Citation Information
Patent Citations
Voice detection method and device, electronic equipment and storage medium
CN111369980A
Leadless group discussion system
CN114881024A
Classroom teaching behavior analysis method and system based on semantic understanding
CN119322819A