Aphasia grading and evaluation system in a multilingual scenario
By collecting and analyzing data in multilingual scenarios, and combining principal component analysis of linguistic and nonlinguistic data, the problem of inaccurate grading in existing aphasia assessment systems has been solved, achieving a more accurate and comprehensive aphasia grading assessment.
Patent Information
- Application Number
- CN202510590532.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Existing aphasia assessment systems fail to fully utilize non-verbal data, leading to inaccurate grading assessments.
A multilingual data acquisition module was used to obtain linguistic and non-linguistic data. A performance analysis module was used to obtain language organization ability coefficients and semantic understanding ability coefficients. Aphasia was then classified using principal component analysis algorithm.
It improves the accuracy and comprehensiveness of aphasia classification assessment by combining language and non-language data, reducing feature redundancy and improving assessment efficiency.
Smart Images

Figure CN120496747B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical data mining, and particularly relates to a system for grading aphasia in a multilingual scenario. BACKGROUND
[0002] Aphasia is a language disorder that is usually caused by damage to the brain, especially the areas related to language function. Patients with aphasia may have difficulty speaking fluently, understanding language, and writing. Aphasia patients often show different symptoms in different communication scenarios, affecting the daily life of patients. Early diagnosis and treatment are crucial for helping patients recover language function.
[0003] Current evaluation systems for aphasia patients are generally based on standardized language tests, which evaluate the symptoms of patients' speech output features. However, aphasia patients may exhibit speech pauses, repetitions, and other behaviors during evaluation, accompanied by non-verbal data such as facial expressions and gestures. These non-verbal data reflect the symptoms of aphasia in patients, and ignoring these factors can lead to inaccurate grading of aphasia. SUMMARY
[0004] To solve the technical problem that the existing evaluation method only focuses on the language features of patients, leading to inaccurate grading, the purpose of the present application is to provide a system for grading aphasia in a multilingual scenario, and the technical solution adopted is as follows:
[0005] Data acquisition module: acquire language data in a conversational scenario and non-verbal data in an instructional scenario for the patient; the non-verbal data includes language feature data and non-verbal cue data;
[0006] Performance analysis module: acquire a language organization ability coefficient based on the fluent language expression performance of the language data of the patient in the conversational scenario; acquire a semantic understanding ability coefficient based on the principal component features of the language feature data of the patient in the instructional scenario, combined with the principal component features of the non-verbal cue data;
[0007] Grading evaluation module: compare the language organization ability coefficient and the semantic understanding ability coefficient with the corresponding preset normal indicators to grade the aphasia of the patient.
[0008] Further, the method for acquiring the language organization ability coefficient includes:
[0009] Extracting language expression fluency-related data for each segment of speech in the language data; acquiring language expression fluency for each segment of speech based on language expression fluency-related data;
[0010] According to the overall characteristic of the language expression fluency of all speech segments of the patient in the dialogue type scene, and in combination with a time interval of each speech segment and a previous sentence speech of a dialogue person, a language organization ability coefficient of the patient is obtained; the overall characteristic of the language expression fluency is positively correlated with the language organization ability coefficient; and the time interval is negatively correlated with the language organization ability coefficient.
[0011] Further, the language expression fluency obtaining method comprises:
[0012] The related data of the language expression fluency comprises a pronunciation speed, a vocabulary repetition rate, a total number of words and a number of pauses; according to the pronunciation speed, the vocabulary repetition rate, the total number of words and the number of pauses of each extracted speech segment, the language expression fluency of each speech segment is obtained; the pronunciation speed and the total number of words are positively correlated with the language expression fluency; and the vocabulary repetition rate and the number of pauses are negatively correlated with the language expression fluency.
[0013] Further, the instruction type scene comprises an instruction feedback action scene and an instruction viewpoint statement scene.
[0014] Further, the semantic understanding ability coefficient obtaining method comprises:
[0015] The principal component scores of the language feature data and the non-language clue data of the patient in different instruction type scenes are fused to obtain the principal component score of the non-language clue data of the patient.
[0016] The product of the principal component score of the language feature data and the principal component score of the non-language clue data is taken as the semantic understanding ability coefficient.
[0017] Further, the non-language clue data at least comprises facial expression changes, action hesitation and slowness of the patient in the instruction feedback action scene.
[0018] Facial expression changes, gesture swing amplitude and logical coherence of the patient in the instruction viewpoint statement scene.
[0019] Further, the principal component score of the non-language clue data obtaining method comprises:
[0020] The principal component scores of the non-language clue data of the patient in the instruction feedback action scene and the principal component scores of the non-language clue data of the patient in the instruction viewpoint statement scene are fused to obtain the principal component score of the non-language clue data of the patient; the principal component scores of the patient in the instruction feedback action scene and in the instruction viewpoint statement scene are negatively correlated with the principal component score of the non-language clue data.
[0021] Further, the language feature data at least includes whether the patient has superfluous conversation in the execution of the instruction in the instruction type scene, whether the instruction answer is fine and smooth, and the vocabulary of the instruction answer.
[0022] Further, the method for classifying the aphasia of the patient comprises:
[0023] When the language organization ability coefficient and the semantic understanding ability coefficient are both greater than or equal to the corresponding preset normal index, the patient is determined to be a 4-level aphasia;
[0024] When the language organization ability coefficient is greater than or equal to the corresponding preset normal index, and the semantic understanding ability coefficient is less than the corresponding preset normal index, the patient is determined to be a 3-level aphasia;
[0025] When the language organization ability coefficient is less than the corresponding preset normal index, and the semantic understanding ability coefficient is greater than or equal to the corresponding preset normal index, the patient is determined to be a 2-level aphasia;
[0026] When the language organization ability coefficient and the semantic understanding ability coefficient are both less than the corresponding preset normal index, the patient is determined to be a 1-level aphasia.
[0027] Further, the conversation type scene comprises a single person daily conversation scene and a multi-person social interaction scene.
[0028] The present application has the following beneficial effects:
[0029] The present application firstly acquires the language data of the patient in the conversation type scene and the non-language data in the instruction type scene through the data acquisition module, to provide a data basis for subsequent analysis; further, in the performance analysis module, the language organization ability coefficient is acquired according to the language expression performance of the language data of the patient in the conversation type scene, to represent the strength of the language organization ability of the patient, to provide a basis for subsequent accurate aphasia classification and evaluation; further, the semantic understanding ability coefficient is acquired according to the principal component features of the language feature data of the patient in the instruction type scene, in combination with the principal component features of the non-language clue data, to acquire the semantic understanding ability coefficient from the non-language feature angle, to provide more basis for the classification and evaluation of the aphasia patient, to improve the accuracy of the evaluation result; meanwhile, in the process of calculating the semantic understanding ability coefficient, the principal component analysis algorithm is used to remove feature redundancy, to reduce the data dimension, to improve the evaluation efficiency; finally, the language organization ability coefficient and the semantic understanding ability coefficient are compared with the corresponding preset normal index through the classification and evaluation module, to classify the aphasia of the patient. The present application supplements and corrects the language expression understanding ability of the patient from the non-language angle by combining the principal component of the non-language data of the patient on the basis of the traditional analysis of the language expression in the language data, to improve the accuracy and comprehensiveness of the classification and evaluation. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, a brief introduction will be given to the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without any creative effort.
[0031] Figure 1 A system block diagram of a system for aphasia grading evaluation in a multi-language scene according to an embodiment of the present application;
[0032] Figure 2 A flowchart of an aphasia grading determination according to an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined purposes, the specific embodiments, structures, features and effects of a system for aphasia grading evaluation in a multi-language scene according to the present application are described in detail below in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0035] The specific scheme of a system for aphasia grading evaluation in a multi-language scene according to the present application is described in detail below in combination with the drawings.
[0036] Please refer to Figure 1 which shows a system block diagram of a system for aphasia grading evaluation in a multi-language scene according to an embodiment of the present application. The system comprises a data acquisition module 101, a performance analysis module 102 and a grading evaluation module 103.
[0037] The data acquisition module 101 acquires language data of the patient in a dialogue scene and non-language data in an instruction scene. The non-language data includes language feature data and non-language clue data.
[0038] In the embodiment of the present application, different language scenes are set, and multi-modal sensors are used to collect language data and non-language data of the patient in different scenes.
[0039] In an embodiment of the present application, the set language scenes include:
[0040] Single-person daily conversation scenario: having a conversation with the patient on topics in daily life (such as family, food, etc.).
[0041] Multi-person social interaction scenario: responding to different roles by having the patient chat with multiple people.
[0042] Instruction feedback action scenario: giving the patient some specific instructions (such as "please pass me that book, etc.).
[0043] Instruction opinion statement scenario: giving an opinion and asking the patient to state their opinion (such as "what do you think").
[0044] Among the conversation scenarios, the single-person daily conversation scenario and the multi-person social interaction scenario focus on evaluating the patient's language organization ability; among the instruction scenarios, the instruction feedback action scenario and the instruction opinion statement scenario focus on evaluating the patient's semantic understanding ability.
[0045] During the evaluation process, the language data (such as voice data) and non-language data of the patient in different scenarios are collected through multi-modal sensors (microphone, camera, etc.), and different behavior change nodes are represented on the time axis. Thus, by combining the performance of the patient in multiple dimensions, the possibility of misjudgment is minimized. The non-language data includes language feature data (such as whether the instruction answer is fine and smooth, the vocabulary of the instruction answer, etc.) and non-language cue data (such as expressions, gestures, etc.), which can be collected manually by relevant personnel to provide a data basis for subsequent analysis.
[0046] It should be noted that the setting method of various language scenarios is a known technical means to those skilled in the art, and the number of conversations or instructions for each scenario can be set by the implementer, but the number of scenarios set for different patients needs to be consistent, which will not be limited here.
[0047] Performance analysis module 102: according to the language expression fluency performance of the language data of the patient in the conversation scenario, the language organization ability coefficient is obtained; according to the principal component features of the language feature data of the patient in the instruction scenario, combined with the principal component features of the non-language cue data, the semantic understanding ability coefficient is obtained.
[0048] The language features of the patient are the most intuitive manifestation of aphasia symptoms and can best demonstrate the patient's language organization ability and semantic understanding ability. The non-language data of the patient is generally used to supplement the language description, so the language organization ability of the patient is first determined according to the voice features and language feedback performance of the patient in different scenarios, and then the language understanding ability of the patient is supplemented and corrected according to the non-language data, thereby improving the accuracy of the aphasia grading evaluation of the patient.
[0049] Considering that the fluent performance of language expression of the patient in the dialogue type scene reflects the language organization ability of the patient, a language organization ability coefficient is first obtained according to the fluent performance of language expression of the patient in the dialogue type scene, to represent the strength of the language organization ability of the patient, and to provide a basis for subsequent accurate aphasia grading assessment.
[0050] Preferably, in an embodiment of the present application, considering that the language data cannot directly quantify the fluent performance of language expression, the related data of the fluent performance of language expression of each voice in the language data is first extracted; the fluent degree of language expression of each voice is obtained based on the related data of the fluent performance of language expression, to provide a basis for analyzing the language organization ability of the patient from the perspective of the fluent degree of language expression, and the greater the fluent degree of language expression, the stronger the language organization ability;
[0051] Considering that the time interval between each voice of the patient and the previous voice of the dialogue personnel represents the reaction time of the patient to the question and the dialogue, and also represents the language organization ability of the patient, the longer the time interval, the longer the reaction time, and the worse the language organization ability, the time interval between each voice and the previous voice of the dialogue personnel is also analyzed;
[0052] Therefore, according to the overall feature of the fluent degree of language expression of all voice segments of the patient in the dialogue type scene, and in combination with the time interval between each voice and the previous voice of the dialogue personnel, the language organization ability coefficient of the patient is obtained; the overall feature of the fluent degree of language expression is positively correlated with the language organization ability coefficient; the time interval is negatively correlated with the language organization ability coefficient.
[0053] As an example, the mean value of the fluent degree of language expression of all voice segments of the patient in the dialogue type scene is taken as the overall fluent degree of expression, to represent the overall feature of the fluent degree of language expression of all voice segments of the patient in the dialogue type scene; the sum value of the time interval between each voice of the patient and the previous voice of the dialogue personnel is taken as the overall reaction time, to represent the time interval of all voice segments; the ratio of the overall fluent degree of expression to the overall reaction time is taken as the language organization ability coefficient of the corresponding patient; and the overall reaction time is in the denominator position.
[0054] It should be noted that the voice of a certain answer or speech of the patient is a voice segment; the difference between the starting time of each voice segment and the ending time of the previous voice of the dialogue personnel is taken as the time interval between each voice segment and the previous voice of the dialogue personnel; in the process of calculating the language organization ability coefficient, only the numerical value of the data is taken;
[0055] In other embodiments of the present application, the implementer can also obtain the overall expression fluency by using weighted summation of the average, mode and median of the language expression fluency, such as weights of 0.5, 0.25 and 0.25; and can further normalize the language organization ability coefficient, which can be specifically normalized by using linear normalization.
[0056] Preferably, in an embodiment of the present application, the faster the pronunciation speed, the lower the vocabulary repetition rate, the higher the total number of words in the speech segment and the fewer the number of pauses, the more fluent the language expression of the patient, and the higher the language expression fluency.
[0057] Therefore, the related data of the language expression fluency includes the pronunciation speed, the vocabulary repetition rate, the total number of words and the number of pauses; the language expression fluency of each speech segment is obtained according to the pronunciation speed, the vocabulary repetition rate, the total number of words and the number of pauses of each speech segment extracted; the pronunciation speed and the total number of words are positively correlated with the language expression fluency; and the vocabulary repetition rate and the number of pauses are negatively correlated with the language expression fluency.
[0058] As an example, the product of the difference between the total number of words and the number of pauses of each speech segment and the ratio of the pronunciation speed to the vocabulary repetition rate is taken as the language expression fluency of each speech segment.
[0059] It should be noted that the speech data is converted into text by using a speech recognition technology such as a Transformer model or a long short-term memory network model, the total number of words is counted, and the ratio of the total number of words to the duration of the speech segment is taken as the pronunciation speed; the text is segmented by using a JieBa segmentation tool, and the repeated words are counted by using a text analysis tool such as a dictionary to obtain the ratio of the number of repeated words to the total number of words as the vocabulary repetition rate; the number of pauses is counted by counting the number of times that the energy is lower than an energy threshold and lasts for more than 200 ms in the speech segment, and the energy threshold can be set as 20% of the highest energy, and the speech energy of each sampling point is the square of the amplitude of each sampling point.
[0060] In other embodiments of the present application, the implementer can also take the product of the pronunciation speed and the total number of words as the numerator, take the sum of the product of the number of pauses and the vocabulary repetition rate and a preset non-zero constant 1 as the denominator, and take the fractional ratio as the language expression fluency; and the difference between the pronunciation speed, the vocabulary repetition rate, the total number of words and the number of pauses can also be fused by using weighted summation, and the algorithm and the weighted summation method used in the language expression fluency data acquisition method are all prior art, and will not be described in detail.
[0061] Then, the non-verbal data of the patient in the instruction type scene is analyzed to evaluate the semantic understanding ability of the patient. The semantic understanding ability is closely related to the language feedback of the patient in the instruction type scene. Since the semantic understanding ability of the patient has different dimensional characteristics, principal component analysis (PCA) is needed to extract the main features of the language characteristics of different patients in the instruction type scene, so as to reduce the interference of redundant information.
[0062] The principal component feature represents the core feature of the language and non-verbal performance of the patient in the instruction type scene, and the redundant information of the data has been removed, and the most important dimension of the semantic understanding ability is highlighted. Therefore, according to the principal component feature of the language feature data of the patient in the instruction type scene, combined with the principal component feature of the non-verbal clue data, the semantic understanding ability coefficient is obtained from the non-verbal feature, which provides more basis for the grading evaluation of aphasia patients and improves the accuracy of the evaluation result. At the same time, in the process of calculating the semantic understanding ability coefficient, the principal component analysis algorithm is used to remove feature redundancy, reduce data dimension, and improve evaluation efficiency.
[0063] Preferably, in an embodiment of the present application, the principal component score directly quantifies the principal component feature, so the principal component score is used to represent the principal component feature; considering that the instruction type scene is divided into instruction feedback action scene and instruction opinion statement scene, the attention points of the two instruction type scenes to the non-verbal clues of the patient are not completely the same, so the principal component scores of the patient in different instruction type scenes are calculated first, and then fused to obtain the principal component score of the non-verbal clue data of the patient.
[0064] As an example, the language feature data at least includes: whether there is redundant conversation in the patient's execution of instructions in the instruction type scene, whether the instruction answer is fine and smooth, and the vocabulary of the instruction answer. The principal component score of the language feature data is obtained by the following method:
[0065] First, collect the language feature data of the patient in each test in the instruction feedback action scene and the instruction opinion statement scene; for example, [0, 0, 12], 1 means yes, 0 means no, 1 means no redundant conversation, 1 means fine and smooth instruction answer, [0, 0, 12] means redundant conversation, the instruction answer is not fine and smooth, and the vocabulary of the instruction answer is 12; the vocabulary can be converted into text by voice recognition technology, and the number of words in the dictionary is obtained by establishing a dictionary after word segmentation; whether there is redundant conversation in the execution of instructions, whether the instruction answer is fine and smooth, and the vocabulary of the instruction answer are recorded by relevant personnel.
[0066] Further, different language feature data is standardized, that is, each feature data is subtracted by its mean value and then divided by its standard deviation to standardize and eliminate the dimension and value range difference of different features.
[0067] Further perform principal component analysis, calculate the covariance matrix and eigenvalues and eigenvectors, select the first N principal components that can explain most of the variation, set the cumulative variance contribution rate of the principal components to reach 70%; Different dimensions (principal components) can be obtained to explain the semantic understanding of aphasia patients:
[0068] The first principal component (PC1) represents the understanding of vocabulary or sentence structure;
[0069] The second principal component (PC2) represents a dimension related to language fluency and grammatical correctness;
[0070] The third principal component (PC3) represents a dimension related to language fluency and vocabulary size;
[0071] Since each principal component corresponds to a dimension of semantic understanding ability, the principal component score of the language feature data represents the patient's ability strength in understanding voice instructions. Specifically, multiply the principal component feature vector by the standardized data matrix, and the product is the score vector of each principal component. The average of the data in the score vector is the average score of each principal component. According to the variance contribution rate, set different principal component weights, and obtain the principal component score by weighting and summing the average score.
[0072] For example, the standardized data matrix The feature vector of the first principal component The feature vector of the second principal component The obtained principal component score vector is The average score corresponding to v1 is 0.37, and the average score corresponding to v2 is -0.09. Assuming that the variance contribution rates of the first principal component and the second principal component are 0.56 and 0.24, respectively, and after normalization, they are 0.7 and 0.3, the obtained principal component score C = 0.37 x 0.7 + (-0.09) x 0.3 = 0.24.
[0073] Preferably, in an embodiment of the present application, non-verbal cues are also considered to help the evaluator understand the patient's intentions. If these non-verbal signals are ignored when evaluating aphasia, the patient's semantic understanding ability may be underestimated, so further consider the non-verbal cues, and supplement the semantic understanding ability.
[0074] In the instruction feedback action scenario, the patient assesses his semantic understanding ability by executing the instruction. Ideally, this scenario does not involve language feedback. If the patient's facial expression shows confusion, eye wandering and other related expressions after receiving the instruction, it means that the patient may not understand the meaning of the instruction and the semantic understanding ability is poor. At the same time, since the patient needs to execute the instruction, if the patient understands the instruction, he will show the intention to prepare to execute the instruction, otherwise the patient may show hesitation and avoidance gestures, so the non-verbal cues that may be generated include expressions and behaviors.
[0075] It should be noted that the patient's understanding behavior is not just the existence of two cases of understanding and not understanding, but also the case of gradually understanding after not understanding at the beginning, so the final result of the patient's execution of the instruction needs to be evaluated.
[0076] As an example, the non-verbal cue data at least includes: facial expression changes of the patient in the instruction feedback action scenario, whether the action is hesitant and slow; for example, facial expression changes (0 means no obvious expression change, 1 means subtle expression change, 2 means obvious expression change such as confusion, tension), whether the action is hesitant and slow (0 means smooth execution, 1 means slight delay in action, 2 means obvious delay or pause in action), [0, 0] means that the patient has no expression change and the action is not delayed, and the patient fully understands the instruction.
[0077] In the same way of obtaining the principal component scores of the language feature data, the non-verbal cue data in the instruction feedback action scenario is input to obtain the principal component scores of the non-verbal cue data of the patient in the instruction feedback action scenario, which represents the richness of the non-verbal cue. The higher the score, the richer the non-verbal cue of the patient in understanding the instruction, which reflects that the patient's understanding of the instruction is lower and the semantic understanding ability is poorer.
[0078] Considering that the instruction opinion statement scenario is an open-ended question and answer, the patient not only needs to understand the instruction, but also needs to actively organize language to express his own opinion. If the patient's expression is natural when answering, it means that the patient can express himself confidently after understanding the instruction, and the semantic understanding ability is strong; on the contrary, if the expression is exaggerated and dull, it means that the patient may have semantic understanding obstacles and needs time to process information, and tries to use expression to compensate for the difficulty of language organization; similarly, if the gesture swing amplitude is too large, random and uncoordinated, it may mean that the patient needs additional body assistance to compensate for the lack of semantic understanding during expression; the logical connection of the patient also shows the information processing and expression ability of the patient;
[0079] Based on this, the patient's facial expression change, gesture swing amplitude and logical coherence in the instruction opinion statement scene. For example, facial expression change (0: natural, 1: slightly exaggerated, 2: very exaggerated / stiff), gesture swing amplitude (0: normal, 1: slightly random, 2: obviously uncoordinated or too large), logical coherence (0: logical clarity, no interruption, 1: less than two interruptions, 2: more than two logical interruptions); [1, 1, 1] indicates that the patient's expression is slightly exaggerated, the gesture is slightly random, and the logic is occasionally interrupted (less than twice).
[0080] Similarly, the non-verbal cue data in the instruction opinion statement scene is input in the same way as obtaining the principal component score of the language feature data, and the principal component score of the non-verbal cue data of the patient in the instruction opinion statement scene is obtained; the higher the principal component score, the richer the non-verbal cue, and the lower the understanding of the patient to the instruction, the worse the corresponding semantic understanding ability.
[0081] Further, the principal component score of the non-verbal cue data of the patient in the instruction feedback action scene is fused with the principal component score of the non-verbal cue data of the patient in the instruction opinion statement scene to obtain the principal component score of the non-verbal cue data of the patient.
[0082] Wherein the principal component scores of the patient in the instruction feedback action scene and in the instruction opinion statement scene are greater, both indicating that the semantic understanding ability of the patient is worse, so the principal component scores of the patient in the instruction feedback action scene and in the instruction opinion statement scene are negatively correlated with the principal component score of the non-verbal cue data.
[0083] As an example, the product of the principal component scores of the non-verbal cues in the two instruction scenes is mapped to the reciprocal negative correlation after the sum of the preset zero normal number 0.1, and the principal component score of the non-verbal cue data is obtained.
[0084] The product of the principal component scores of the language feature data and the principal component scores of the non-verbal cue data is taken as the semantic understanding ability coefficient.
[0085] It should be noted that in other embodiments of the present application, the principal component scores of the language feature data and the principal component scores of the non-verbal cue data can also be fused by weighted sum and the like to obtain the semantic understanding ability coefficient; the semantic understanding ability coefficient can also be further normalized, such as linear normalization; the variance contribution rate can be selected between 70% and 90% to select the principal component;
[0086] The implementer can also set other language feature data, such as the number of correct sentences, the number of valid vocabularies (excluding stop words, catchphrases, etc.), the number of redundant dialogues (such as repeated instruction dialogues), and the like, and replace or add them; in the instruction feedback action scene, other non-verbal cue data such as the number of expression changes, the number of eye wandering, the number of action pauses, and whether to actively execute can also be set, replaced or added; in the instruction viewpoint statement scene, other non-verbal cue data such as the number of logical interruptions and the number of expression changes can also be set, replaced or added.
[0087] In another embodiment of the present application, the implementer can also multiply the feature vector of the principal component with the standardized data matrix, sum all the score vectors corresponding to the principal components, and take the modulus to obtain the score of the corresponding principal component.
[0088] Considering that the more the actions of the patient in the instruction feedback action scene conform to the requirements of the instruction, the shorter the thinking or brewing time of the patient in the instruction viewpoint statement scene, the stronger the semantic understanding ability of the patient, the implementer can also take the product of the number of conforming instructions, the inverse of the brewing time and value of all viewpoints, the principal component score of the language feature data, and the principal component score of the non-verbal cue data as the semantic understanding ability coefficient. The weighted sum can also be used to fuse to obtain the semantic understanding ability coefficient.
[0089] The hierarchical assessment module 103: compares the language organization ability coefficient and the semantic understanding ability coefficient with the corresponding preset normal indicators to grade the aphasia of the patient.
[0090] In the performance analysis module 102, the language data and non-verbal data of the patient in different scenes are analyzed, the language organization ability coefficient and the semantic understanding ability coefficient of the patient are evaluated from the language expression angle and the non-verbal cue angle, and the aphasia grading basis is obtained, so the language organization ability coefficient and the semantic understanding ability coefficient are further compared with the corresponding preset normal indicators, and the patient is automatically graded for aphasia.
[0091] Preferably, in one embodiment of the present application, the method for grading the aphasia of the patient comprises: Figure 2 which shows a kind of aphasia grading determination flow chart provided by one embodiment of the present application, specifically comprising:
[0092] When the language organization ability coefficient and the semantic understanding ability coefficient are greater than or equal to the corresponding preset normal indicators, the patient is determined to be a 4-level aphasia;
[0093] When the language organization ability coefficient is greater than or equal to the corresponding preset normal indicators, and the semantic understanding ability coefficient is less than the corresponding preset normal indicators, the patient is determined to be a 3-level aphasia;
[0094] When the language organization ability coefficient is less than the corresponding preset normal index, and the semantic understanding ability coefficient is greater than or equal to the corresponding preset normal index, the patient is determined to be a 2-level aphasia;
[0095] When the language organization ability coefficient and the semantic understanding ability coefficient are both less than the corresponding preset normal index, the patient is determined to be a 1-level aphasia.
[0096] Figure 2 In the formula, A represents the difference between the language organization ability coefficient and the corresponding preset normal index, and B represents the difference between the semantic understanding ability coefficient and the corresponding preset normal index; the aphasia levels are in descending order, and the symptoms are in ascending order, that is, the 4-level aphasia has the lightest symptoms, and the 1-level aphasia has the heaviest symptoms.
[0097] It should be noted that the two preset normal indexes are obtained by collecting language data and non-language data of a plurality of normal persons, obtaining the language organization ability coefficient and the semantic understanding ability coefficient, and setting the mean value of a certain proportion of the minimum data as the preset normal index, for example, collecting data of 100 normal persons, and setting the mean value of the minimum 10% data as the preset normal index of the corresponding type of data.
[0098] It should be noted that the obtained aphasia grading result is only for reference by doctors and other relevant personnel, and provides some data support for the intervention or treatment of the patient by the relevant personnel.
[0099] To sum up, in view of the technical problem that the existing evaluation method only focuses on the language characteristics of the patient, leading to inaccurate grading evaluation, the present application provides an aphasia grading evaluation system in a multi-language scene. The present application first obtains the language data of the patient in a dialogue type scene and the non-language data in an instruction type scene through a data acquisition module; further obtains the language organization ability coefficient according to the fluent language expression performance of the patient in the dialogue type scene in the performance analysis module; obtains the semantic understanding ability coefficient according to the principal component features of the language characteristic data of the patient in the instruction type scene, in combination with the principal component features of the non-language clue data; and finally compares the language organization ability coefficient and the semantic understanding ability coefficient with the corresponding preset normal index through a grading evaluation module to grade the aphasia of the patient. The present application supplements and corrects the language expression understanding ability of the patient from the non-language perspective by combining the principal components of the non-language data of the patient on the basis of the traditional analysis of the language expression in the language data, thereby improving the accuracy and comprehensiveness of the evaluation.
[0100] It should be noted that the above-mentioned embodiment order of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or can be advantageous.
[0101] The various embodiments described in this specification are presented by way of example, and each embodiment is not inherently more important than any other embodiment. To the extent that any embodiment is directed to a distinct, independently applicable inventive concept, it is to be understood that the inventive concept(s) can be realized in a multitude of alternative ways. Each embodiment is presented for the purpose of illustrating one or more aspects of the inventive concept(s), and the application should not be construed as requiring that the inventive concept(s) be limited to only those embodiments.
Claims
1. A multilingual aphasia grading assessment system, characterized by: The system comprises: a data acquisition module: acquiring language data of a patient in a dialogue scene and non-language data in an instruction scene; the non-language data comprises language feature data and non-language clue data; a performance analysis module: acquiring a language organization ability coefficient according to language expression fluency performance of the language data of the patient in the dialogue scene; acquiring a semantic understanding ability coefficient according to principal component features of the language feature data of the patient in the instruction scene, in combination with principal component features of the non-language clue data; a grading evaluation module: comparing the language organization ability coefficient and the semantic understanding ability coefficient with corresponding preset normal indicators to grade aphasia of the patient; the instruction scene comprises an instruction feedback action scene and an instruction viewpoint statement scene; the method for acquiring the semantic understanding ability coefficient comprises: fusing principal component scores of the patient in different instruction scenes to obtain principal component scores of the non-language clue data of the patient; and taking a product of the principal component scores of the language feature data and the principal component scores of the non-language clue data as a semantic understanding ability coefficient; the non-language clue data at least comprises facial expression changes, action hesitation and slowness of the patient in the instruction feedback action scene; facial expression changes, gesture swing amplitude and logical coherence of the patient in the instruction viewpoint statement scene; the language feature data at least comprises whether the patient has superfluous dialogue in executing an instruction, whether instruction answering is fine and fluent, and a vocabulary of instruction answering; the method for grading aphasia of the patient comprises: when the language organization ability coefficient and the semantic understanding ability coefficient are both greater than or equal to corresponding preset normal indicators, determining that the patient is a 4th-grade aphasia patient; when the language organization ability coefficient is greater than or equal to a corresponding preset normal indicator and the semantic understanding ability coefficient is less than a corresponding preset normal indicator, determining that the patient is a 3rd-grade aphasia patient; when the language organization ability coefficient is less than a corresponding preset normal indicator and the semantic understanding ability coefficient is greater than or equal to a corresponding preset normal indicator, determining that the patient is a 2nd-grade aphasia patient; and when the language organization ability coefficient and the semantic understanding ability coefficient are both less than corresponding preset normal indicators, determining that the patient is a 1st-grade aphasia patient.
2. The system for aphasia grading evaluation in a multi-lingual scenario as claimed in claim 1 wherein, the method for acquiring the language organization ability coefficient comprises: extracting language expression fluency related data of each voice segment in the language data; and acquiring language expression fluency of each voice segment based on the language expression fluency related data; acquiring the language organization ability coefficient of the patient according to overall features of the language expression fluency of all voice segments of the patient in the dialogue scene, in combination with time intervals between each voice segment and a previous voice segment of a dialogue person; the overall features of the language expression fluency are positively correlated with the language organization ability coefficient; and the time intervals are negatively correlated with the language organization ability coefficient.
3. The system for aphasia grading evaluation in a multi-lingual scenario as claimed in claim 2 wherein, the method for acquiring the language expression fluency comprises: The language expression fluency related data includes pronunciation speed, vocabulary repetition rate, total word number and pause number; according to the pronunciation speed, vocabulary repetition rate, total word number and pause number of each extracted speech, the language expression fluency of each speech is obtained; the pronunciation speed and the total word number are positively correlated with the language expression fluency; the vocabulary repetition rate and the pause number are negatively correlated with the language expression fluency.
4. The system for aphasia grading evaluation in multi-language scenarios according to claim 1, wherein, The method for obtaining the principal component score of the non-verbal cue data comprises: The principal component score of the non-verbal cue data of the patient in the instruction feedback action scene and the principal component score of the non-verbal cue data of the patient in the instruction opinion statement scene are fused to obtain the principal component score of the non-verbal cue data of the patient; the principal component score of the patient in the instruction feedback action scene and the principal component score of the patient in the instruction opinion statement scene are negatively correlated with the principal component score of the non-verbal cue data.
5. The system for aphasia grading evaluation in multi-language scenarios according to claim 1, wherein, The dialogue type scene includes a single person daily conversation scene and a multi-person social interaction scene.
Citation Information
Patent Citations
System for diagnosing post-stroke aphasia
CN113598712A
System and method for detection of cognitive and speech impairment based on temporal visual facial feature
US20200237290A1