Method and device for rapid identification based on stroke symptoms

Through a rapid identification method based on video and audio data, using a variety of biometric information to evaluate stroke symptoms, the problems of long diagnosis time and high cost in the prior art are solved, and rapid, accurate and economical stroke screening are achieved.

CN119132611BActive Publication Date: 2025-05-09TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411311971.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-05-09
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

Existing stroke diagnosis methods rely on the experience of professional doctors and complex medical examinations, resulting in long diagnosis, high cost and unsuitable for large-scale screening.

Method used

Through a rapid identification method based on video and audio data, using pose estimation, MFCC algorithm, Dlib detector and other technologies, subjects' stroke symptoms are evaluated from multiple angles, including fluency in both hands trajectory movement, speech disorders, and facial feature point analysis.

Benefits of technology

The preliminary screening of stroke symptoms is achieved in a short period of time, which reduces the cost of equipment, is suitable for large-scale screening, and does not require direct contact with the subject, which improves the accuracy and reliability of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132611B_ABST
    Figure CN119132611B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for rapid identification of stroke symptoms, and a computing device, wherein the method comprises: extracting the coordinates of the key points of the hands of the subject from the video to be detected according to the posture estimation method; calculating the smoothness of the movement of the hand trajectory according to the coordinates of the key points of the hands; extracting the MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; identifying the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression; obtaining facial feature points including face and mouth from the video to be detected by the Dlib detector; calculating the distance between the center point of the mouth corner and the lip and the distortion angle of the mouth corner according to the facial feature points; judging whether the subject has stroke symptoms according to the smoothness of the movement of the hand trajectory, the speech features, the center point distance and the distortion angle of the mouth corner. The present invention realizes the rapid judgment of stroke symptoms without labeled samples, and improves the efficiency and accuracy of recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical auxiliary diagnosis, and in particular to a method and device for rapid identification of stroke symptoms, and a computing device. Background Art

[0002] Stroke, also known as cerebral infarction, is a type of cerebral blood circulation disorder with sudden fainting, unconsciousness, crooked mouth, speech impairment and hemiplegia as the main symptoms. Due to the high incidence, mortality, disability, recurrence and complications of stroke, the medical community lists it as one of the three major diseases that threaten human health, along with coronary heart disease and cancer. Stroke poses a great threat to human health and life, and brings great pain to patients. For stroke patients, time is life. If symptoms can be discovered in a short period of time and timely treatment is given, the disability and mortality rates caused by stroke can be greatly reduced.

[0003] Existing stroke diagnosis methods mainly rely on the experience of clinicians and a series of neurological function tests, such as NIHSS (National Institutes of Health Stroke Scale) scores, CT or MRI imaging examinations, etc. Although it has high accuracy in diagnosing stroke, it takes a long time to complete a series of tests and imaging examinations, which may delay the best treatment time for acute stroke patients. In addition, the diagnostic process is highly dependent on the experience and technology of professional doctors, which is costly and difficult to popularize to ordinary users. Alternatively, based on a large number of labeled stroke samples, stroke symptoms are predicted through deep learning models, but labeling stroke samples requires professional doctors or specially trained personnel, resulting in high labeling costs.

[0004] To solve the above problems, the present invention proposes a rapid recognition method based on stroke symptoms to achieve rapid stroke symptom judgment without labeled samples and improve recognition efficiency and accuracy. Summary of the invention

[0005] 1. Technical issues to be resolved

[0006] The present invention provides a method for quickly identifying stroke symptoms, which can quickly identify stroke symptoms without labeling stroke samples.

[0007] (II) Technical solution

[0008] According to one aspect of the present invention, a method for rapid identification based on stroke symptoms is provided, comprising:

[0009] Extracting the coordinates of the key points of the hands of the subject from the video to be detected according to the posture estimation method; calculating the smoothness of the movement of the hand trajectory according to the coordinates of the key points of the hands;

[0010] Extracting MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; identifying the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression;

[0011] Obtaining facial feature points including the face and the mouth from the video to be detected by using a Dlib detector; calculating the distance between the corner of the mouth and the center point of the lips and the distortion angle of the corner of the mouth according to the facial feature points;

[0012] Whether the subject has stroke symptoms is determined based on the smoothness of the two-hand trajectory movement, the speech characteristics, the center point distance, and the distortion angle of the mouth corners.

[0013] In an optional manner, the calculation formula for the smoothness of the two-hand trajectory movement is:

[0014]

[0015] Among them, k is the scale of fluency; ∈ is used to control the maximum value of fluency;

[0016] Var L , Var R are the variances of the acceleration changes of the left and right hands, respectively, reflecting the degree of fluctuation of the acceleration change;

[0017]

[0018]

[0019] Δa L (t) = a L (t+Δt)-a L (t)

[0020] Δa R (t) = a R (t+Δt)-a R (t)

[0021] Where n is the number of points in the acceleration sequence used to calculate the difference; μ L and μ R are the average values ​​of the acceleration changes of the left and right hands respectively; a L (t) is the acceleration of the left hand at time t; a R (t) is the acceleration of the right hand at time t; Δt is the time difference between two adjacent frames.

[0022] In an optional manner, the approximate expression of the acceleration of the left hand at time t is:

[0023]

[0024] Among them, x L (t) is the position coordinate of the left hand at time t;

[0025] The approximate expression for the acceleration of the right hand at time t is:

[0026]

[0027] Among them, x R (t) is the position coordinate of the right hand at time t.

[0028] In an optional manner, identifying the speech disorder of the subject according to the speech feature vector sequence further includes:

[0029] Calculating the coefficient variance of the speech feature vector sequence, and if the coefficient variance deviates from a preset normal range, determining that the subject's speech disorder is unclear speech;

[0030] The zero-crossing rate of the speech feature vector sequence is calculated, and if the zero-crossing rate is higher than a preset threshold, it is determined that the subject's speech disorder is expression difficulty.

[0031] In an optional manner, the expression of the zero crossing rate is:

[0032]

[0033] Among them, x″ t is the speech feature vector x t The second derivative at time index t, x t is the audio signal value at time point t; Ind{} is an indicator function that returns 1 when the condition is met, otherwise it returns 0.

[0034] In an optional manner, calculating the distance between the mouth corner and the center point of the lips and the distortion angle of the mouth corner according to the facial feature points further includes:

[0035] Estimate the average position of the mouth corner points according to the positions of the left and right mouth corners to obtain the lip center point; calculate the distance between the mouth corner and the lip center point using the Euclidean distance formula;

[0036] The distortion angle of the mouth corner is obtained by the angle between the direction of the line connecting the mouth corner point and the lip center point and the horizontal direction.

[0037] In an optional manner, the distance between the mouth corner and the lip center point is:

[0038]

[0039] Where n is the number of feature points; P i is the coordinate of the ith feature point; C is the coordinate of the center point of the lips; (P i , C) is the distance from the i-th feature point to the lip center point C; w i is the offset coefficient of the i-th feature point, |y i -y C | is the feature point P i The vertical coordinate y i The ordinate y of the center point C C The absolute offset difference.

[0040] In an optional manner, the expression for the distortion angle of the mouth corner is:

[0041] Δθ=|θ R -θ L |

[0042] in, It is the angle between the line connecting the right corner of the mouth and the center point and the horizontal direction; It is the angle between the line connecting the left corner of the mouth and the center point and the horizontal direction.

[0043] According to another aspect of the present invention, there is provided a device for rapid identification based on stroke symptoms, comprising:

[0044] The arm analysis module is used to extract the coordinates of the key points of the hands of the subject from the video to be detected according to the posture estimation method; and calculate the smoothness of the movement of the two hands trajectory according to the coordinates of the key points of the two hands;

[0045] A speech disorder analysis module, used to extract MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; and identify the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression;

[0046] A facial feature analysis module is used to obtain facial feature points including the face and the mouth from the video to be detected through a Dlib detector; and calculate the distance between the corner of the mouth and the center point of the lips and the distortion angle of the corner of the mouth according to the facial feature points;

[0047] The stroke symptom judgment module is used to judge whether the subject has stroke symptoms according to the smoothness of the two-hand trajectory movement, the voice characteristics, the center point distance and the distortion angle of the mouth corners.

[0048] According to another aspect of the present invention, there is provided a computing device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus;

[0049] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute operations corresponding to the above-mentioned method for rapid identification of stroke symptoms.

[0050] (III) Beneficial effects

[0051] (1) By integrating multiple biometric information such as gesture analysis, voice feature extraction, and facial feature detection, the subject's stroke symptoms can be comprehensively assessed from multiple angles, thereby improving the accuracy and reliability of recognition.

[0052] (2) It can complete the initial screening of stroke symptoms in a short period of time without the need for complex medical examination equipment. It can quickly draw conclusions through real-time processing of video and audio data, shorten the diagnosis time, and gain precious time for timely treatment.

[0053] (3) The present invention mainly relies on video and audio data and does not require direct contact with the subject. Compared with traditional imaging examinations such as CT or MRI, it has less physical burden on patients and is more suitable for large-scale screening.

[0054] (4) Only an ordinary mobile phone APP is needed to perform the test, which reduces the equipment cost. It is not only suitable for the professional environment of medical institutions, but also suitable for homes, schools and other places, and has a wider applicability.

[0055] (5) By extracting the coordinates of the key points of both hands through the posture estimation method, the details of hand movements can be accurately captured, improving the accuracy of gesture analysis. The movement smoothness is evaluated by calculating the variance of the acceleration change, the speech disorder is identified by calculating the coefficient variance and zero crossing rate of the speech feature vector sequence, and the facial symmetry is evaluated by calculating the distance between the mouth corner and the lip center point and the distortion angle, further improving the accuracy of stroke recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0057] Figure 1 A schematic diagram showing a flow chart of a method for rapid identification of stroke symptoms according to an embodiment of the present invention is shown;

[0058] Figure 2A schematic diagram of MFCC features according to an embodiment of the present invention is shown;

[0059] Figure 3 A schematic diagram of speech impairment according to an embodiment of the present invention is shown;

[0060] Figure 4 A schematic diagram showing the distorted angle of the mouth corner according to an embodiment of the present invention is shown;

[0061] Figure 5 A schematic diagram of the trajectory movement of both hands according to an embodiment of the present invention is shown;

[0062] Figure 6 A structural block diagram of a device for rapid identification of stroke symptoms according to an embodiment of the present invention is shown;

[0063] Figure 7 A schematic diagram of the structure of a computing device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0064] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.

[0065] Figure 1 The flowchart of the method for rapid identification of stroke symptoms according to an embodiment of the present invention is shown. Specifically, the method includes the following steps.

[0066] Step S101, extracting the coordinates of the key points of the hands of the subject from the video to be detected according to the posture estimation method; and calculating the smoothness of the movement trajectory of the hands according to the coordinates of the key points of the hands.

[0067] For example, open the "Stroke Quick Check" app and enter the main interface. An animation or text prompt is displayed on the screen, such as: "Please raise your hands and keep them parallel to your chest for 10 seconds." The user raises his hands according to the prompt and keeps them for a while. After the gesture analysis is completed, the app prompts the user to perform a voice test: "Please say 'ah' loudly for 5 seconds." The user says "ah" according to the prompt, and the app collects the user's voice signal through the microphone.

[0068] Specifically, the coordinates of the key points of the hands are extracted using posture estimation methods such as OpenPose and MediaPipe, which can maintain high recognition accuracy even in complex backgrounds. By calculating the variance of the acceleration change to evaluate the smoothness of the movement, abnormal hand movements can be more accurately identified. Figure 5As shown, the video file can be read through cv2.VideoCapture, the Hands module of MediaPipe is used for posture estimation, the coordinates of the key points of the subject's hands (including the positions of the finger tips, wrists, etc.) are extracted from the video frame, and the trajectory list is updated.

[0069] In this embodiment, the calculation formula for the smoothness of the two-hand trajectory movement is:

[0070]

[0071] Among them, k is the scale of fluency; ∈ is used to control the maximum value of fluency;

[0072] Var L , Var R are the variances of the acceleration changes of the left and right hands, respectively, reflecting the degree of fluctuation of the acceleration change;

[0073]

[0074]

[0075] Δa L (t) = a L (t+Δt)-a L (t)

[0076] Δa R (t) = a R (t+Δt)-a R (t)

[0077] Where n is the number of points in the acceleration sequence used to calculate the difference; μ L and μ R are the average values ​​of the acceleration changes of the left and right hands respectively; a L (t) is the acceleration of the left hand at time t; a R (t) is the acceleration of the right hand at time t; Δt is the time difference between two adjacent frames.

[0078] In this embodiment, the approximate expression of the acceleration of the left hand at time t is:

[0079]

[0080] Among them, x L (t) is the position coordinate of the left hand at time t;

[0081] The approximate expression for the acceleration of the right hand at time t is:

[0082]

[0083] Among them, xR (t) is the position coordinate of the right hand at time t.

[0084] By calculating the approximate value of acceleration, it is possible to more accurately capture the dynamic changes of the hand, especially the subtle acceleration changes of the hand, which helps to identify potential stroke symptoms.

[0085] Step S102, extracting MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; identifying the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression.

[0086] MFCC (Mel Frequency Cepstral Coefficient) can capture the spectral characteristics of speech signals. Even in a noisy environment, MFCC features can provide reliable speech feature representation, thereby more accurately identifying speech disorders. : The MFCC algorithm has a high computational efficiency and can complete the extraction of speech features in a short time. The specific steps of MFCC feature extraction are as follows:

[0087] The speech signal is pre-emphasized to enhance the high frequency part. The speech signal is divided into multiple short-time frames. Each frame is Fourier transformed to obtain a spectrum. The spectrum is filtered using a Mel filter bank to obtain an energy spectrum. The energy spectrum is discrete cosine transformed (DCT) to obtain MFCC coefficients. The MFCC coefficients of each frame are combined into a feature vector sequence. The speech disorder of the subject is identified based on the speech feature vector sequence.

[0088] In this embodiment, identifying the speech disorder of the subject according to the speech feature vector sequence further includes:

[0089] Calculating the coefficient variance of the speech feature vector sequence, and if the coefficient variance deviates from a preset normal range, determining that the subject's speech disorder is unclear speech;

[0090] The zero-crossing rate of the speech feature vector sequence is calculated, and if the zero-crossing rate is higher than a preset threshold, it is determined that the subject's speech disorder is expression difficulty.

[0091] By calculating the coefficient variance and zero-crossing rate, a reliable representation of speech features can be provided even in a noisy environment, thereby more accurately capturing the feature changes in the speech signal. In addition, the calculation of the coefficient variance and zero-crossing rate is relatively simple and can be completed in a short time. Figure 2 Shown is the frequency of speech impairment, Figure 3 Shown are the frequencies of difficulties in expression.

[0092] In this embodiment, the expression of the zero crossing rate is:

[0093]

[0094] Among them, x″ t is the speech feature vector x t The second derivative at time index t, x t is the audio signal value at time point t; Ind{} is an indicator function that returns 1 when the condition is met, otherwise it returns 0.

[0095] Step S103, obtaining facial feature points including the face and the mouth from the video to be detected through the Dlib detector; and calculating the distance between the corner of the mouth and the center point of the lips and the distortion angle of the corner of the mouth according to the facial feature points.

[0096] The Dlib detector has good robustness and can accurately detect facial landmarks under different lighting conditions and angles. For example, the shape_predictor_68_face_landmarks.dat model detects 68 facial landmarks, which are enough to locate the corners of the mouth and the center of the lips as well as the angle of the distortion of the corners of the mouth.

[0097] In this embodiment, calculating the distance between the mouth corner and the center point of the lips and the distortion angle of the mouth corner according to the facial feature points further includes:

[0098] Estimate the average position of the mouth corner points according to the positions of the left and right mouth corners to obtain the lip center point; calculate the distance between the mouth corner and the lip center point using the Euclidean distance formula;

[0099] The distortion angle of the mouth corner is obtained by the angle between the direction of the line connecting the mouth corner point and the lip center point and the horizontal direction.

[0100] In Dlib's 68-point model, the mouth corners usually correspond to the 48th point (left mouth corner) and the 54th point (right mouth corner), as shown in Figure 4 As shown, the distance from each mouth corner to the lip center is calculated using the Euclidean distance formula. The distortion angle is obtained by calculating the angle between the line connecting the mouth corner point and the lip center point and the horizontal direction.

[0101] Wherein, the distance between the corner of the mouth and the center point of the lips is:

[0102]

[0103] Where n is the number of feature points; P i is the coordinate of the ith feature point; C is the coordinate of the center point of the lips; (P i , C) is the distance from the i-th feature point to the lip center point C; w i is the offset coefficient of the i-th feature point, |y i-y C | is the feature point P i The vertical coordinate y i The ordinate y of the center point C C The absolute offset difference.

[0104] The expression of the distortion angle of the mouth corner is:

[0105] Δθ=|θ R -θ L |

[0106] in, It is the angle between the line connecting the right corner of the mouth and the center point and the horizontal direction; It is the angle between the line connecting the left corner of the mouth and the center point and the horizontal direction.

[0107] Step S104, judging whether the subject has stroke symptoms according to the smoothness of the two-hand trajectory movement, the voice characteristics, the center point distance and the distortion angle of the mouth corners.

[0108] In this embodiment, different weights are set for the smoothness of the two-hand trajectory movement, voice features, center point distance, and the distortion angle of the mouth corners, and feature standardization (such as z-score standardization) is performed. When the comprehensive score exceeds the preset threshold, it is judged that there are stroke symptoms. Optionally, data analysis methods such as recursive feature elimination (RFE) and feature importance ranking are used to find out which features are more sensitive to the identification of stroke symptoms, thereby adjusting the weight distribution. In addition to the existing features, features such as changes in facial expressions and eye movement trajectories can also be introduced to further identify stroke symptoms.

[0109] The solution provided by the above embodiment of the present invention extracts the MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; the speech disorder of the subject is identified according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression; the facial feature points including the face and the mouth are obtained from the video to be detected by the Dlib detector; the distance between the center point of the mouth corner and the lips and the distortion angle of the mouth corner are calculated according to the facial feature points; and whether the subject has stroke symptoms is determined according to the smoothness of the movement of the two-hand trajectory, the speech features, the center point distance and the distortion angle of the mouth corner. The present invention can comprehensively evaluate the stroke symptoms of the subject from multiple angles by integrating multiple biometric information such as gesture analysis, speech feature extraction and facial feature detection, thereby improving the accuracy and reliability of recognition. The preliminary screening of stroke symptoms can be completed in a short time without the need for complex medical examination equipment, and conclusions can be quickly drawn through real-time processing of video and audio data, shortening the diagnosis time and winning precious time for timely treatment. It mainly relies on video and audio data, does not require direct contact with the subject, and has less physical burden on the patient than traditional imaging examinations such as CT or MRI, and is more suitable for large-scale screening. Only an ordinary mobile phone APP is needed for detection, which reduces the cost of equipment. It is not only suitable for professional environments of medical institutions, but also suitable for homes, schools and other places, and has a wider applicability. The coordinates of the key points of both hands are extracted through the posture estimation method, which can accurately capture the details of hand movement and improve the accuracy of gesture analysis. The smoothness of movement is evaluated by calculating the variance of the acceleration change, the coefficient variance and zero crossing rate of the speech feature vector sequence are calculated to identify speech disorders, and the distance between the corner of the mouth and the center of the lip and the distortion angle are calculated to evaluate facial symmetry, further improving the accuracy of stroke recognition.

[0110] Figure 6 The structural block diagram of the device for rapid identification of stroke symptoms according to an embodiment of the present invention is shown. The device for rapid identification of stroke symptoms 600 comprises:

[0111] The arm analysis module 601 is used to extract the coordinates of the key points of the hands of the subject from the video to be detected according to the posture estimation method; and calculate the smoothness of the movement of the two hands trajectory according to the coordinates of the key points of the two hands;

[0112] The speech disorder analysis module 602 is used to extract the MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; and identify the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression;

[0113] The facial feature analysis module 603 is used to obtain facial feature points including the face and the mouth from the video to be detected through the Dlib detector; and calculate the distance between the corner of the mouth and the center point of the lips and the distortion angle of the corner of the mouth according to the facial feature points;

[0114] The stroke symptom judgment module 604 is used to judge whether the subject has stroke symptoms according to the smoothness of the two-hand trajectory movement, the voice characteristics, the center point distance and the distortion angle of the mouth corners.

[0115] Figure 7 The schematic diagram of the structure of the computing device embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the computing device.

[0116] like Figure 7 As shown, the computing device may include: a processor (processor) 702 , a communications interface (Communications Interface) 704 , a memory (memory) 706 , and a communication bus 708 .

[0117] The processor 702, the communication interface 704, and the memory 706 communicate with each other via a communication bus 708. The communication interface 704 is used to communicate with other devices such as a client or other server network elements. The processor 702 is used to execute a program 710, which can specifically execute the relevant steps in the above-mentioned embodiment of the method for rapid identification of stroke symptoms.

[0118] Specifically, the program 710 may include program codes, which include computer operation instructions.

[0119] The processor 702 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the computing device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0120] The memory 706 is used to store the program 710. The memory 706 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0121] The solution provided by the above embodiment of the present invention extracts the MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; the speech disorder of the subject is identified according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression; the facial feature points including the face and the mouth are obtained from the video to be detected by the Dlib detector; the distance between the center point of the mouth corner and the lips and the distortion angle of the mouth corner are calculated according to the facial feature points; and whether the subject has stroke symptoms is determined according to the smoothness of the movement of the two-hand trajectory, the speech features, the center point distance and the distortion angle of the mouth corner. The present invention can comprehensively evaluate the stroke symptoms of the subject from multiple angles by integrating multiple biometric information such as gesture analysis, speech feature extraction and facial feature detection, thereby improving the accuracy and reliability of recognition. The preliminary screening of stroke symptoms can be completed in a short time without the need for complex medical examination equipment, and conclusions can be quickly drawn through real-time processing of video and audio data, shortening the diagnosis time and winning precious time for timely treatment. It mainly relies on video and audio data, does not require direct contact with the subject, and has less physical burden on the patient than traditional imaging examinations such as CT or MRI, and is more suitable for large-scale screening. Only an ordinary mobile phone APP is needed for detection, which reduces the cost of equipment. It is not only suitable for professional environments of medical institutions, but also suitable for homes, schools and other places, and has a wider applicability. The coordinates of the key points of both hands are extracted through the posture estimation method, which can accurately capture the details of hand movement and improve the accuracy of gesture analysis. The smoothness of movement is evaluated by calculating the variance of the acceleration change, the coefficient variance and zero crossing rate of the speech feature vector sequence are calculated to identify speech disorders, and the distance between the corner of the mouth and the center of the lip and the distortion angle are calculated to evaluate facial symmetry, further improving the accuracy of stroke recognition.

[0122] The algorithm or display provided herein is not inherently related to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious to construct the structure required for this type of system. In addition, the embodiment of the present invention is not directed to any specific programming language yet. It should be understood that various programming languages ​​can be utilized to realize the content of the present invention described herein, and the description made to specific languages ​​above is for disclosing the best mode of the present invention.

[0123] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0124] Similarly, it should be understood that in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, the various features of the embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following intention: that the claimed invention requires more features than the features explicitly recited in each claim. More specifically, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, with each claim itself serving as a separate embodiment of the present invention.

[0125] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0126] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.

[0127] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all functions of some or all components according to embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., computer program and computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0128] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be understood as limitations on the order of execution.

Claims

1. A method for rapid identification of stroke symptoms, characterized in that: include: Extracting the coordinates of the key points of the subject's hands from the video to be detected according to the posture estimation method; The smoothness of the movement of the two-hand trajectory is calculated according to the coordinates of the key points of the two-hand; Extracting MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; identifying the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression; Obtaining facial feature points including the face and the mouth from the video to be detected by using a Dlib detector; calculating the distance between the corner of the mouth and the center point of the lips and the distortion angle of the corner of the mouth according to the facial feature points; Determining whether the subject has stroke symptoms according to the smoothness of the two-hand trajectory movement, the speech characteristics, the center point distance, and the distortion angle of the mouth corners; The calculation formula for the smoothness of the two-hand trajectory movement is: Among them, k is the scale of fluency; ∈ is used to control the maximum value of fluency; Var L , Var R are the variances of the acceleration changes of the left and right hands, respectively, reflecting the degree of fluctuation of the acceleration change; Δa L (t)=a L (t+Δt)-a L (t) Δa R (t)=a R (t+Δt)-a R (t) Where n is the number of points in the acceleration sequence used to calculate the difference; μ L and μ R are the average values ​​of the acceleration changes of the left and right hands respectively; a L (t) is the acceleration of the left hand at time t; a R (t) is the acceleration of the right hand at time t; Δt is the time difference between two adjacent frames; The approximate expression of the acceleration of the left hand at time t is: Among them, x L (t) is the position coordinate of the left hand at time t; The approximate expression for the acceleration of the right hand at time t is: Among them, x R (t) is the position coordinate of the right hand at time t.

2. The method for rapid identification of stroke symptoms according to claim 1, characterized in that: Identifying the subject's speech disorder according to the speech feature vector sequence further includes: Calculating the coefficient variance of the speech feature vector sequence, and if the coefficient variance deviates from a preset normal range, determining that the subject's speech disorder is unclear speech; The zero-crossing rate of the speech feature vector sequence is calculated, and if the zero-crossing rate is higher than a preset threshold, it is determined that the subject's speech disorder is expression difficulty.

3. The method for rapid identification of stroke symptoms according to claim 2, characterized in that: The expression of the zero crossing rate is: Among them, x” t is the speech feature vector x t The second derivative at time index t, x t is the audio signal value at time point t; Ind{} is an indicator function that returns 1 when the condition is met, otherwise it returns 0.

4. The method for rapid identification of stroke symptoms according to claim 1, characterized in that: Calculating the distance between the center point of the mouth corner and the lips and the distortion angle of the mouth corner according to the facial feature points further includes: Estimate the average position of the mouth corner points according to the positions of the left and right mouth corners to obtain the lip center point; calculate the distance between the mouth corner and the lip center point using the Euclidean distance formula; The distortion angle of the mouth corner is obtained by the angle between the direction of the line connecting the mouth corner point and the lip center point and the horizontal direction.

5. The method for rapid identification of stroke symptoms according to claim 4, characterized in that: The distance between the corner of the mouth and the center of the lips is: Where n is the number of feature points; P i is the coordinate of the i-th feature point; C is the coordinate of the center point of the lips; d(P i , C) is the distance from the i-th feature point to the lip center point C; w i is the offset coefficient of the i-th feature point, |y i -y C | is the feature point P i The vertical coordinate y i The ordinate y of the center point C C The absolute offset difference.

6. The method for rapid identification of stroke symptoms according to claim 4, characterized in that: The expression of the distortion angle of the mouth corner is: Δθ=|θ R -θ L | in, It is the angle between the line connecting the right corner of the mouth and the center point and the horizontal direction; It is the angle between the line connecting the left corner of the mouth and the center point and the horizontal direction.

7. A device for rapid identification of stroke symptoms, comprising: An arm analysis module is used to extract the coordinates of the key points of the subject's hands from the video to be detected according to the posture estimation method; The smoothness of the movement of the two-hand trajectory is calculated according to the coordinates of the key points of the two-hand; A speech disorder analysis module, used to extract the MFCC features of each frame in the video to be detected according to the MFCC algorithm to obtain a speech feature vector sequence; and identify the speech disorder of the subject according to the speech feature vector sequence, wherein the speech disorder includes unclear speech and difficulty in expression; A facial feature analysis module is used to obtain facial feature points including the face and the mouth from the video to be detected through a Dlib detector; and calculate the distance between the corner of the mouth and the center point of the lips and the distortion angle of the corner of the mouth according to the facial feature points; A stroke symptom judgment module, used to judge whether the subject has stroke symptoms according to the smoothness of the two-hand trajectory movement, the voice characteristics, the center point distance and the distortion angle of the mouth corners; The calculation formula for the smoothness of the two-hand trajectory movement is: Among them, k is the scale of fluency; ∈ is used to control the maximum value of fluency; Var L , Var R are the variances of the acceleration changes of the left and right hands, respectively, reflecting the degree of fluctuation of the acceleration change; Δa L (t)=a L (t+Δt)-a L (t) Δa R (t)=a R (t+Δt)-a R (t) Where n is the number of points in the acceleration sequence used to calculate the difference; μ L and μ R are the average values ​​of the acceleration changes of the left and right hands respectively; a L (t) is the acceleration of the left hand at time t; a R (t) is the acceleration of the right hand at time t; Δt is the time difference between two adjacent frames; The approximate expression of the acceleration of the left hand at time t is: Among them, x L (t) is the position coordinate of the left hand at time t; The approximate expression for the acceleration of the right hand at time t is: Among them, x R (t) is the position coordinate of the right hand at time t.

8. A computing device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the method for rapid identification of stroke symptoms according to any one of claims 1-6.

Citation Information

Patent Citations

  • Classification method of hepatolenticular degeneration (HLD) speech disorders based on compressed sensing

    CN110211566A

  • Intelligent cerebral apoplexy diagnosis system

    CN111312389A