Method, server and computer-readable medium for detecting cognitive and language impairments

Through computer vision analysis of facial feature points and mouth movements, and measuring pause frequency and repetition patterns, the problem of inaccurate privacy leakage and evaluation in the prior art is solved, and the accurate assessment of aphasia is achieved under the premise of protecting privacy.

CN111489819BActive Publication Date: 2025-08-15FUJIFILM BUSINESS INNOVATION CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911347183.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-29
Filing Date
2019-12-24
Publication Date
2025-08-15
Estimated Expiration
2039-12-24

AI Technical Summary

Technical Problem

The prior art has problems with the risk of privacy leakage and inaccurate assessment when assessing language and cognitive impairments, especially when detecting aphasia, and the need to protect subjects’ privacy and provide an accurate assessment of competence.

Method used

Computer vision methods are used to analyze facial feature points and mouth movements, measure pause frequency, repetition patterns and vocabulary changes, generate total scores, and use decision trees or support vector machines to predict to avoid recognition of language content.

Benefits of technology

It realizes accurate assessment of language and cognitive impairments, especially aphasia, while protecting privacy, and provides continuous and representative assessments, suitable for automated assessments outside the clinical setting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111489819B_ABST
    Figure CN111489819B_ABST
Patent Text Reader

Abstract

A method, server, and computer-readable medium for detecting cognitive and language disorders. A computer-implemented method for assessing the presence of a condition in a subject is provided, the method comprising: generating facial landmarks on a received input video to define points associated with regions of interest on the subject's face; defining speech periods based on respective positions of the defined landmarks; measuring pause frequency, repetition patterns, and lexical variation during the speech periods without determining or applying semantic information associated with the speech to generate an overall score; and generating a prediction of the presence of the condition in the subject, associated with the overall score, based on the overall score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of example implementations relate to methods, systems, and user experiences associated with detecting language and cognitive impairments based on video sequences when using only visual facial features without detecting the user's language content. Background Art

[0002] The individuality (for example, experimenter) with neurological condition may experience language disorder as the symptom of those neurological conditions.For example, aphasia is the neurological condition that affects the ability of individuality to understand or produce language.Aphasia may be caused by apoplexy or other brain damage, and may improve or worsen over time.The degree and type of the language disorder associated with aphasia may span a wide range.For example, there may be a slightly slurred speech (for example, pauses, repeated words or limited vocabulary) in the individuality with aphasia, or the severe limitation of only being able to say a few words or sounds.

[0003] In the prior art, an assessment of the abilities of an individual with aphasia can be performed. According to one prior art technique, a doctor or therapist can perform a manual determination of aphasia. Prior art assessments can range from a rough classification of the subject's abilities to a detailed analysis of symptoms based on clinical interview records. For example, where medical personnel use clinical interview records to analyze symptoms associated with aphasia, a significant amount of time is required, and the results may not be representative of the subject's abilities outside of a clinical setting, as the subject may not react in the same way in a clinical setting as in a non-clinical setting.

[0004] Existing technology assessments may be inaccurate. Furthermore, interview transcripts may reveal language content (e.g., private or sensitive information), which could expose the user's private information. Consequently, the user may be forced to either risk leaking their private information to clinicians or others, or be denied treatment or management of their aphasia.

[0005] According to one prior art approach, audio information can be used to infer language proficiency. However, the audio approach requires the system to detect the subject's conversational content. This approach can lead to prior art issues regarding the disclosure of confidential information and the subject's privacy.

[0006] Therefore, there is a pressing need for capabilities that allow for the assessment of subjects over time in a privacy-preserving and confidentiality-protecting manner that also avoids prior art errors associated with analyses performed in clinical settings. Summary of the Invention

[0007] According to aspects of an example implementation, a computer-implemented method for assessing the presence of a condition in a subject is provided, comprising the steps of generating facial feature points on a received input video to define points associated with regions of interest on the subject's face; defining speech periods based on the respective positions of the defined points; measuring pause frequency, repetition patterns, and vocabulary variations during the speech periods without determining or applying semantic information associated with the speech to generate an overall score; and generating a prediction of the presence of the condition in the subject associated with the overall score based on the overall score.

[0008] According to another aspect, the step of generating facial landmarks includes defining, for a region of interest including a mouth of the subject, defined points outlining the lips of the subject, and measuring a temporal similarity metric of the defined points over time, and removing body movement and head movement from the temporal similarity metric to generate a temporal dissimilarity metric of the mouth.

[0009] According to another aspect, the step of defining a speech period includes filtering out mouth movements associated with out-of-plane rotations and jitters of the mouth, measuring a vertical distance between an upper lip and a lower lip, and performing a closing operation to generate a speech score indicating a speech period.

[0010] According to another aspect, the step of measuring the pause frequency includes applying a threshold to the speech score and registering periods of speech inactivity during the speech period as the pause frequency.

[0011] According to another aspect, the step of measuring the repetition pattern includes defining first and second mouth movement patterns having certain lengths around first and second time intervals, respectively, during the speech period, and performing a dissimilarity comparison to obtain the number of repetitions over the time period.

[0012] According to another aspect, the step of measuring vocabulary variation includes: using clustering to collect and aggregate a fixed number of vocabulary patterns throughout the input video in a fixed number of clusters by selecting recurring patterns in the patterns, reconstructing mouth movements on the fixed number of vocabulary patterns to generate a reconstruction cost of a score indicative of mouth movement variation.

[0013] According to another aspect, lexical variation is measured without identifying the language of the lexicon.

[0014] According to another aspect, the step of generating a prediction further comprises applying a decision tree or a support vector machine to learn a separating function between subjects with the condition and subjects without the condition.

[0015] Example implementations may also include a non-transitory computer-readable medium having a storage device and a processor capable of executing instructions for assessing a subject for the presence or absence of a condition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] This patent or application file contains at least one drawing drawn in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0017] Figure 1(a) to Figure 1(d) Example implementations of an incoming feature image with facial landmarks and measurements are shown according to various example implementations.

[0018] Figure 2(a) to Figure 2(b) Shown are pauses determined according to an example implementation for a person with aphasia compared to members of a control group.

[0019] Figure 3(a) to Figure 3(b) A comparison of repetitive mouth patterns of a person with aphasia and members of a control group, determined according to an example implementation, is shown.

[0020] Figure 4 An example system diagram is shown according to an example implementation.

[0021] Figure 5 An example process according to an example implementation is shown.

[0022] Figure 6 An example computing environment is shown with an example computer device suitable for use with some example implementations.

[0023] Figure 7 An example environment suitable for some example implementations is shown. DETAILED DESCRIPTION

[0024] The following detailed description provides further details of the drawings and example implementations of the present application. For clarity, the numbers and descriptions of redundant elements between the drawings are omitted. The terms used throughout the specification are provided as examples and are not intended to be limiting.

[0025] There is a pressing need for automated assessment tools that track language abilities more frequently over time and outside of clinical settings. This approach would allow for more effective tailoring of therapy for individuals with aphasia or other cognitive and language impairments. According to one approach, assessment tools could be integrated with daily video calls to provide a more continuous and representative assessment of language impairment and detect improvement or deterioration over time.

[0026] Various aspects of example implementations relate to methods and systems for detecting speech and cognitive impairments based on video sequences. More specifically, a cross-media approach is employed that applies only visual facial features to detect speech properties without using the audio content of the speech. For example, facial landmark detection results can be used to measure facial motion over time. Thus, speech and pause instances are detected based on temporal lip shape analysis, and a dynamic time warping mechanism is used to identify repetitive mouth patterns. In one example implementation, the above-mentioned features associated with pause frequency, mouth pattern repetition, and vocabulary pattern variation are applied to detect symptoms associated with medical conditions such as aphasia.

[0027] Computer vision is used to model conditions related to language and cognition. Compared to prior art methods that track facial points to recognize, identify, or detect facial gestures (e.g., gaze or emotion), example implementations involve modeling language and cognition. Furthermore, and also distinct from prior art audio / visual speech recognition that uses visual signals to directly record the words spoken by a subject, example implementations characterize the properties of speech (e.g., speech rate, degree of motor repetition, and vocabulary variation employed) without revealing the content of the speech itself to infer the presence of conditions such as aphasia.

[0028] According to the example implementations described herein, the cross-media method uses only visual information to infer language-related properties, thereby protecting the subject's privacy by not detecting or determining the semantic content of their language. In addition, the privacy-preserving methods of the example implementations can increase the acceptance of automated assessment systems and provide further possible implementations for diagnosis and treatment. In addition, the example implementations can provide for continuous assessment of subject abilities during specific conversations or video conferences with medical personnel, self-assessment by subjects over time, or aggregation with the subject's consent for use in large-scale medical research.

[0029] According to an example implementation, for a subject with the ability to be assessed for aphasia, an initial registration of facial landmarks is performed while the subject is speaking. Based on the initial registration of facial landmarks associated with the subject's face, temporal features for speech and pause detection are generated, as well as repeated facial patterns detected and total facial pattern changes measured directly based on the facial landmarks. In addition, the temporal features are associated with the actual language-related symptoms of the aphasic subject, including but not limited to disfluency, repeated speech, and limited vocabulary.

[0030] In more detail, an example implementation involves analyzing a subject's mouth (more specifically, the shape and movement of the mouth) to generate features related to the actual nature of speech. Point detection is employed, which outlines the subject's mouth in a received video and compares its shape over time. These point detections can reveal speaking turns and brief pauses within speaking turns. Furthermore, the temporal sequence of mouth shapes is grouped into mouth patterns, and different patterns are compared to each other to identify recurring patterns that occur during the subject's speech. Additionally, changes in the subject's mouth movements are measured by comparing the patterns to a small vocabulary of observed patterns.

[0031] According to an example implementation, an incoming video sequence of a subject (e.g., a subject speaking) is used as starting data. 2D feature points are registered for this video sequence. The example implementation employs a method using a convolutional neural network to obtain a 70-point model of characteristic facial points. While 70 points are used in the examples shown herein, the example implementation is not limited thereto, and those skilled in the art will appreciate that any number of points associated with facial features may be used.

[0032] For example, FIG1( a) shows a stage 101 of a subject with facial points 103, showing a frame of an incoming video sequence. Those skilled in the art will appreciate that the facial points are obtained using well-known off-the-shelf software executed on a processor such as a GPU or CPU. A significant proportion of the facial points do not change during the subject's speech in a manner meaningful for analysis of the aphasic subject's abilities. Therefore, the example implementation focuses on a subset of facial points associated with changes that occur during the subject's speech. These points are shown at 105 in FIG1( a) as points associated with inner and outer lip features of the lips. For example, in the case of 70 facial points, facial points that outline the lips (e.g., approximately 20 facial points) can be used for analysis. While the above numbers and proportions are provided for example purposes, the inventive concept is not limited to these parameters, and those skilled in the art will appreciate that other parameters may be used.

[0033] Therefore, the set of 2D facial points that change meaningfully when the subject speaks at a specific time point t in the video is represented as follows:

[0034]

[0035] Furthermore, it is noted that video material may be normalized to a standard rate, such as 30 frames per second. Thus, any point in time t may be specified by a frame index.

[0036] The analysis of mouth configurations and their movement over time is based on a temporal similarity measure of 2D facial points (e.g., mouths) associated with the subject and direct measurement of prescribed facial features (e.g., mouth opening). To compare how much the mouth shape has changed over time, the difference between two mouth configurations is measured based on their point squared difference, for example as shown in Equation (2):

[0037]

[0038] Changes in mouth configurations over time may be caused by body movement, head movement, or intra-facial motion. To account for changes based on intra-facial motion, arbitrary scaling, 2D rotation, and translation are performed to map one of the compared mouth configurations onto the other as closely as possible, and the remaining differences are used to represent the actual shape differences. In addition, the example implementation directly maps the multiple mouth configurations to be compared without the need for an intermediate mapping onto a frontal template view of the mouth. Since the information of interest is the difference between temporally adjacent mouth configurations in the video, the example implementation avoids additional errors that may be introduced by prior art methods using frontalization. For example, FIG1(b) shows temporal mouth dissimilarity for different time windows. Thus, the example implementation provides for performing shape analysis by providing a mouth dissimilarity metric as shown in Equation (3):

[0039]

[0040] Since msim depends on the scale of the first operand, an asymmetric and scale-invariant function is provided in equation (4) as follows:

[0041]

[0042] Therefore, msim norm is used to measure intra-facial motion by comparing the mouth configurations of the subjects in a time window Δt. To allow for the determination of which value of Δt will provide the optimal metric, a set of different time windows W is employed. Thus, the final temporal self-dissimilarity metric for the mouth is provided in Equation (5) as follows:

[0043]

[0044] The above operations are executed as instructions stored in a storage device and executed as operations on a processor such as a GPU or a CPU. Once the 2D feature points on the subject's face are registered using a model such as that described in the example implementation above, an analysis of the subject's condition regarding language ability (e.g., speech detection, pause frequency, and repetitive pattern analysis) can be performed, for example, for aphasia.

[0045] Example implementations provide for detecting when a subject speaks in a temporal window of a received input video to infer different properties of the subject's language ability. Speech occurs when the mouth is open and moving, and causes the mouth to move, such that periods of speech are revealed as regions of high activity (e.g., dissimilarity) in, for example, Equation (5), which captures mouth movement during speech, as well as jitter in point detections or miss-detections and changes due to out-of-plane rotations of the mouth (e.g., when the subject nods or shakes their head). To account for jitter, all temporal instances of facial landmarks need to be detected with sufficient confidence.

[0046] To account for changes due to out-of-plane rotations of the mouth (e.g., nodding or shaking the head), registration errors that occur in this situation are filtered based on the fact that speech can only occur when the mouth is open at some point during the window. For example, the vertical distance between points on the upper and lower lips can be measured, for example using o(t) as the vertical opening of the inner mouth, which is normalized to the scale of the full face. Note that o(t) will change frequently during speech, so a closing operation can be applied to eliminate gaps up to Δt=50. Using a combined minimum and maximum filter, the vertical distance is obtained as follows in Equation (6):

[0047]

[0048] As shown in Figure 1(c) above, on the sample video, o(t) and o(t) are shown across the time window. max (t). The gaps in the mouth that are temporarily closed during speech are filled, while clear boundaries are maintained for periods of silence. Therefore, it can be defined as a premise. max A threshold is set for (t) to establish that speech is detected. In addition, movements without speech or mouth opening are filtered out, such as when a person only nods. The final speech score (e.g., mouth movement multiplied by mouth opening) is provided as shown in the following equation (7):

[0049] talk(t)=d w (t)·o max (t) (7)

[0050] Furthermore, the mouth self-dissimilarity measure as described above for Equation (5) and the vertical distance measure defined above for Equation (6) are normalized to the maximum value [0, 1]. In addition, a speech threshold τ may be used. talk A hard {0,1} assignment is performed, and very short speech intervals can be removed. In addition, closely adjacent speech intervals separated only by very short gaps can be combined to generate speech instance detections, as shown in Figure 1(d). The above operations are executed as instructions stored in a storage device and executed as operations on a processor such as a GPU or CPU.

[0051] Once it can be determined whether the subject is speaking, the frequency of pauses can be determined. For example, for subjects with aphasia, dysfluent speech, which manifests as unexpected pauses during speech, is a defining symptom. Thus, an example implementation involves developing a metric for pauses as a direct measure of fluency. In an example implementation, pauses are detected by using the speech score as described above for equation (7), and then applying a more restrictive threshold τ pause , and also registers all inactive regions during previously detected speech instances. The above operations are performed as instructions stored in a storage device and executed as operations on a processor such as a GPU or a CPU.

[0052] As mentioned above, Figure 1(d) shows an example of these pauses during a speech period.Despite providing pauses that the subject intentionally provides as part of the normal course of speaking, the overall pause frequency can still be correlated with the subject's language fluency.

[0053] As shown in FIG2(a) and FIG2(b) respectively, a qualitative example can be provided which shows the difference in the pause frequency of people with aphasia compared to members of the control group.

[0054] In addition to dysfluent speech, aphasic subjects also have the characteristic of frequently repeating sounds, words, or sentence fragments. These repetitions may occur when forming the next word or attempting to correct the previous word. An example implementation involves detecting repetitions of mouth movements associated with speech repetition. Although repeated mouth movements may not indicate repetition of the semantic meaning of the language content, information associated with the visual representation can determine repetitive behavior.

[0055] In order to detect visual repetition of mouth movements, a mouth movement pattern of length l around time point t is defined in Equation (8):

[0056] p t =(m t-l / 2 ,...,m t ,...,m t+l / 2 ) (8)

[0057] Thus, the example implementation uses dynamic time warping (e.g., frame-by-frame analysis of speech segments) to compare two observed patterns of arbitrary length. A first pattern is transformed into a second pattern by transforming a first mouth configuration into a second mouth configuration, or by allowing insertions or deletions in the first configuration. As described above, the cost of a direct transformation operation is msim between two transformed mouth configurations. norm, with its maximum value normalized to [0, 1]. Greater dissimilarity is associated with a higher cost. Furthermore, insertion and deletion operations are assigned a maximum cost of 1. Because the same mouth movement may not always be performed at the same speed, temporal warping using insertions and deletions provides for matching similar patterns of varying lengths. Consequently, the total pattern matching cost is the sum of the optimal sequence costs of the transformation operations.

[0058] To find possible repetitions, reference patterns around the local unique mouth configuration are extracted. For example, d w Then, for example but not by way of limitation, search for matches with a cost less than a threshold τ in the vicinity (e.g., plus or minus 5 seconds) of match Matching pattern.

[0059] Figures 3(a) and 3(b) respectively illustrate examples of pattern matching for subjects with aphasia who frequently repeat a single word, compared to individuals without aphasia. While repetitive patterns are present in both cases, subjects with aphasia exhibit fewer, more direct repetitions, compared to the many, highly interwoven repetitions of individuals without aphasia. Therefore, direct repetitions that are absent or separated by few other patterns can indicate direct word repetitions, allowing occurrences of direct repetitions to be counted and normalized by total speech time to obtain a measure of visual repetitions per second, an indicator of aphasia. The operations described above are executed as instructions stored in a storage device and executed as operations on a processor, such as a GPU or CPU.

[0060] In addition to detecting patterns that appear or repeat within a time window, patterns can be collected over the entire video to assess language variation, which is directly related to changes in mouth movements or expressions and, therefore, actual language variation. For example, a person who can only express a few words or sounds will show less mouth movement variation than an able-bodied person with a normal vocabulary.

[0061] According to an example implementation, the method considers a longer time window and evaluates whether clusters exist across an entire speech event (e.g., a phone conversation or video conference), and can perform a diagnosis of the vocabulary used by a subject without knowing the identity of the actual words in the language. This is accomplished by obtaining a measure of vocabulary variation, as described herein.

[0062] More specifically, a measure of vocabulary variation is obtained by constructing a visual vocabulary of mouth patterns for each person, collecting a fixed number of patterns throughout the video, and aggregating the patterns using clustering in a fixed number of clusters. Patterns that repeat at least once are selected. However, those skilled in the art will appreciate that a different threshold can be selected, such as two repetitions, or some other value for pattern repetition.

[0063] The representatives of all cluster centers form a vocabulary. According to the example implementation, k-medioids clustering is used, which has a predefined vocabulary size k. However, other clustering heuristics may be substituted for it without departing from the scope of the invention of the example implementation. For smaller vocabulary sizes k, measurements are performed to determine to what extent the complete variation of the subject's mouth movements is represented. The speech in the video is divided into fixed-size blocks of mouth movements. Each block has the same length as the pattern in the vocabulary. Each block is assigned the vocabulary element with the lowest pattern matching cost, as explained above for the repeated pattern determination. Therefore, the complete mouth movement is reconstructed during the speech only by concatenating the most suitable vocabulary elements. Based on this reconstruction, the matching cost of the reconstruction cost between each block and its assigned pattern is measured. The total reconstruction cost is then calculated as a block average. The above operations are performed as instructions stored in a storage device and as operations on a processor such as a GPU or CPU.

[0064] In cases where the variation in mouth movements is limited, a small vocabulary is good enough to describe the total movement and leads to good reconstructions. A reconstruction cost is calculated for each of multiple values of k, and the mean of the reconstruction costs is the final score for mouth movement variation, which is an indicator of the subject's aphasia.

[0065] In addition to not detecting the actual semantics of the words in the conversation itself, and thus protecting the subject's privacy with respect to the content of the conversation, the example implementation may also have additional benefits and advantages. As an example of such a benefit, there is no need for the tool or clinician to understand the subject's grammar in order to apply the tool. For example, a clinician does not need to know the subject's language in order to treat or study aphasia in one or more subjects. Similarly, for more extensive studies that aggregate the subject's information, a larger group of people can participate in these studies because there is no requirement for language-specific tools. One such additional feature will be that the example implementation is language-independent. Therefore, the example implementation can be used regardless of the language spoken by the subject.

[0066] As mentioned above, information associated with pauses, repetitions and vocabulary is obtained for the subject. Application of this information associated with pauses, repetitions and vocabulary provides a comprehensive prediction for whether aphasia exists and the degree or type of aphasia. Although example implementations relate to aphasia, other cognitive or language disorders can also be predicted. A training set of videos of a subject with a given disorder can be provided, as well as examples of control subjects who do not show the disorder, so that a combination of characteristics that predict the presence of the disorder can be learned using learning techniques (e.g., machine learning or neural networks, etc.).

[0067] In the example implementations described above, machine learning can be used to generate predictions based on example videos, such as publicly or privately stored videos showing examples of facial movements of one or more aphasia subjects. Using this past data as a basis, machine learning can be used to generate predictions of aphasia. General machine learning tools known to those skilled in the art can be used. In addition, and as also described herein, decision trees and support vector machines can also be executed as instructions on a non-transitory computer-readable medium with a processor (e.g., a CPU or GPU). In addition, by using a processor interconnected with the subject's sensors (e.g., via a network), the example implementations can be executed independently of language, country, location, etc.

[0068] For example, but not by way of limitation, on a few minutes of video, features can be extracted along with statistics of their distribution and general classifiers, including but not limited to decision trees or support vector machines, which can be applied to learn a separation function between two classes (i.e., subjects with a given disorder and control subjects).

[0069] Figure 4 An example system diagram 400 associated with an example implementation is shown. According to the example system diagram 400, an input video is provided at 401. At 403, pose estimation is performed to generate local mouth similarities at 405 and global mouth similarities at 407. At 409, the local mouth similarities generated at 405 are input to detect pause frequency. At 411, the results of the pause frequency detection at 409 and the global mouth similarities provided at 407 are provided as input to determine the presence of repetitive patterns. Additionally, at 413, the results of the pause frequency detection at 409 and the global mouth similarities generated at 407 are provided as input to provide visual clustering determination of lexical patterns to assess mouth movement changes as described above. At 415, the outputs of pause determination, repetition detection, and visual lexical pattern recognition are input to a modeling and prediction tool, which performs aggregation and prediction for the presence of aphasia as described above.

[0070] Figure 5 An example process 500 is shown according to an example implementation. The example process 500 may be performed on one or more devices as described above.

[0071] At 501, an input video is received. For example, but not by way of limitation, the input video may be generated by a subject outside of a clinical setting (e.g., during a video conference). The subject may use a mobile communication device (e.g., a smartphone, laptop, tablet, or other device, whether integrated or detachable), that includes an input device such as a camera. Example implementations are not limited to video cameras, and any other sensing device capable of performing the function of sensing a facial area of a subject may be substituted. For example, but not by way of limitation, according to example implementations, a 3D camera may be used to sense facial features and generate signals for processing.

[0072] For the purpose of example implementation, it is not required that the input device include a microphone because audio or other language semantic output is not analyzed. Optionally, the subject also can receive input from another user with whom he or she communicates, such as a video picture and / or audio output (e.g., a speaker or headphones). For the purpose of example implementation, the subject's activity is captured by the input device, and the input is sent to one or more processors (e.g., a server) that receive the input video.

[0073] At 503, a lateral map of the subject's face is generated using the input video. As described above, characteristic facial points are determined, and a subset of characteristic facial points associated with the mouth is identified as the focused facial points.

[0074] At 505, it is determined whether the facial motion in the input video is speech. As described above, non-speech movements (e.g., shaking, nodding, shaking) of a subset of facial points are filtered out, and a closing operation is performed to obtain a speech score indicating speech periods.

[0075] At 507, for the speech period, the frequency of pauses is measured. For example, but not by way of limitation, inactive regions within the speech instance obtained at 505 may be measured as pauses, and the frequency of the measured pauses determined.

[0076] At 509, repetitive patterns are determined. More specifically, based on mouth movement patterns, the similarity of multiple observed patterns is measured. The occurrence of direct repetitions is measured and normalized by the total speech time to obtain a measure of visual repetitions per second, for example.

[0077] The overall variation of mouth movements is measured to determine vocabulary variation without identifying the content of the vocabulary at 511. Mouth patterns and clustering are implemented to calculate a final score for mouth movement variation.

[0078] At 513 , the measures and scores obtained from 507 for pause frequency, 509 for repetition pattern, and 511 for lexical variation are combined (eg, aggregated).

[0079] At 515, the aggregated measurements and scores are applied to generate a model that indicates whether the subject has a medical condition (eg, aphasia).For example, a training set of videos can be used as described above to learn characteristic combinations that are characteristic of a medical condition.

[0080] Figure 6 An example computing environment 600 is shown with an example computing device 605 suitable for use with some example implementations. The computing device 605 in the computing environment 600 may include one or more processing units, cores, or processors 610, memory 615 (e.g., RAM, ROM, etc.), internal storage 620 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interfaces 625, any of which may be coupled on a communication mechanism or bus 630 for communicating information or embedded in the computing device 605.

[0081] The computing device 605 may be communicatively coupled to an input / interface 635 and an output device / interface 640. Either or both of the input / interface 635 and the output device / interface 640 may be wired or wireless interfaces and may be detachable. The input / interface 635 may include any device, component, sensor, or interface (physical or virtual) that can be used to provide input (e.g., buttons, a touch screen interface, a keyboard, a pointing / cursor control, a microphone, a camera, Braille, a motion sensor, an optical reader, etc.).

[0082] Output devices / interfaces 640 may include displays, televisions, monitors, printers, speakers, Braille, etc. In some example implementations, input / interface 635 (e.g., a user interface) and output devices / interface 640 may be embedded in or physically coupled to computing device 605. In other example implementations, other computing devices may serve as or provide functionality for input / interface 635 and output devices / interface 640 of computing device 605.

[0083] Examples of computing devices 605 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices in vehicles and other machines, devices carried by people and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, server devices, other computers, kiosks, televisions embedded with and / or coupled to one or more processors, radios, etc.).

[0084] Computing device 605 can be communicatively coupled to external storage 645 and network 650 (e.g., via I / O interface 625) for communicating with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. Computing device 605 or any connected computing device can function as, provide services for, or be referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or another label. For example, and not by way of limitation, network 650 can include a blockchain network and / or a cloud.

[0085] I / O interface 625 may include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11xs, universal system bus, WiMAX, modems, cellular network protocols, etc.) for communicating information to and / or from at least all connected components, devices, and networks in computing environment 600. Network 650 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).

[0086] The computing device 605 may use and / or communicate with computer-usable or computer-readable media (including transitory and non-transitory media). Transitory media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-transitory media include magnetic media (e.g., magnetic disks and tapes), optical media (e.g., CD ROMs, digital video disks, Blu-ray discs), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage devices), and other non-volatile storage devices or memories.

[0087] The computing device 605 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some example computing environments. The computer-executable instructions can be retrieved from a transient medium, as well as stored on and retrieved from a non-transitory medium. The executable instructions can originate from one or more of any programming, scripting, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0088] The processor 610 can execute under any operating system (OS) (not shown) in a local or virtual environment. One or more applications can be deployed, including a logic unit 655, an application programming interface (API) unit 660, an input unit 665, an output unit 670, a facial feature point unit 675, a temporal feature analysis unit 680, a modeling / prediction unit 685, and an inter-unit communication mechanism 695 (not shown) for different units to communicate with each other, with the OS, and with other applications.

[0089] For example, the facial feature point unit 675, the temporal feature analysis unit 680, and the modeling / prediction unit 685 may implement one or more of the processes described above for the above structures. The design, function, configuration, or implementation of the described units and elements may vary and are not limited to the description provided.

[0090] In some example implementations, when information is received or instructions are executed via the API unit 660, they may be communicated to one or more other units (e.g., the logic unit 655, the input unit 665, the facial landmark unit 675, the temporal feature analysis unit 680, and the modeling / prediction unit 685).

[0091] For example, the facial landmark unit 675 can receive and process the input video and register the 2D landmarks on the subject's facial image to generate a model of characteristic facial points, as described in more detail above. The output of the facial landmark unit 675 is provided to the temporal feature analysis unit 680, which performs analysis to detect visual vocabulary of speech, pause frequency, repetitive patterns, and mouth patterns, as described in more detail above. The output of the temporal feature analysis unit 680 is provided to the modeling / prediction unit 685, which generates a model and provides inferences about the subject's condition as to whether a medical condition such as aphasia is present.

[0092] In some cases, in some of the example implementations described above, the logic unit 655 may be configured to control the flow of information between the units and direct the services provided by the API unit 660, the input unit 665, the facial landmark unit 675, the temporal feature analysis unit 680, and the modeling / prediction unit 685. For example, the flow of one or more processes or implementations may be controlled by the logic unit 655 alone or in combination with the API unit 660.

[0093] Figure 7 An example environment suitable for some example implementations is shown. Environment 700 includes devices 705-745, and each device is communicatively connected to at least one other device via, for example, a network 760 (e.g., via a wired and / or wireless connection). Some devices may be communicatively connected to one or more storage devices 730 and 745.

[0094] Examples of one or more devices 705-745 may be Figure 6 7. Devices 705-745 may include, but are not limited to, a computer 705 (e.g., a laptop computing device) with a monitor and associated webcam as described above, a mobile device 710 (e.g., a smartphone or tablet), a television 715, a device associated with a vehicle 720, a server computer 725, computing devices 735-740, and storage devices 730 and 745.

[0095] In some implementations, devices 705-720 can be considered user devices associated with users of an enterprise. Devices 725-745 can be devices associated with a service provider (e.g., used by an external host to provide services as described above and with respect to the various figures, and / or to store data such as web pages, text, portions of text, images, portions of images, audio, audio clips, videos, video clips, and / or related information thereof).

[0096] While the example implementations described above relate to aphasia, any other disorder indicated by movement of facial points may also be modeled and predicted as described above. For example, and not by way of limitation, cameras may be installed in medical facilities or in the homes of subjects at high risk for a condition that use computer vision to obtain facial points and perform analysis on a relevant subset of the facial points to address the presence of the condition. In some conditions, such as stroke, certain facial movements may occur before other detectable movements, serving as an early indicator of a potentially life-threatening event, such as a stroke. Having one or more cameras capable of detecting such facial point movements at an early stage may be used to aid in early detection and enable potential stroke subjects to receive early treatment. Such a tool may be integrated with a network, a local security system, a mobile device associated with a user that has a camera, or even an object in the room that is noticeable to the user (e.g., a television or display screen).

[0097] According to another example implementation, an online application for speech therapy can use the implementations described herein. For example, but not by way of limitation, a user can use a speech therapy application downloaded to a device such as a tablet, laptop, television, smartphone, or other device where a video camera is present so that the user can obtain feedback on their progress. In addition, a user who performs therapy activities based on instructions (from a clinician or the online application itself) can use tools to track progress and receive further feedback regarding additional therapy or relevant medical information.

[0098] In addition, example implementations can be integrated into a communication application (e.g., a telephone conference or video conferencing system) between a subject and one or more other users. In such a system, an online application can use example implementations as a third-party tool or an integrated tool or plug-in to allow the subject to self-assess aphasia while communicating with other people. This may also be beneficial to the subject, because if the subject has the condition of the aphasia presented, they can adjust their communication schedule, track progress, etc. On the other hand, with the subject's consent, the user outside the subject communicating with the subject can be provided with information that the subject has aphasia, to avoid potential frustration, embarrassment, or awkward situations between the subject and other users.

[0099] It is also noted that the example implementations described above can be used in combination with prior art audio systems that also perform language content detection. However, the example implementations do not require prior art audio systems and, as described above, can be used without any audio system or other system that provides privacy-preserving information to another party.

[0100] At a broader research level, analyses associated with aphasia can be used to identify patterns across a large number of subjects. Identifying patterns in when and how symptoms appear (e.g., due to a certain trigger or after a specific time) can be used to determine and assess the extent or type of aphasia. With the subject's consent, this information can be used in large studies.

[0101] The example implementations described above may have various benefits and advantages over the prior art. For example, and not by way of limitation, according to the example implementations, the subject's privacy is protected while allowing for the determination and prediction of whether the subject has aphasia, as well as the severity of the aphasia. Furthermore, as also described above, the example implementations are language-independent and can be employed across a variety of different languages without modification. Furthermore, the example implementations allow for the identification of language properties without revealing the content of the language itself, and focus on temporal features of facial movements to ultimately derive properties of language patterns without detecting or determining actual language content.

[0102] Although some example implementations have been shown and described, these example implementations are provided in order to convey the subject matter described herein to those familiar with the art. It should be understood that the subject matter described herein can be implemented in various forms and is not limited to the example implementations described. The subject matter described herein can be practiced without those specifically defined or described matters or with other or different elements or matters that are not described. Those familiar with the art will understand that these exemplary implementations can be changed without departing from the subject matter defined in the appended claims and their equivalents as described herein.

Claims

1. A computer-implemented method for assessing whether a subject has aphasia, the method comprising the following steps: generating facial landmarks on the received input video to define points associated with regions of interest on the face of the subject; defining speech periods based on respective positions of the points associated with the region of interest; measuring pause frequency, repetition patterns, and lexical variation during the speech period without determining or applying semantic information associated with the speech to generate a total score; as well as Based on the total score, a prediction is generated that the subject has the aphasia condition that is associated with the total score.

2. The method according to claim 1, wherein The step of generating the facial landmarks includes defining, for the region of interest including the subject's mouth, defined points outlining the subject's lips, measuring a temporal similarity metric of the defined points over time, and removing body movement and head movement from the temporal similarity metric to generate a temporal dissimilarity metric for the mouth.

3. The method according to claim 2, wherein: The step of defining the speech period includes filtering out mouth movements associated with out-of-plane rotation and jitter of the mouth, measuring a vertical distance between an upper lip and a lower lip of the lips, and performing a closing operation to generate a speech score indicative of the speech period.

4. The method according to claim 3, wherein: The step of measuring the pause frequency includes applying a threshold to the speech score and registering speech inactivity periods during the speech period as the pause frequency.

5. The method according to claim 1, wherein The step of measuring the repetition pattern includes defining a first mouth movement pattern and a second mouth movement pattern having certain lengths around a first time interval and a second time interval, respectively, during the speech period, and performing a dissimilarity comparison to obtain the number of repetitions over the time period.

6. The method according to claim 1, wherein The step of measuring the vocabulary variation includes: collecting and aggregating a fixed number of vocabulary patterns throughout the input video using clustering by selecting recurring patterns in the patterns in a fixed number of clusters, reconstructing mouth movements on the fixed number of vocabulary patterns to generate a reconstruction cost of a score indicative of mouth movement variation.

7. The method according to claim 6, wherein: The vocabulary variation is measured without identifying the language of the vocabulary.

8. The method according to claim 1, wherein Generating the prediction further comprises applying a decision tree or a support vector machine to learn a separating function between subjects with the aphasic condition and subjects without the aphasic condition.

9. A server capable of determining an aphasia condition of a subject, the server being configured to: receiving an input video and performing a facial landmark generation operation on the received input video to define points associated with a region of interest; defining speech periods based on respective positions of the points associated with the region of interest; measuring pause frequency, repetition patterns, and lexical variation during the speech period without determining or applying semantic information associated with the speech; as well as The measures are aggregated to generate a prediction that the subject has the aphasia condition.

10. The server according to claim 9, wherein: The operation of generating the facial landmarks includes defining, for the region of interest including the subject's mouth, defined points outlining the subject's lips, measuring a temporal similarity metric of the defined points over time, and removing body movement and head movement from the temporal similarity metric to generate a temporal dissimilarity metric for the mouth.

11. The server according to claim 10, wherein: The operations of defining the speech period include filtering out mouth movements associated with out-of-plane rotations and jitters of the mouth, measuring a vertical distance between an upper lip and a lower lip, and performing a closing operation to generate a speech score indicative of the speech period.

12. The server according to claim 11, wherein: The operation of measuring the pause frequency includes applying a threshold to the speech score and registering a speech inactivity period during the speech period as the pause frequency.

13. The server according to claim 9, wherein: The operation of measuring the repetition pattern includes defining first and second mouth movement patterns having certain lengths around first and second time intervals, respectively, during the speech period, and performing a dissimilarity comparison to obtain the number of repetitions over the time period.

14. The server according to claim 9, wherein: The operation of measuring the vocabulary variation includes: collecting and aggregating a fixed number of vocabulary patterns throughout the input video using clustering by selecting recurring patterns in the patterns in a fixed number of clusters, and reconstructing mouth movements on the fixed number of vocabulary patterns to generate a reconstruction cost of a score indicating the mouth movement variation.

15. The server according to claim 14, wherein: The vocabulary variation is measured without identifying the language of the vocabulary.

16. The server according to claim 9, wherein: Generating the prediction further includes applying a decision tree or a support vector machine to learn a separating function between subjects with the aphasic condition and subjects without the aphasic condition.

17. A non-transitory computer-readable medium having a storage device storing instructions for execution by a processor, the instructions comprising: receiving an input video of a subject, and performing a facial landmark generation operation on the received input video to define points associated with a region of interest; defining speech periods based on respective positions of the points associated with the region of interest; measuring pause frequency, repetition patterns, and lexical variation during the speech period without determining or applying semantic information associated with the speech; as well as The measures are aggregated to generate a prediction that an aphasia condition is present in the subject associated with the measures.

18. The non-transitory computer-readable medium of claim 17, wherein: Generating the facial landmarks includes: defining, for the region of interest including the subject's mouth, defined points outlining the subject's lips, measuring a temporal similarity metric of the defined points over time, and removing body movement and head movement from the temporal similarity metric to generate a temporal dissimilarity metric for the mouth, Defining the speech period includes filtering out mouth movements associated with out-of-plane rotations and jitters of the mouth, measuring a vertical distance between an upper lip and a lower lip, and performing a closing operation to generate a speech score indicative of the speech period, Measuring the pause frequency includes applying a threshold to the speech score and registering speech inactivity periods during the speech period as the pause frequency, Measuring the repetitive pattern includes defining a first mouth movement pattern and a second mouth movement pattern having certain lengths around a first time interval and a second time interval, respectively, during the speech period, and performing a dissimilarity comparison to obtain a number of repetitions over the time periods, and Measuring the vocabulary variation includes collecting and aggregating a fixed number of vocabulary patterns across the input video using clustering by selecting recurring patterns in the patterns in a fixed number of clusters, reconstructing mouth movements on the fixed number of vocabulary patterns to generate a reconstruction cost that is a score indicative of mouth movement variation.

19. The non-transitory computer-readable medium of claim 18, wherein: The vocabulary variation is measured without identifying the language of the vocabulary.

20. The non-transitory computer-readable medium of claim 17, wherein: Generating the prediction further includes applying a decision tree or a support vector machine to learn a separating function between subjects with the aphasic condition and subjects without the aphasic condition.

Citation Information

Patent Citations

  • Cognitive and linguistic assessment using eye tracking

    CN102245085A

  • Psycholinguistic evaluation method and system for Chinese aphasia

    CN108113651A