Vocal assessment

The method analyzes speech patterns to quickly and accurately identify mental health conditions, addressing the limitations of current assessment methods by providing a more efficient and accessible diagnostic tool.

WO2025093889A1PCT designated stage expired Publication Date: 2025-05-08PSYRIN LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/052784
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-11-01
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current assessment methods for psychiatric conditions, such as the Comprehensive Assessment of At-Risk Mental States (CAARMS), are time-consuming, require extensive training, and limit availability, making early identification and diagnosis challenging.

Method used

A computer-implemented method analyzing speech patterns in a speech sample to determine mental characteristics by assessing coherence, connectivity graphs, and syntactic assessments, allowing for quick and accurate identification of mental health conditions without direct interaction with a healthcare professional.

Benefits of technology

The method enables rapid and accurate identification of mental characteristics and conditions, such as major depression, schizophrenia, and bipolar disorder, improving early diagnosis and treatment accessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024052784_08052025_PF_FP_ABST
    Figure GB2024052784_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A method for analysing speech patterns in a speech sample obtained from a user, the method comprising: analysing the speech sample to determine different measures of coherence relating to the semantic arrangement of words contained within the speech sample; analysing the speech sample to determine different types of connectivity graphs for words contained within the speech sample; analysing the speech sample to determine a syntactic assessment of words contained within the speech sample; using the coherence values, the features derived from connectivity graphs and the syntactic assessment of the speech sample to determine one or more mental characteristics associated with the user from which the speech sample was obtained; and providing an output indicating the one or more identified mental characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] VOCAL ASSESSMENT

[0002] Field of the invention

[0003] This invention relates to a computer implemented method for analysing speech patterns in a speech sample obtained from a user, a data processing device configured to carry out such a method, a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry such a method, and a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out such a method.

[0004] Background

[0005] Assessment and triage for psychiatric conditions is typically complex and subjective, for example it takes, on average, up to 10 years to get a bipolar disorder diagnosis in the United Kingdom (UK). Early identification and diagnosis of psychiatric conditions, and of any potential changes in a psychiatric condition, can be key to ensuring effective treatment.

[0006] Assessment tools such as the Comprehensive Assessment of At-Risk Mental States (CAARMS) are presently used by mental health professionals and researchers to identify help-seeking young people who are at ultra-high risk (UHR) of developing psychosis. This is the present "gold standard" test used in the UK. However, using the CAARMS involves a 90-minute assessment and requires extensive training for the practitioner to be able to perform the assessment. The training required to be able to perform the assessment, and the time required to perform it each time, limits the availability of such assessments.

[0007] Summary of the invention

[0008] According to a first aspect of the invention, there is provided a method for analysing speech patterns in a speech sample obtained from a user, the method comprising: analysing the speech sample to determine different measures of coherence relating to the semantic arrangement of words contained within the speech sample; analysing the speech sample to determine different types of connectivity graphs for words contained within the speech sample; analysing the speech sample to determine a syntactic assessment of words contained within the speech sample; using the coherence values, the features derived from connectivity graphs and the syntactic assessment of the speech sample to determine one or more mental characteristics associated with the user from which the speech sample was obtained; and providing an output indicating the one or more identified mental characteristics.

[0009] Analysing a speech sample for speech patterns in order to identify one or more mental characteristics is advantageous because it can allow for quick and accurate identification of the mental characteristics. The method is also advantageous in that it can allow for mental characteristics to be identified without the need for direct interaction between the user and a trained healthcare professional. This could allow for the method to be made available for use by a far wider range of users, and more frequently, than would be possible for traditional assessment methods.

[0010] Analysing the speech sample for coherence, connectivity and syntactics allows for more accurate identification of mental characteristics that can indicate potential neurological disorders in a user, by determining particular features of the speech sample. This particular combination of analysis steps and methods improves the accuracy with which particular mental characteristics can be identified and differentiated between, as indicated by the determined features of speech sample. Identifying such mental characteristics can help to identify and differentiate between specific neurological conditions the user may be experiencing and can allow for multiple different psychiatric conditions to be identified and / or differentiated between.

[0011] The combination of analysing the speech sample for coherence, connectivity and syntactics allows for more accurate identification of mental characteristics, and for better differentiation between different mental characteristics. This is because this particular combination of analysis techniques captures changes in both the coherence and the syntactics of the speech sample, while also considering the link between the coherence and syntactics using the connectivity analysis. By determining the connectivity of the speech sample, the deviation in coherence and syntactics and how they relate / depend on each other can be considered.

[0012] This combination of determined features of the speech sample allows for a broader and more-layered sample of language characteristics to be determined in the speech sample, than would be possible with other methods. The complementary determined feature structure captures characteristics at different levels of language / speech and captures the connection between these levels; this is a particularly advantageous benefit of the method in terms of ability to identify mental characteristic specific speech alterations in a user. In a computer implementation of the method, this particular combination of analysis steps can also advantageously reduce the computational resources that may be needed to train the identification model used to implement the method.

[0013] Preferably, the connectivity graph is determined by assessing one or more of: semantic, contextual and grammatical relationships between words contained within the speech sample. Each of these advantageously indicate in a reliable manner the connectivity of the speech sample, thereby helping to characterise the connectivity of the speech. This can advantageously improve determination of the link between the coherence and syntactics of the speech sample.

[0014] Preferably, the syntactic assessment is determined by assessing the sophistication and complexity of grammatical resources exhibited in the words contained within the speech sample. These assessment methods advantageously indicate, in a reliable manner, the syntactics of the speech sample.

[0015] Preferably, the method further comprises analysing the speech sample to determine acoustic features. Analysing acoustic features is advantageous as it adds further levels of information regarding the speech sample, thereby allowing for more accurate identification and differentiation of mental characteristics that can indicate potential neurological disorders in the user from which the speech sample was obtained.

[0016] Preferably, the acoustic features comprise one or more of: spectral parameters, temporal parameters, frequency parameters, phonation parameters and volume- related parameters. Each of these parameters advantageously indicate, in a reliable manner, particular acoustic features of the speech sample that can help to indicate and differentiate between mental characteristics of the user from which the speech sample was obtained.

[0017] Preferably, the output of the method can be used to identify the presence of one or more mental health and / or neurological conditions in the user from which the speech sample was obtained. The output of the method comprises one or more identified mental characteristics that have been identified by the method. Advantageously, these identified mental characteristics can be used to determine mental health and / or neurological conditions in the user from which the speech sample was obtained. Further advantageously, these mental characteristics can be used to differentiate between one or more potential mental health / neurological conditions in the user. Preferably, the one or more mental health conditions that can be identified comprise: major depression, schizophrenia, bipolar disorder and clinical risk of psychosis. Each of these are severe mental health conditions, and so it is very advantageous to be able to specifically identify one or more of these conditions using only a speech sample obtained from a user.

[0018] Preferably, the output can be used to identify a risk of developing a mental health disorder in the user from which the speech sample was obtained. This is advantageous as the method can be used before the user has developed a severe mental health disorder, thereby allowing for preventative actions to be taken. This is particularly advantageous as the method can be used in circumstances where access to a healthcare professional may not be available, since the user has not yet fully developed the potential mental health disorder.

[0019] Preferably, the output of the method can be used to predict a transition from clinical high risk of psychosis to full psychosis in the user from which the speech sample was obtained. This is advantageous as it can identify particularly high-risk individuals so that available treatment resources can be allocated more effectively.

[0020] Preferably, the output of the method can be used to predict a relapse episode of psychosis in the user from which the speech sample was obtained, if the user has previously been diagnosed with a severe mental illness. This is advantageous as it can identify particularly high-risk individuals so that available treatment resources can be allocated more effectively. This is further advantageous as the method can predict a relapse before one takes place, potentially allowing for preventative actions to be taken.

[0021] Preferably, the output of the method can be used for monitoring of changes in symptomatology in the user from which the speech sample was obtained, if the user has been previously diagnosed with a severe mental illness. This is advantageous as it can provide healthcare professionals with updates on the condition of a user without the need for direct interaction between the healthcare professional and the user. This can allow for more frequent updates on the condition of the user to be recorded than may be possible if a direct interaction between the user and healthcare professional were required, particularly in situations where access to healthcare professionals may be limited. Preferably, the output of the method can be used for selecting, from one or more predetermined treatment options, the treatment option most likely to illicit a beneficial response in the user from which the speech sample was obtained, if the user has already been diagnosed with a severe mental illness. This is advantageous as it can allow for predetermined treatment options to be selected without the need for direct input from a healthcare professional.

[0022] According to a second aspect of the invention, there is provided a data processing device configured to carry out the method discussed above. This is advantageous as it can allow for the method to be implemented automatically using computer hardware.

[0023] According to a third aspect of the invention, there is provided a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method discussed above. This is advantageous as it can allow for the method to be implemented automatically using computer hardware.

[0024] According to a fourth aspect of the invention, there is provided a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method discussed above. This is advantageous as it can allow for the method to be implemented automatically using computer hardware.

[0025] Brief description of the figures

[0026] There now follows a brief description of preferred embodiments of the invention, by way of non-limiting examples, with reference being made to the following figures in which:

[0027] Figure 1 illustrates a schematic view of a method of analysing speech patterns in a speech sample obtained from a user to determine one or more mental characteristics associated with the user from which the speech sample was obtained;

[0028] Figure 2a illustrates an example connectivity graph that may be used to determine mental characteristics associated with a healthy user;

[0029] Figure 2b illustrates an example connectivity graph that may be used to determine mental characteristics associated with a user suffering from schizophrenia;

[0030] Figure 2c illustrates an example connectivity graph that may be used to determine mental characteristics associated with a user suffering from bipolar disorder;

[0031] Figure 3a illustrates a schematic diagram of a computer system suitable for implementing the method of the present disclosure; and

[0032] Figure 3b illustrates a schematic block diagram of a user device suitable for implementing the method. Detailed description

[0033] A method of analysing speech patterns in a speech sample obtained from a user to determine one or more mental characteristics associated with the user from which the speech sample was obtained, is referred to generally by reference numeral 100, as shown in Figure 1.

[0034] As shown schematically the method 100 of the invention comprises the following main steps for analysing speech patterns in a speech sample obtained from a user: analysing 101 the speech sample to determine different measures of coherence relating to the semantic arrangement of words contained within the speech sample; analysing 102 the speech sample to determine different types of connectivity graphs for words contained within the speech sample; analysing 103 the speech sample to determine a syntactic assessment of words contained within the speech sample; using 104 the coherence values, the features derived from connectivity graphs and the syntactic assessment of the speech sample to determine one or more mental characteristics associated with the user from which the speech sample was obtained; and providing 105 an output indicating the one or more identified mental characteristics.

[0035] In the embodiment method shown, three distinct analysis steps (101, 102, 103) are used in order to determine one or more mental characteristics associated with the user from which the speech sample was obtained. The inventors have determined that these three analysis steps in combination provide an improved method for identifying, and differentiating between, mental characteristics of the user. These mental characteristics can be used to identify and differentiate between mental health and / or neurological conditions in the user.

[0036] This particular combination of analysis steps and techniques offer a clear improvement over previous attempts to identify mental characteristics from speech in that multiple different mental characteristics can be identified and differentiated between simultaneously. This allows for the method 100 to be used to assist with diagnosing multiple mental health and / or neurological conditions simultaneously using an obtained speech sample, with a high degree of accuracy. The method 100 is used to analyse speech patterns in a speech sample obtained from a user. The speech sample may be any suitable sample of speech, comprising an audio recording and / or a transcription. In some examples, the speech sample may be recorded in response to picture description and dream recall tasks. This is advantageous as such tasks lead the user to respond in free speech, rather than directed speech (for example, by reading written text). Analysing a sample of free speech, using the method 100, may result in more accurate identification of mental characteristics.

[0037] In examples where the method is implemented using a computer system, an audio recording of speech obtained from a user may first be transcribed in order to convert the audio recording into words. This may be performed using any suitable speech recognition system. In some examples, this step may be used in order to produce the speech sample used in the method.

[0038] The first step 101 involves analysing the speech sample to determine different measures of coherence relating to the semantic arrangement of words contained within the speech sample. Coherence refers to the logical arrangement of how an individual speaks, deriving features that relate to how each word in a phrase is connected with other parts of the phrase. Coherent speech is indicative of the ability to speak with normal levels of continuity, rate and effort, and the linkage of ideas and language together to form sensical, connected speech.

[0039] The second step 102 involves analysing the speech sample to determine different types of connectivity graphs for words contained within the speech sample. Analysing the connectivity of speech refers to the analysis of networks created by the transformation of text into graphical representations, based on semantic, contextual and grammatical relationships between the words included in the speech. For example, and as illustrated in figures 3a-3c, in such graphical representations words 303 can be represented by nodes 301, and their use together can be represented by edges 302. In order to implement the method 100, it should be appreciated that in some examples graphical representations of the connectivity may not actually be visually represented. For example, in a computer implementation the connectivity graph may instead be produced only in computer memory.

[0040] Figure 3a illustrates an example of a connectivity graph that may be produced during the second step 102 of analysing the speech sample. In some examples, the speech sample may be a recording of free speech produced by the user in response to a visual prompt, such as an image. In the graphical representation 300 of the connectivity graph, words 303 from the speech sample are represented by nodes 301 and connections between words as determined by analysis of the speech sample are illustrated by edges 302. As discussed above, a connectivity graph as illustrated in figures 3a-3c may be used to determine one or more mental characteristics associated with the user from which the speech sample was obtained. In the case of the graphical representation 300 illustrated in figure 3a, the mental characteristics determined from this connectivity graph could indicate that the speech sample was obtained from a healthy user.

[0041] Figure 3b illustrates a further example of a connectivity graph that may be produced during the second step 102 of analysing the speech sample. In the case of the graphical representation 310 illustrated in figure 3b, the mental characteristics determined from this connectivity graph could indicate that the speech sample was obtained from a user suffering from schizophrenia.

[0042] Figure 3c illustrates a further example of a connectivity graph that may be produced during the second step 102 of analysing the speech sample. In the case of the graphical representation 320 illustrated in figure 3c, the mental characteristics determined from this connectivity graph could indicate that the speech sample was obtained from a user suffering from bipolar disorder.

[0043] The third step 103 involves analysing the speech sample to determine a syntactic assessment of words contained within the speech sample. This refers to the range and the sophistication of grammatical resources exhibited in the language used in the speech sample.

[0044] The fourth step 104 involves using the coherence values, the features derived from connectivity graphs and the syntactic assessment of the speech sample to determine one or more mental characteristics associated with the user from which the speech sample was obtained. The combination of these three analysis techniques allows for accurate determination of particular mental characteristics in the user from which the speech sample was obtained. For example, an assessment of the coherence of the speech sample may be sufficient to identify mental characteristics that could indicate that the user is suffering from (or may be at risk of developing) either bipolar disorder or schizophrenia. However, by using coherence values, the features derived from connectivity graphs and the syntactic assessment, the method 100 can determine mental characteristics with sufficient accuracy determine which of these two conditions is indicated.

[0045] In the fifth step 105, an output indicating the one or more identified mental characteristics is provided. This output can be used, as discussed above, to determine one or more of the following :

[0046] - the presence of one or more mental health and / or neurological conditions in the user from which the speech sample was obtained;

[0047] - whether the user is at risk of developing one or more mental health and / or neurological conditions; if a user is at risk of a transition from clinical high risk of psychosis to full psychosis; and monitoring of changes in symptomatology in the user from which the speech sample was obtained, if the user has been previously diagnosed with a severe mental illness.

[0048] The output can also be used for selecting, from one or more predetermined treatment options, the treatment option most likely to illicit a beneficial response in the user from which the speech sample was obtained, if the user has already been diagnosed with a severe mental illness.

[0049] In some examples, further analysis steps may be undertaken in addition to those outlined in steps one to three above (101, 102, 103). For example, a further step may include analysing acoustic features. This may include, for example, extracting different acoustic and vocal features from the speech sample. In such examples, this analysis is performed on a speech sample comprising an audio recording obtained from a user. The different acoustic features can fall under several categories, including spectral parameters, frequency parameters, and volume. In some examples, the acoustic parameters may include one or more of: spectral parameters, temporal parameters, frequency parameters, phonation parameters and volume-related parameters. In examples where the method is implemented on a computer, signal processing software may be used to extract one or more of these features from the speech sample.

[0050] In some examples, the method 100 may be partially, or completely, implemented using computer hardware. Figure 2a illustrates a schematic diagram of a computer system 200 suitable for implementing the method 100. The system 200 comprises a user device 201 connected to a server 203 via a network 202. The user device 201 is configured to implement the method 100. The system 200 may be utilised to implement one, some, or all of the techniques or features described herein. In some examples, the method 100 may be fully implemented on a user device 201, without the need for access to a network 202 and / or a server 203. In other examples, the user device 201 may require a connection to a network 202 and / or a server 203 in order to implement the method 100.

[0051] The user device 201 may be implemented using bespoke hardware, or any conventional computer system such as, for example, a laptop, cellular telephone, desktop computer or tablet computer. Having readily available access to computer implementation of the method, provided via the system 200, can be beneficial as it may allow for the user to make use of the method more easily.

[0052] Figure 2b illustrates a schematic block diagram of a user device 210 suitable for implementing the method 100. The user device 210 comprises one or more processors 212 in communication with memory 214. The memory 214 is an example of a computer readable storage medium. The one or more processors 212 are also in communication with one or more input devices 216 and one or more output devices 218. The various components of the user device 210 may be implemented using generic means for computing known in the art. For example, the input devices 216 may comprise a keyboard or mouse and the output devices 218 may comprise a monitor or display, and an audio output device such as a speaker. In addition, the user device 210 comprises network access circuitry, such as a modem or network adaptor, to provide access to the network to obtain the software application and data. Alternatively, a different data entry means, such as a portable disc drive or data interface may be provided in addition to, or instead of, the network adaptor 220 to provide data to the user device 210.

[0053] It will be appreciated that, in some examples, the various hardware units may be integrated with one another. For example, where the user device 210 is provided by a tablet computer or cellular telephone, a display of the device 210 may provide both the display of the output device 218 and a touch sensor providing the input device 216.

Claims

CLAIMS1. A method for analysing speech patterns in a speech sample obtained from a user, the method comprising: analysing the speech sample to determine different measures of coherence relating to the semantic arrangement of words contained within the speech sample; analysing the speech sample to determine different types of connectivity graphs for words contained within the speech sample; analysing the speech sample to determine a syntactic assessment of words contained within the speech sample; using the coherence values, the features derived from connectivity graphs and the syntactic assessment of the speech sample to determine one or more mental characteristics associated with the user from which the speech sample was obtained; and providing an output indicating the one or more identified mental characteristics.

2. A method according to claim 1, wherein the connectivity graph is determined by assessing one or more of: semantic, contextual and grammatical relationships between words contained within the speech sample.

3. A method according to claims 1 or 2, wherein the syntactic assessment is determined by assessing the sophistication and complexity of grammatical resources exhibited in the words contained within the speech sample.

4. A method according to any preceding claim, wherein the method further comprises analysing the speech sample to determine acoustic features.

5. A method according to claim 4, wherein the acoustic features comprise one or more of: spectral parameters, temporal parameters, frequency parameters, phonation parameters and volume-related parameters.

6. A method according to any preceding claim, wherein the output can be used to identify the presence of one or more mental health conditions in the user from which the speech sample was obtained.

7. A method according to claim 6, wherein the one or more mental health conditions comprise: major depression, schizophrenia, bipolar disorder and clinical risk of psychosis.

8. A method according to any preceding claim, wherein the output can be used to identify a risk of developing a mental health disorder in the user from which the speech sample was obtained.

9. A method according to any preceding claim, wherein the output can be used to predict a transition from clinical high risk of psychosis to full psychosis in the user from which the speech sample was obtained.

10. A method according to any preceding claim, wherein the output can be used to predict a relapse episode of psychosis in the user from which the speech sample was obtained, if the user has previously been diagnosed with a severe mental illness.

11. A method according to any preceding claim, wherein the output can be used for monitoring of changes in symptomatology in the user from which the speech sample was obtained, if the user has been previously diagnosed with a severe mental illness.

12. A method according to any preceding claim, wherein the output can be used for selecting, from one or more predetermined treatment options, the treatment option most likely to illicit a beneficial response in the user from which the speech sample was obtained, if the user has already been diagnosed with a severe mental illness.

13. A data processing device configured to carry out the method of any preceding claim.

14. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 12.

15. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any of claims 1 to 12.