Automated patient communication ability assessment

An automated method using a machine learning model processes patient-conversation audio to provide objective and reliable communication ability assessments, addressing the inefficiencies and subjectivity of traditional methods, particularly benefiting patients with limited mobility and enhancing clinical trial monitoring.

WO2025190705A1PCT designated stage Publication Date: 2025-09-18F HOFFMANN LA ROCHE & CO AG
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/EP2025/055670
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-11
Filing Date
2025-03-03
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing methods for assessing communication abilities of patients are time-consuming, subjective, and lack test-retest reliability, especially when conducted outside a patient's regular environment, and are difficult for patients with limited mobility.

Method used

An automated method using a machine learning model, such as a large language model like GPT4, processes audio signals from patient-conversation recordings to generate objective communication ability scores and diagnoses, enhancing objectivity and reliability.

Benefits of technology

The method provides accurate, efficient, and reliable communication ability assessments, enabling monitoring of health conditions like ASD and dementia, and supports clinical trials with improved test-retest reliability and reduced travel requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025055670_18092025_PF_FP_ABST
    Figure EP2025055670_18092025_PF_FP_ABST
Patent Text Reader

Abstract

In various examples there is a computer-implemented method of assessing communication ability of a patient in order to diagnose a health condition. The method comprises receiving an audio signal of a conversation between a patient and a study partner. The audio signal, or data derived from the audio signal, is input to a model. The model is large language model. A prompt is input to the model, the prompt requesting the model to generate an output comprising a numerical score in a specified range, and requesting the numerical score to quantify a health condition of the patient using a specified clinical scale. The method comprises receiving output from the model comprising the numerical score.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATED PATIENT COMMUNICATION ABILITY ASSESSMENTFIELD

[0001] The present invention relates to health conditions affecting the speech of patients and the use of machine learning models for automated assessment of communication abilities of patients for medical diagnosis or symptom measurement.BACKGROUND

[0002] Communication abilities of patients are assessed by medical practitioners for diagnosis and monitoring of various health conditions including but not limited to Autism Spectrum Disorder (ASD), Dementia, and Major Depressive Disorder (MDD). Medical practitioners typically assess communication abilities of patients through observation of the patient which is time consuming and difficult to keep objective. Test- retest reliability is typically low. When patients are being assessed outside of their regular environment, such as in a medical clinic or medical setting, the patient’s behaviour may be influenced by the setting so that it is difficult to take an accurate measurement. It can also be difficult for patients to travel long distances to attend the medical clinic for assessment so that many patients are not properly assessed. Where patients have limited mobility, travel to a clinic for assessment is not straightforward.

[0003] In the case of clinical trials where communication abilities of patients are to be monitored, assessment is typically done manually by having a caregiver, teacher or clinician complete a survey. The surveys are time-consuming, subjective and suffer from poor test-retest reliability.

[0004] The embodiments described below are not limited to implementations which solve any or all of the disadvantages of known ways of assessment of communication abilities of patients for medical diagnosis of health conditions affecting speech.SUMMARY

[0005] The following presents a simplified summary of the disclosure in order to provide a basic understanding to the reader. This summary is not intended to identify key features or essential features of the claimed subject matter nor is it intended to be used to limit the scope of the claimed subject matter. Its sole purpose is to present a selection ofconcepts disclosed herein in a simplified form as a prelude to the more detailed description that is presented later.

[0006] In various examples there is a computer-implemented method of assessing communication ability of a patient in order to diagnose a health condition. The method comprises receiving an audio signal of a conversation between a patient and a study partner. The audio signal, or data derived from the audio signal, is input to a model. The model is a large language model. A prompt is input to the model, the prompt requesting the model to generate an output comprising a numerical score in a specified range, and requesting the numerical score to quantify a health condition of the patient using a specified clinical scale, taking into account the audio signal or data derived from the audio signal. The method comprises receiving output from the model comprising the numerical score.

[0007] Many of the attendant features will be more readily appreciated as the same becomes better understood by reference to the following detailed description considered in connection with the accompanying drawings.DESCRIPTION OF THE DRAWINGS

[0008] The present description will be better understood from the following detailed description read in light of the accompanying drawings, wherein:FIG. l is a schematic diagram of a machine learning model being used to compute a communication ability score for a variety of purposes;FIG. 2 is a schematic diagram of the machine learning model of FIG. 1 and illustrating determining whether the model has convergent validity with a human doctor;FIG. 4 is a graph of data obtained from a plurality of patients, with the x axis representing an expressive communication score predicted by generative pretrained transformer 4 (GPT 4) for a specified clinical scale, and the y axis representing the expressive communication score assigned by a human doctor expert in using the specified clinical scale;FIG. 5 A is a graph of Vineland Adaptive Behaviour Scale (VABS) expressive communication scores assigned by a human doctor for each of a plurality of patients, on a first occasion (x axis) and a second occasion (y axis);FIG. 5B is a graph of VABS expressive communication scores predicted by GPT 4 for each of the plurality of patients of FIG. 5 A, on a first occasion (x axis) and a second occasion (y axis);FIG. 6 A is a graph as in FIG. 4 except that each data point is a mean taken over 6 score instances of the same patient conversation;FIG. 6B is a graph as in FIG. 6A where rather than a mean a 50% statistic is plotted;FIG. 7 is a schematic diagram of an a model for use in the method of FIG. 1;FIG. 8 is a schematic diagram of a transformer-based machine learning language model;FIG. 9 illustrates an exemplary computing-based device for implementing a communication assessment and diagnosis tool;FIG. 10A is a graph of correlation between actual VABS expressive communication scores (y-axis) and predicted scores from GPT4 (x-axis), where each dot is a participant (n=52 and for a case when concatenating all conversations from each participant together, and averaging the results across 3 iterations;FIG. 10B is a is a graph of correlation between actual VABS expressive communication scores (y-axis) and predicted scores from GPT4 (x-axis), where each dot is a participant (n=52) and for the case when using individual conversations, and taking the median score across all conversations across 3 iterations;FIG. 11 A is a graph of correlation between actual VABS expressive communication scores (y-axis) and predicted scores from GPT4 (x-axis), where each dot is a participant (n=52) and when training examples in the prompt;FIG. 1 IB is a graph of correlation between actual VABS expressive communication scores (y-axis) and predicted scores from GPT4 (x-axis), where each dot is a participant (n=52) and when using no training examples in the prompt;FIG. 12A is a graph of correlation between GPT4 predictions of the VABS expressive communication scores (y-axis) versus Words per Sentence (WpS; x- axis);FIG. 12B is a graph of correlation between GPT4 predictions of the VABS expressive communication scores (y-axis) versus Utterance Duration (UD; x-axis);FIG. 12C is a graph of correlation between GPT4 predictions of the VABS expressive communication scores (y-axis) versus Age of Acquisition (AoA; x- axis);FIG. 13 is a box plot of GPT4 predictions of VABS expressive communication scores (y-axis) for ASD (intelligence quotient (IQ) < 70) participants, ASD (IQ >70) participants and neurotypical control (NTC) participants within each age group comprising Children, Adolescents and Adults (x-axis).Like reference numerals are used to designate like parts in the accompanying drawings.DETAILED DESCRIPTION

[0009] The detailed description provided below in connection with the appended drawings is intended as a description of the present examples and is not intended to represent the only forms in which the present examples are constructed or utilized. The description sets forth the functions of the examples and the sequence of operations for constructing and operating the examples. However, the same or equivalent functions and sequences may be accomplished by different examples.

[0010] As mentioned above, existing ways of assessing communication abilities of patients are time consuming and have poor test-retest reliability. Strong test-retest reliability means that the same patient conversation (such as in a transcript, audio or video recording), when assessed by the same medical practitioner on separate occasions, is given similar communication ability scores. Poor test-retest reliability means that the same patient conversation, when assessed by the same medical practitioner on separate occasions, is given dissimilar communication ability scores. As a result it is difficult to monitor communication ability accurately in order to monitor progression of a disease such as dementia or ASD.

[0011] Traditional methods for symptom assessment, such as clinical scales, are often hampered by subjectivity, rater biases, time-intensiveness, and reliance on episodic recall of events. These limitations can compromise their reliability and sensitivity, particularly in the context of clinical trials.

[0012] Health conditions affecting the speech of patients are numerous and include but are not limited to: neurological disorders encompassing bothneurodevel opmental and neurodegenerative conditions such as autism, Alzheimer’s and Parkinson’s disease, mental and mood disorders such as major depressive disorder, or other disorders such as burnout. Methods for measuring speech ability of patients are crucial for monitoring changes over time, particularly in the context of clinical trials. Traditional symptom measurement methodologies, such as the Vineland Adaptive Behaviour Scales (VABS), Autism Impact Measure (AIM), and the Autism Diagnostic Observation Schedule (ADOS) are heavily reliant on subjective evaluations by parents, teachers, or clinicians. While these methods are the cornerstone of neuropsychiatric evaluation, they are time-consuming, subjective, and dependent on memory recall of past events. This can lead to poor inter-rater reliability and placebo effects, demanding larger samples to observe therapeutic responses. This lack of objectivity and reliability can render these methods insensitive to subtle changes in symptomatology over time, a critical factor when evaluating the efficacy of a drug during a clinical trial.

[0013] The inventors have recognized that an automated way to monitor communication ability, particularly speech ability, of a patient brings significant benefits. Medical practitioners typically assess communication abilities of patients through observation of the patient which is time consuming and difficult to keep objective; whereas using an automated method will enhance objectivity and improve efficiency. An automated method that can be used in a patient’s home is particularly beneficial since when patients are being assessed outside of their regular environment, such as in a medical clinic or medical setting, the patient’s behaviour may be influenced by the setting so that it is difficult to take an accurate measurement of the patient’s everyday state. Automated assessments at home avoid the need for patients to travel long distances to attend the medical clinic for assessment and increases patient access to medical services.

[0014] In the case of clinical trials where communication abilities of patients are to be monitored, assessment is typically done manually by having a caregiver, teacher or clinician complete a survey. Using an automated assessment reduces the time needed and increases test-retest reliability as now explained.

[0015] FIG. 1 is a schematic diagram of a model 104 being used to compute a communication ability score 118 for a variety of purposes. A patient 114 and a study partner 112 are in a domestic setting having a conversation 100. The conversation is recorded by a microphone in a tablet computer 116 in the domestic setting. The tablet computer has an application installed on it enabling the tablet computer to display a userinterface guiding the study partner 112 how to record the conversation 100 and giving a topic for the conversation and optionally some opening statements to be read out by the patient 114 and / or study partner 112. The conversation 100 is recorded by the tablet computer 116 and sent to the model 104 with appropriate permissions and consent from the study partner 112 and the patient 114. The conversation 100 is an audio signal which is optionally compressed and encrypted before being sent to the model 104 via a wired or wireless connection. In some examples, the audio signal is converted to text using speech to text functionality at the tablet computer 116 before being sent to the model 104.

[0016] The model 104 is a machine learning model which is a large language model. A language model is functionality to generate text or speech in response to a prompt. A machine learning model is a representation of a plurality of training examples which may be used to generate new examples, or to make a prediction about an example which was not in the training examples. In some examples, the language model is a large language model, which is a language model having over one billion parameters where a parameter is a neural network weight. In some cases a language model is a neural network model able to compute contextualized embeddings of tokens where a token is a piece of text. In some examples the language model is a transformer-based language model. A transformer is a type of machine learning model that has a parallel multi-head attention mechanism. The term “transformer-based” refers to a model that comprises a parallel multi-head attention mechanism. In various examples the model 104 is any of: generative pre-trained transformer GPT 4, LlaMA, Claude, PaLM, BLOOM.

[0017] In some cases the model 104 is a pre-trained model 104 which has been adapted to a particular health condition using few shot learning or other adaptation techniques. In the case of few shot learning the model 104 is a pre-trained model such as any of GPT 4, LlaMA, Claude, PaLM, BLOOM and is adapted by training it using tens of examples of conversations having known communication ability scores with the specified clinical scale and known medical diagnoses for the specified medical condition. Once the model has been adapted to a particular health condition the prompt may omit details about which clinical scale to use and which health condition to diagnose.

[0018] As explained with reference to FIG. 2 below, the model 104 has been determined to have convergent validity with a human doctor working in the field of a health condition to be diagnosed by the model 104. The term “convergent validity” is explained in more detail with reference to FIG. 2.

[0019] The model 104 is configured to receive as input a prompt 106 as well as the audio signal 102 or data derived from the audio signal 102 In some cases the audio signal comprises a plurality of audio signals of conversations between the patient and the study partner, where the audio signals are concatenated. Using a plurality of concatenated conversations (either as audio signals or as text transcriptions) is found to improve performance.

[0020] Using the inputs the model 104 computes an output comprising a communication ability score 118 which is a number in a specified range, where the range is specified in the prompt 106. The model 104 optionally also outputs a diagnosis 120 as to whether or not the patient 114 has a specified health condition such as ASD, dementia or depression. Optionally the model 104 outputs examples of symptoms of a health condition observed in the conversations. In some cases the prompt 106 omits a request to give examples of symptoms of a health condition observed in the conversation(s) and this is unexpectedly found to give improved accuracy of the communication ability score 118.

[0021] The communication ability score 118 is made available to a medical doctor 110 using a computing device 108 such as by displaying the communication ability score in a graphical user interface having tools to assist the medical doctor in checking and assessing the communication ability score 118. The diagnosis 120, where available, is also made available by display at the graphical user interface. Where the prompt includes a request to give examples of symptoms of a health condition observed in the conversation(s) these examples are made available by display at the graphical user interface as candidates for a medical practitioner to select for saving in a patient record or for input to a computer.

[0022] In some cases the communication ability score 118 is used as part of a clinical trial 122 whereby the process indicated in FIG. 1 is repeated at intervals through the clinical trial to monitor progress of the patient through the clinical trial. In some cases the communication ability score 118 is used as part of a clinical trial 122 wherein a specified range of the communication ability score 118 is used as a selection criteria for participants to be included in the clinical trial 122. In some cases the communication ability score 118 is used as part of a dose regime 124 determination by monitoring how the patient responds to a pharmacological treatment for the health condition and adjusting a dose of the pharmacological treatment accordingly. In some cases the communication ability score 118 is used to control a digital intervention 126 to treat the patient, wherebythe communication ability score is used to adjust a frequency or length of time, or time of day at which the digital intervention is available to the patient 114. In some examples the digital intervention is a computing device deploying a conversation game between the patient and an avatar in augmented reality or virtual reality.

[0023] Thus FIG. 1 shows a computer-implemented method of assessing communication ability of a patient 114 in order to diagnose a health condition affecting speech such as a neurological disorder, or to measure a symptom of a neurological disorder. An audio signal 102 of a conversation 100 between a patient 114 and a study partner 112 is received at model 104. By using an audio signal it is possible to capture observations of the patient in an unobtrusive manner and which enables high fidelity, that is the observations are accurate since the audio signal is recorded automatically and any errors introduced by background noise or sensor noise of the microphone can be ameliorated through digital signal processing techniques.The audio signal 102, or data derived from the audio signal, is input to the model 104, the model being a large language model. Using a large language model is found to give particularly good convergent validity with a human doctor as explained in more detail below.A prompt 106 is input to the model, the prompt requesting the model to generate an output comprising a numerical score in a specified range, and requesting the numerical score to quantify a health condition of the patient using a specified clinical scale. By using a prompt in this way it is possible to use different ranges for the numerical score or to use different clinical scales, in an accurate and efficient manner. An output from the model is received comprising at least the numerical score. The numerical score is stored and / or presented in a graphical user interface of a tool used by a medical doctor 110 at computing device 108. The method gives a fully automated, efficient and accurate method of assessing communication ability of a patient for use in diagnosis of a health condition or measuring the symptom of a health condition. The health condition affects speech and may be an neurological disorder such as ASD or dementia, a mental or mood condition such as MDD or other conditions such as burnout.

[0024] In some cases the communication ability is an expressive communication ability and the clinical scale is Vineland Adaptive Behaviour Scale. Particularly good convergent validity is found in this case as explained below.

[0025] In some cases the prompt requests the model to diagnose the health condition and the output from the model also comprises a diagnosis of whether or not the patient has the health condition. This gives the benefit of an automated diagnosis which is computed using medical knowledge encoded in the model 104. Medical knowledge is encoded in the model 104 during training of the model 104 through training data comprising medical text books, medical scientific papers, medical training manuals, patient records and other medical knowledge.

[0026] In the example of FIG. 1 the method comprises receiving the audio signal 102 by using a microphone in a portable or wearable computing device to capture the audio signal 102. In some cases the method comprises defining a topic of the conversation via a user interface of the portable or wearable computing device such as tablet computer 116 of FIG. 1.

[0027] In an example with particularly good convergent validity with human doctors, the model is any of: generative pre-trained transformer 4, GPT4, versions 0.2.5- 0.2.7, the clinical scale is Vineland Adaptive Behaviour Scale, and the neurological disorder is Autistic Spectrum Disorder.

[0028] In various examples the prompt 106 directs the model to respond as: an expert and experienced consultant psychiatrist that specializes in autism spectrum disorder; or an expert and experienced medical doctor that specializes in dementia. In this way the output from the model 104 takes into account medical knowledge held by the specified experts, about the specified neurological disorder; since knowledge which is semantically similar to the prompt, and which was encoded into the model during training, is taken into account when generating the output from the model.

[0029] In some examples, the prompt comprises one or more labelled training examples in addition a request for the model to generate an output comprising a numerical score in a specified range, and requesting the numerical score to quantify a health condition of the patient using a specified clinical scale. A labelled training example is an example conversation between a person and a study partner (where the person is not the same person as the patient and the study partner is either the same study partner or a different study partner) and where the numerical score on the specified clinical scale is specified in the prompt either exactly or as a category (such as very high score, high score, medium score, low score, very low score).

[0030] The inventors have unexpectedly found that the model 104 is able to process and analyze unstructured conversations, in order to capture a broad spectrum of linguistic and paralinguistic features, including those that are not predefined by researchers. A non-exhaustive list of linguistic features and paralinguistic features is: the number of words spoken per sentence, utterance duration, lexical maturity, turn-taking behaviour.

[0031] The inventors have found that where the prompt comprises one or more labelled training examples in addition to the request for the model to generate a numerical score, the convergent validity of the model is improved (as explained in more detail below).

[0032] As illustrated in FIG. 1, output from the model 104 is provided to a human medical doctor 110 at a user interface of a computing device 108, together with medical records of the patient 114, to enable the human medical doctor 110 to confirm the automated diagnosis, or the automated communication ability score. In some cases the automated diagnosis or the automated communication ability score is presented to the medical doctor as a candidate at the user interface. The medical doctor is able to review the patient records, conversation transcript, audio or video recording of the conversation, and accept or decline the candidate diagnosis. The medical doctor is able to edit a rationale for the candidate diagnosis produced by the model 104.

[0033] In some cases the patient 114 is a member of a clinical trial of a pharmaceutical treatment for the health condition, and the method of FIG. 1 is repeated during the clinical trial in order to assess whether communication ability of the patient 114 is influenced by the pharmaceutical treatment. Because the method of FIG. 1 exhibits strong test-retest reliability the method of FIG. 1 is especially useful for monitoring progression of the neurological disorder, such as by measuring a symptom of the neurological disorder, in a clinical trial or otherwise.

[0034] In some cases the patient is being treated with a pharmaceutical treatment for the health condition. The method of FIG. 1 is repeated during the pharmaceutical treatment and a dose regime 124 of the pharmaceutical treatment is automatically adjusted, using rules, according to changes in the numerical score. The dose regime is stored and output to the medical doctor at the graphical user interface on computing device 108.

[0035] In some cases a digital intervention 126 is controlled automatically using the communication ability score from the method of FIG. 1. In an example, the digital intervention is deployed using an augmented reality or virtual reality computing device. The automated method comprises generating a time limited password to the augmented reality or virtual reality computing device and configuring the time limited password at the augmented reality or virtual reality computing device.

[0036] FIG. 2 is a schematic diagram of the machine learning model of FIG. 1 and illustrating determining whether the model has convergent validity with a human doctor. The term “convergent validity” refers to a metric for indicating how well a measure correlates with other measures of the same quantity. In the example of FIG. 2 the quantity is communication ability and two metrics being compared are the model 104 and a human doctor 110. The model is either prompted to act as a doctor expert in the field of the neurological disorder, or has been adapted to act in such a manner using few-shot learning or other adaptation techniques. The human doctor 110 is an expert in the field of the health condition . The conversation is observed by the medical doctor 110. A communication ability score from the medical doctor 200 is input to a convergent validity assessment 204. A diagnosis from the human doctor 202 is also input to the convergent validity assessment 204.

[0037] The audio signal of the conversation 100, or information derived from the audio signal 102, is input to the model 104. The model 104 also receives a prompt 106 as explained with reference to FIG. 1. In some cases the prompt 106 comprises one or more labelled training examples in addition to a request for a numerical score as explained above. The prompt 106 may be stored in memory and input at the appropriate time. Using the prompt 106 and the audio signal 102, or information derived from the audio signal, the model 104 generates a communication ability score 118 (which is an example of a measurement of a symptom) and optionally a diagnosis 120. These are input to the convergent validity assessment 204.

[0038] The process of FIG. 2 is repeated for several patients to obtain a pair of values for each patient where the pair comprises a score from the model 104 and a corresponding score from the human doctor 110. FIG. 4 has an example of such a graph. In FIG. 4 each pair of values comprises a VABS score from a human doctor (plotted on the y axis) and a VABS score from the model 104, where the model is GPT 4, (plotted on the x axis). It is seen that the points on the graph generally follow a straight lightindicating strong convergent validity. The data in the graph of FIG. 4 was obtained empirically using patients known to be autistic. The data in the graph of FIG. 4 has a coefficient of correlation, r, of 0.78 and a statistical significance p-value of 3.325 e'7or lower. Intra-rater reliability was found to be strong: intraclass correlation coefficient ICC = 0.86 (similar answers over multiple requests).

[0039] The process of FIG. 2 is carried out, prior to inputting the prompt to the model in the process of FIG. 1, to determining that the model has convergent validity with a human doctor expert in diagnosing the health condition. Convergent validity occurs where a linear relationship between communication ability scores on the clinical scale from the model for a plurality of patients and communication ability scores on the clinical scale for the same patients from the human doctor has a coefficient of correlation, r, with a modulus above 5 and a statistical significance p-value of 0.05 or lower.

[0040] FIG. 4 is a graph of data obtained from a plurality of patients, with the x axis representing an expressive communication score predicted by generative pre-trained transformer 4 (GPT 4) for a specified clinical scale, and the y axis representing the expressive communication score assigned by a human doctor expert in using the specified clinical scale.

[0041] FIG. 5 A is a graph of Vineland Adaptive Behaviour Scale (VABS) expressive communication scores assigned by a human doctor for each of a plurality of patients, on a first occasion (x axis) and a second occasion (y axis). The coefficient of correlation r is 0.89 and the statistical significance p-value is p<lxl0'5.

[0042] FIG. 5B is a graph of VABS expressive communication scores predicted by GPT 4 for each of the plurality of patients of FIG. 5 A, on a first occasion (x axis) and a second occasion (y axis). The coefficient of correlation r is 0.98 and the statistical significance p-value is p<lxl0'5. Thus FIGs. 5A and 5B demonstrate that the model 104 is able to have test-retest ability on a par with a human doctor.

[0043] FIG. 6 A is a graph as in FIG. 4 except that each data point is a mean taken over 6 score instances of the same patient conversation. The coefficient of correlation r is 0.69 and the statistical significance p-value is 1.082 e'7.

[0044] FIG. 6B is a graph as in FIG. 6A where rather than a mean a 50% statistic is plotted. The coefficient of correlation r is 0.71 and the statistical significance p-value is 2.61 e8.

[0045] FIG. 7 is a schematic diagram of an example of the model 104 of FIG. 1 in the case that the model 104 is able to take audio signals as input directly (without first converting the audio signals to text). The model 104 comprises an audio encoder such as but not limited to Wav2Vec2 and HuBERT, a projection mechanism W and a language model. The language model is a pre-trained large language model such as GPT 4, BLOOM, LLaMA or large language model. An example of a language model is described with reference to FIG. 8.

[0046] The audio encoder takes audio signals and divides them into segments. The contents of each segment is represented by a vector (by linearly mapping signal values to a vector) which is referred to as an input embedding. For each segment, the location of the segment in the audio signal is encoded using a position embedding which is concatenated with the input embedding. The vectors may be discretized to remove noise.

[0047] The projection mechanism W is a trainable projection matrix to convert an embedding (denoted Zv) of an audio signal (denoted Xv) into language embedding tokens (denoted Hq) which have the same dimensionality of a word embedding space in the language model. When an audio signal Xv and language prompt (denoted Xq) are received these are converted into a sequence of embeddings of tokens. In FIG. 7 the sequence comprises three embeddings of prompt tokens Hq followed by three embeddings of audio segments Hv. The sequence of embeddings are input to the language model which generates a language response (denoted Xa) comprising embeddings which are converted into text. The language response is text comprising a communication ability score and / or a diagnosis and a rationale for the diagnosis.

[0048] The projection mechanism W is trained by freezing the weights of the audio encoder and the language model, and updating only trainable parameters of the projection mechanism W (which is a matrix). The training data comprises a plurality of training examples such as those from FIG. 3 item 300.

[0049] Subsequent to training the projection mechanism, fine tuning is carried out end-to-end. During this stage of the training only the audio encoder weights are kept frozen. The pre-trained weights of the large language model and the weights of the projection mechanism W are updated. The training is supervised training using any suitable training algorithm such as back propagation with gradient descent.

[0050] FIG. 8 shows an exemplary architecture of a language model which is a transformer-based machine learning language model. The example of FIG. 8 is one example of an architecture of model 104 and is not intended to be limiting. In some examples, the transformer-based machine learning language model is an example of a large language model 104 as described previously. However, it is also possible to use other large language models which do not have a transformer-based architecture.

[0051] The model 104 comprises a plurality of layers of nodes interconnected by edges. There may be as many as several hundred layers of nodes in some examples. The model 800 comprises a plurality of transformers implementing self-attention and comprises a plurality of attention heads. Attention heads are used to direct the neural network to focus on a subset of features or tokens in an input sequence thereby learning different representations from the different positions of the tokens in an input sequence. Attention heads and transformers provide the model with a better capability to learn the task at hand thereby generating more accurate predictions of anomalies and / or trends in visual representations of telemetry data.

[0052] The model 800 contains one or more encoder blocks 802 coupled to one or more decoder blocks 804. The initial inputs to an encoder block 802 are the input embeddings 806 of an input sequence of a training dataset. In order to retain the order of the tokens in the input embedding 806, positional embeddings 808 are added to the input embedding 806 forming a context tensor 809. The initial inputs to the decoder block 804 are a shifted sequence of the output embeddings 818 from a previous time step to which the positional embeddings 820 are added forming context tensor 819.

[0053] An encoder block 802 consists of at least two layers. The first layer includes a multi -head attention component 810 followed by layer normalization component 812. The second layer includes a feed-forward neural network 814 followed by a layer normalization component 816. The context tensor 809 is input into the multihead attention component 810 of the first encoder block 802 with a residual connection to the layer normalization component 812. The output of the layer normalization component 812 is input to the feed-forward neural network 814 with another residual connection to layer normalization component 816. The output of the encoder block 802 is a set of hidden representations 817. The set of hidden representations 817 is then sent through additional encoder blocks. At the last encoder block, the set of hidden representations 817 is sent to the decoder 804.

[0054] Attention is used to decide which parts of the input embedding are important for each token, especially when decoding long sequences since the encoder is limited to encoding a fixed-size vector. Attention mechanisms gather information about the relevant context of a given token and then encode that context into a vector which represents the token. It is used to identity the relationships between tokens in the long sequence while ignoring other tokens that do not have much bearing on a given prediction.

[0055] The multi-head attention component 810 takes a context tensor 809 and weighs the relevance of each token represented in the context tensor 809 to each other by generating attention weights for each token in the input embedding 806. In one aspect, the attention function is scaled dot-product attention.

[0056] In order to reduce the training time of the neural network transformer model, layer normalization is used between the layers. The layer normalization components 812, 816 normalize the inputs across the features. In an example, the mean and standard deviation is computed across the feature dimensions.

[0057] The feed-forward neural network 814 processes each output encoding separately. The output of the top encoder block is a set of attention vectors K and V 817 which is used by the encoder-decoder multi-head attention layer 826 of the decoder block 804.

[0058] The decoder block 804 predicts each token in the output text one-by-one at each time step conditioned on all previously-generated target tokens. A decoder block 804 consists of three layers. The first layer includes a masked multi-head attention component 822 followed by a layer normalization component 824. The output of the layer normalization component 825 is input into the encoder-decoder multi-head attention component 826 with a residual connection to layer normalization component 828. The second layer includes an encoder-decoder multi-head attention component 826 followed by a layer normalization component 828. The third layer includes a feed-forward neural network 830 followed by a layer normalization component 832. The output of layer normalization component 828 is input into the feed-forward neural network 830 with a residual connection to layer normalization component 832.

[0059] The masked multi-head attention component 822 receives the output embeddings of the previous timestep. The masked multi-head attention component 822 masks the output embeddings from future time steps. The encoder-decoder multi-headattention layer 822 receives queries from the previous decoder layer and the memory keys and values 817 from the output of the encoder block 802. In this manner, the decoder block 804 can attend to every position of the input sequence. The feed-forward neural network 830 processes each output encoding separately. A layer normalization component 824, 828, 832 is used between the layers in order to normalizes the inputs across the features.

[0060] In one example of the model of FIG. 8, the model contains a stack of six encoder blocks and a stack of six decoder blocks which are aggregated into a neural transformer block. However, other numbers of encoder and decoder blocks may be used. The output of each encoder block is passed onto the next encoder block and processed. Each decoder block receives the attention weights computed from the last encoder block. The use of multiple stacked encoder blocks and decoder blocks increases the model’s capacity allowing the model to learn increasing levels of abstraction.

[0061] In an example, a model such as that illustrated in FIG. 8 is trained using training examples comprising text obtained from a corpus of documents having medical knowledge about health conditions which include health conditions affecting speech, neurological disorders such as ASD, dementia, Parkinson’s disease, Alzheimer’s disease. The corpus also has huge numbers of documents about other topics. In examples the corpus of documents comprises the VABS, documentation about how to use VABS, training materials for medical doctors and medical practitioners working with patients having neurological disorders. Training examples are formed from the corpus of documents by taking text from a document in the corpus and masking out part of the text. The text with and without the mask becomes a labelled training example. Millions or more training examples are obtained in this way and are used to train a large language model, such as a model as described with reference to FIG. 8, using supervised learning via backpropagation. The model is then finetuned using reinforcement learning from human feedback or reinforcement learning from artificial intelligence feedback.

[0062] FIG. 9 illustrates various components of an exemplary computing-based device 900 which are implemented as any form of a computing and / or electronic device, and in which communication assessment and / or diagnosis tools are implemented in some examples.

[0063] Computing-based device 900 comprises one or more processors 902 which are microprocessors, controllers or any other suitable type of processors for processingcomputer executable instructions to control the operation of the device in order to do any of: automatically diagnose a medical condition, automatically assess communication ability of a patient, automatically determine a model has convergent validity with a human medical doctor, measure efficacy of a pharmacological treatment, measure efficacy of a treatment being assessed in a clinical trial, determine a dose regime for a patient, control a computing device used in digital intervention for a medical condition, generate a time limited password to a computer game apparatus having an intervention for the medical condition, generate an automated diagnostic tool, compute a numerical score of expressive communication ability of a patient on the VABS clinical scale. In some examples, for example where a system on a chip architecture is used, the processors 902 include one or more fixed function blocks (also referred to as accelerators) which implement a part of the method of any of FIGs. 1, 2, 3, 7, 8 in hardware (rather than software or firmware). Platform software comprising an operating system 908 or any other suitable platform software is provided at the computing-based device to enable application software 910 to be executed on the device. In various examples, an instance of application software is a communication assessment and diagnosis tool 912. A model 914 may be stored at the computing-based device (such as second model 302 of FIG. 3) or the computing-based device may access a model (such as model 104 of FIG. 1) via communication interface 904 such as over the Internet. Data store 916 holds audio signals, text, , scores, diagnoses, prompts and other data.

[0064] The computer executable instructions are provided using any computer- readable media that is accessible by computing based device 900. Computer-readable media includes, for example, computer storage media such as memory 906 and communications media. Computer storage media, such as memory 906, includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or the like. Computer storage media includes, but is not limited to, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), electronic erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that is used to store information for access by acomputing device. In contrast, communication media embody computer readable instructions, data structures, program modules, or the like in a modulated data signal, such as a carrier wave, or other transport mechanism. As defined herein, computer storage media does not include communication media. Therefore, a computer storage medium should not be interpreted to be a propagating signal per se. Although the computer storage media (memory 906) is shown within the computing-based device 900 it will be appreciated that the storage is, in some examples, distributed or located remotely and accessed via a network or other communication link (e.g. using communication interface 904).

[0065] The computing-based device 900 is arranged to output display information to a display device 918 which may be separate from or integral to the computing-based device 900. The display information may provide a graphical user interface such as to display scores, diagnoses, conversations, and other data, such as for use by a medical doctor. The computing-based device 900 also arranged to receive and process input from one or more devices, such as a user input device (e.g. a mouse, keyboard, camera, microphone or other sensors 920). Other sensors 920 are microphones, smart watches, body worn sensors in some cases.

[0066] The term 'computer' or 'computing-based device' is used herein to refer to any device with processing capability such that it executes instructions. Those skilled in the art will realize that such processing capabilities are incorporated into many different devices and therefore the terms 'computer' and 'computing-based device' each include personal computers (PCs), servers, mobile telephones (including smart phones), tablet computers, smart watches, wearable computers, and many other devices.

[0067] The methods described herein are performed, in some examples, by software in machine readable form on a tangible storage medium e.g. in the form of a computer program comprising computer program code means adapted to perform all the operations of one or more of the methods described herein when the program is run on a computer and where the computer program may be embodied on a computer readable medium. The software is suitable for execution on a parallel processor or a serial processor such that the method operations may be carried out in any suitable order, or simultaneously.

[0068] Those skilled in the art will realize that storage devices utilized to store program instructions are optionally distributed across a network. For example, a remotecomputer is able to store an example of the process described as software. A local or terminal computer is able to access the remote computer and download a part or all of the software to run the program. Alternatively, the local computer may download pieces of the software as needed, or execute some software instructions at the local terminal and some at the remote computer (or computer network). Those skilled in the art will also realize that by utilizing conventional techniques known to those skilled in the art that all, or a portion of the software instructions may be carried out by a dedicated circuit, such as a digital signal processor (DSP), programmable logic array, or the like.

[0069] Clinical trial example

[0070] A study explored the potential of using GPT4 to analyze natural conversations between autistic individuals and their caregivers, with the aim of predicting the individuals' VABS communication scores for facilitating medical diagnosis or for symptom measurement. To do so, over 500 conversations were analyzed, from 54 ASD and 18 neuro-typical control (NTC) participants, spanning in age from 5 to 45 years old, in an observational (non-drug) clinical trial. The conversations were manually transcribed by trained professionals, and the resulting transcripts were supplied to GPT4 for analysis and score prediction. This study has demonstrated that GPT4 can serve as a powerful tool for assessing expressive communication in individuals with autism. The strong correlation between GPT4's predictions and actual VABS scores, coupled with the model's high test- retest reliability, suggests that GPT4 can provide objective and sensitive assessments of communicative abilities. Whilst the study used GPT4 there are good theoretical reasons to assume that equivalent or better results will be obtained with other large language models developed in the future or existing already.

[0071] Participant demographics

[0072] Conversation recordings were obtained from 54 ASD participants, including children (5-12 years) adolescents (13-17 years), and adults (18-45 years) as part of an observational clinical trial.

[0073] Each participant was required to have a single study partner with whom they would record their conversations. Participants were stratified by their intelligence quotient IQ (<70 and >69) as measured by the Abbreviated Battery IQ (AB IQ) of the Stanford-Binet intelligence scales fifth edition.

[0074] While 135 participants were initially recruited into the trial, only 54 autistic participants completed the conversation task at least once. Table 1 displays thedemographic information and number of conversations available for the participants included in the final analysis.

[0075] Table 1 : Participant Demographics. The number of participants and the number of conversations (in brackets) for each cohort. Sex (M=male, F=female). Recorded = the total number of participants (conversations) with recorded audio. Transcribed = the number of participants (conversations) with manually transcribed audio. ASD = autism spectrum disorder, NTC = neuro-typical controls, Children = 5-12yr, Adolescents = 13-17yr, and Adults = 18-45yr.

[0076] Inclusion and exclusion criteria for the clinical trialThe inclusion criteria for participants with ASD were as follows:• Diagnosis of ASD based on the Diagnostic and Statistical Manual of Mental Disorders (DSM-5), and the Autism Diagnostic Observation Schedule (ADOS-2).• Children's Yale-Brown Obsessive Compulsive Scale modified for ASD (CY- BOCS-ASD) total score of at least 12.• Clinical Global Impression-Severity (CGI-S) score of at least 4 about participant's current autism severity.• Intelligence quotient (IQ) score of 50 or above as assessed by the Abbreviated Intelligence Quotient (AB IQ) SB 5 scale.• English proficiency compatible with the study measurements as judged by the investigator.• Hearing, vision, and speech compatible with the study measurements as judged by the Investigator.• All medications and treatments were expected to be stable for the duration of the study.

[0077] The diagnostic evaluations were completed at the study site by research staff and supervised by a licensed psychologist.

[0078] The exclusion criteria for all participants were as follows:• Participation in an investigational drug or device study within 4 weeks, or five times the half-life (if it is a drug study) of the investigational molecule (whichever is longer), prior to screening and the participant is expected not to enroll in any other trial during the study.• Co-occurring disease or condition that could interfere with, or treatment which might interfere with, the conduct of the study, or that would, in the opinion of the Investigator, pose an unacceptable risk to the participant in this study.• Unstable or uncontrolled clinically significant psychiatric and / or neurological disorder that may interfere with the objectives of the study.• Participants with known "syndromic" ASD (e.g., Fragile-X syndrome, Angelman Syndrome, Prader-Willi, Rett's syndrome, tuberous sclerosis, Dupl5q syndrome).• History of alcohol misuse and / or illicit drug use during the last 12 months prior to screening.

[0079] The inclusion criteria for the study partners were as follows:• A staff member of the residential home can be the caregiver if this person spends sufficient time with the participant. In the opinion of the Investigator, the caregiver must be able to reliably assess the participant's mental status, activities, and behavior, and report on the participant's adherence and health. This would normally be possible when the caregiver spends a few hours each day with the participant.• A family member living at the participant's home can be the caregiver if the participant returns home every night. When the participant returns home only over the weekend, a family member can only be the caregiver if they have intensive interaction with the participant during the week e.g., via phone calls, calls via Skype, SMS messages, etc. The quality of these interactions between caregiver and participant needs to be assessed for each participant to determine whether they are sufficient.

[0080] Participant demographics are show in table 2 as follows:

[0081] Table 2 shows additional demographic information showing the sex, age, and IQ of each cohort. ASD = autism spectrum disorder. NTC = neurotypical control. In addition, the ethnicity breakdown across all cohorts is as follows: 72.2% white, 7% Asian, 18% black or African American, and 2.7% multiple.

[0082] Audio recordings and transcriptions

[0083] As part of the study, participants and study partners were instructed to record natural conversations once per week while being recorded on a cell phone that was provided at the start of the trial. The study partners were given 6 suggested topics to discuss: school, work, free-time, what they did today, dreams, and things they like or dislike. While the conversations were instructed to take place in a quiet environment, they often included background noise from everyday living activities, as well as speech from other members of the household. They were instructed to be at least 5 minutes in length, but were often shorter in practice. No other instructions were given; therefore the conversations were free-flowing and unstructured in every other regard.

[0084] VABS

[0085] The Vineland Adaptive Behaviour Scale (VABS) is commonly used for assessing various adaptive behaviors in neurodevelopmental disorders, and has been shown to be effective in assessing ASD2. The VABS-II interview was completed for each ASD participant in this study. The communication portion of the VABS contains the subdomains of expressive, receptive and written communication. Expressivecommunication refers to the ability to express wants and needs through verbal and nonverbal communication. It is the ability to put thoughts into words and sentences in a way that makes sense and is grammatically correct. Receptive communication refers to the ability to understand and comprehend spoken language that is heard or read. Written communication was not assessed in this study.

[0086] The raw VABS scores in each sub-domain of VABS are calculated by assigning a score between 0 and 2 for a vast range of questions. Each subdomain has a different number of questions, with the Expressive and Receptive subdomains containing 55 and 20 questions respectively, resulting in theoretical maximum scores of 110 and 40. However, the specific number of questions varies depending on the age and communication abilities of the individual being assessed.

[0087] The model

[0088] The model 104 used in the study was Generative Pre-trained Transformer 4 (GPT4), which is a state-of-the-art large language model developed by OpenAI. It is trained on a diverse range of internet text, so it can perform a wide variety of tasks without task-specific training, and can effectively process and generate human-like text.

[0089] There are two at least two settings when using GPT4. The first is the temperature parameter (range: 0-1), which controls the randomness of the output. A lower temperature means the model will produce more predictable and conservative text, while a higher temperature results in more varied and sometimes less predictable outputs. In order to ensure consistency and reliability in the responses a temperature of 0.1 was used.

[0090] The second setting is the model prompt 106. This refers to the initial context and instructions given to the model to guide its text generation. In the study the model prompt 106 was used to define the style and format of the conversation transcripts, the specific task of GPT4, and the format that the output should have. Specifically, the prompt 106 requested that GPT4 provide an estimate of the expressive and receptive communication abilities of the participant based on the VABS scale. However, due to the variability in the number of questions in the VABS survey (see previous section), and the undefined minimum and maximum scores, the prompt 106 was configured to ask GPT to produce a score between 0 and 100. Therefore, the study was not looking to precisely predict the exact raw score of the VABS, but rather to produce a score that is monotonically correlated with it. An example full prompt used in the study is as follows:

[0091] Example PromptYou will be provided with multiple conversations between two people: a person with autism, and their caregiver.BACKGROUND ON THE FORMA T OF THE CONVERSA TIONS:The person with autism is labelled as [participant], and the caregiver is labelled as [study partner [ However, the study partner might sometimes be labelled as [Fl] or [Ml].There may or may not be other people in the conversation, labelled as [F2] or [M2] etc... any names in the conversation have been redacted for privacy concerns, and replaced with <redacted>.YOUR JOB:Your job is to provide a score between 0 and 100 of both the expressive and receptive communication abilities of the participant based on the Vineland Adaptive Behaviour Scales clinical scale.OUTPUT FORMAT:This is VERY IMPORTANT, the format of the final score should be as follows: Final Expressive Communication Score: XFinal Receptive Communication Score: Y << >> >>

[0092] The study experimented with using two example conversations as training data, one from an autistic individual with a very low VABS expressive communication score, and one from a neurotypical control without a VABS score. These labelled training examples were used in the prompt to guide GPT4 on the style and format of the conversations, and what should be expected in terms of VABS scores. If the study used these example conversations, the following lines were added to the prompt:<< >> >>Here is an example of a conversation that should score very low on both scales: [study parlnei] Hi, how was school today? [participant] Good.And here is an example of a conversation that should score very high on both scales: [study partner I Ok so what should we talk about today? [participant] Oh I don ’t know, how about the weather?

[0093] Participant transcripts are omitted from this document due to privacy concerns, so the above are non-limiting examples of how two conversations start. In practice transcripts of the full conversations and examples were used from both individuals.

[0094] The study did not use data from the same participant for training and testing. The training example from the NTC participant was used (because that individual did not have a VABS score, and therefore the study never tried to predict it), but the study swapped the labelled training example from the participant with a low VABS score with another participant with a similarly low score when appropriate.

[0095] The conversations were input for analysis to GPT4 for three separate iterations. Each request was submitted from scratch, with no memory of previous iterations or participant conversations.

[0096] Linguistic and temporal features

[0097] In addition to comparing GPT4 evaluations with the VABS, the study also compared GPT4 with simple, quantifiable linguistic and temporal features that have been shown to be indicative of communicative competence. These features were the number of words spoken per sentence (WpS), which reflects syntactic complexity and the ability to form coherent, structured expressions. Utterance duration (UD: the length of time that the participant would speak for within their turn) is considered to capture the fluency and temporal aspects of speech, providing insight into the speaker's verbal delivery and pacing. Lastly, the age of acquisition (AoA) for the words used in the communication samples is included, as it serves as a proxy for the developmental aspect of language use, indicating the maturity of the individual's lexicon. These features were chosen for their empirical relevance to language proficiency and their established correlation with expressive communication abilities in the literature. For each feature, the study calculated the mean value across all conversations to result in a single value for each participant. By comparing the performance of GPT-4 against this baseline, it was possible to ascertain the added value of the model's advanced linguistic and contextual processing capabilities.

[0098] Intraclass correlation

[0099] To assess the test-retest reliability of each measure, the study calculated the intraclass correlation coefficient (ICC) using the two-way random, single measures, absolute agreement method (ICC). The study used a Bayesian model using the No-U-Turn Sampler (NUTS) algorithm from PyMClO.

[0100] Results

[0101] Predicting the VABS Expressive Communication Score

[0102] The core findings of the study are now presented, focusing on the performance of the GPT4 model in predicting the VABS expressive and receptive communication scores based on the transcribed conversations. Instead of aiming to precisely predict the raw VABS scores, the prompt bound GPT4’s predictions between 0- 100, in the hopes of finding a monotonic relationship between the GPT4 predictions and the VABS scores from human medical practitioners.

[0103] Figure 10A depicts the relationship between the scores predicted by GPT4 (x-axis) and the actual VABS expressive communication scores (y-axis) determined by human medical practitioners such as medical doctors. For each participant, the study first tested performance by concatenating all of their conversations together within a single prompt and submitting the prompt to the model three separate times. The mean value of the VABS score predicted by the model across the 3 iterations was computed and these mean values are shown in FIG. 10 A.

[0104] FIG 10B is a graph of correlation between the actual VABS expressive communication scores (y-axis) and the predicted scores from GPT4 (x-axis). Each dot is a participant (n=52). Individual conversations were used and taking the median score across all conversations per patient and across 3 iterations. The study also assessed performance on the individual conversations, taking the median value across all conversations and iterations (FIG 10B) per patient. The correlation plots reveal a strong positive linear relationship (Pearson’s r >= 0.65, p<lxl0-5).

[0105] Asking GPT4 for symptom descriptions and conversation summaries

[0106] Rather than just asking GPT4 to provide a single prediction of the VABS scores, the study also experimented with asking for additional information on the conversations, such as summarising them and providing a description of any autism symptoms that are evident. The resulting outputs were always interesting, with GPT4 discussing things like repetitive speech usage, the ability to maintain a conversation, andif the participant’s speech contained incomprehensible phrases, or non-verbal sounds. However, given the inherent natural language style of these outputs, it’s difficult to quantify these results objectively. Unexpectedly, the final predicted scores had a lower correlation with the true VABS scores when using a prompt asking for a description of autism symptoms in addition to a VABS score, versus the simpler prompt without a request for a symptom description (r = 0.59 versus r = 0.65), and a lower ICC value (0.85 [95% CI: 0.79, 0.91] versus 0.95 [95% CI: 0.93, 0.97]).

[0107] The effect of providing training examples in the prompt

[0108] FIG. 11 A is a graph of correlation between actual VABS expressive communication scores (y-axis) and predicted scores from GPT4 (x-axis), where each dot is a participant (n=52) and when training examples were used in the prompt. There were two training examples in the prompt, one from an autistic individual with a very low VABS expressive communication score, and one from a neurotypical control (and who therefore did not have a VABS score). It was found that providing training examples in the prompt provides a slight increase in correlation (r=0.65 vs. 0.6), with the primary effect being that the scores are more dispersed along the 0-100 range (mean+ / -SD: x + / -y, a+ / -b for with and without training examples, respectively).

[0109] FIG. 1 IB is a graph of correlation between actual VABS expressive communication scores (y-axis) and predicted scores from GPT4 (x-axis), where each dot is a participant (n=52) and when using no training examples in the prompt.[001 10] Linguistic and temporal feature comparison[001 1 1 ] In order to compare the performance of GPT4 with simpler features computed from the conversations manually, the study opted to use three linguistic and temporal features that have been shown to correlate with communicative abilities: words per sentence (WpS), utterance duration (UD) and Age of Acquisition (AoA).[001 12] FIG. 12A is a graph of correlation between GPT4 predictions of the VABS expressive communication scores (y-axis) versus Words per Sentence (WpS; x-axis) in the conversations. FIG. 12A illustrates the comparison between the GPT4 predictions of the VABS expressive communication scores (y-axis) and the mean number of Words per Sentence (WpS; x-axis), revealing a strong correlation (Pearson: r=0.62, p<lxl0'5).[001 1 3] FIG. 12B is a graph of correlation between GPT 4 predictions of the VABS expressive communication scores (y-axis) versus Utterance Duration (UD; x-axis) in the conversations. FIG. 12B illustrates the comparison between the GPT4 predictions of theVABS expressive communication scores (y-axis) and the Utterance Duration (UD; x- axis), revealing a moderate correlation (Pearson: r=0.41, p<lxlO'5).[001 14] FIG. 12C is a graph of correlation between GPT 4 predictions of the VABS expressive communication scores (y-axis) versus Age of Acquisition (AoA; x-axis) in the conversations. FIG. 12C illustrates the comparison between the GPT4 predictions of the VABS expressive communication scores (y-axis) and the Age of Acquisition (AoA; x- axis), revealing a strong correlation (Pearson: r=0.66, p<lxl0'5). Table 3 shows a correlation between each feature and GPT4, revealing significant correlations between each of them.[001 1 5] FIG. 13 is a box plot of GPT4 predictions of VABS expressive communication scores (y-axis) for ASD (IQ < 70) participants, ASD (IQ >70) participants and neurotypical control (NTC) participants within each age group comprising Children, Adolescents and Adults (x-axis). FIG. 13 illustrates the comparison of GPT4 predictions of the VABS expressive communication scores (y-axis) between ASD (IQ<70), ASD (IQ >70) and NTC participants between different age groups (x-axis) comprising Children, Adolescents and Adults, revealing that the GPT4 predictions of VABS expressive communication scores for NTC participants generally trended higher than those for the ASD groups. FIG. 13 also illustrates within the ASD cohorts that participants with IQ >70 received GPT4 predictions of the VABS expressive communication scores that were more closely aligned with the NTC group, whereas participants with IQ<70 had a broader range of predicted scores which were generally lower.[001 16] To investigate how much unique information is provided by each feature a partial correlation analysis was performed. When doing so (Table 3), GPT4 provides stronger correlations than any other combination of feature covariate combinations demonstrating that using a large language model gives better performance than using manual assessment of the individual features. Table 3 is given here. In table 3 the letters Na denotes not a number and indicates that no meaningful prediction was obtained.Cross-feature correlationTable 2Partial CorrelationCovariateTable 3[001 1 7] Intraclass correlation[001 1 8] In order to assess the test-retest reliability of GPT4, the study assessed the intraclass correlation coefficient (ICC) when using both the concatenated and individual conversations. When using the concatenated conversations, this resulted in 52x3 observations (participants x iterations) and an ICC score of 0.95 [95% CI: 0.93, 0.97], Thus a remarkable consistency in GPT4’s predictions is shown.[001 1 9] When using the individual conversations, the study is not only testing the consistency of GPT4, but also the consistency of individual conversations to reveal sufficient information about the communication abilities of each participant. Therefore, this resulted in 297x3 observations (conversations x iterations) and an ICC of 0.78 [95% CI: 0.71, 0.85], FIG. 13A is a graph of test-retest reliability when using a human doctor to determine a VABS score. FIG. 13B is a graph of test-retest reliability when using GPT-4 to determine a VABS score.[001 20] The application of GPT4 to predict Vineland Adaptive Behaviour Scales (VABS) communication scores from natural conversations signifies a substantial advancement in the assessment of communicative abilities in individuals with autism. The results of this study reveal a robust correlation between the expressive communication scores predicted by GPT4 and the actual VABS scores. This outcome suggests that GPT4 can capture nuanced aspects of expressive communication, surpassing the capabilities of traditional linguistic and temporal features.[001 21 ] The high intraclass correlation coefficient (ICC) observed in this study for GPT4's predictions, particularly when analysing concatenated conversations, underscores the model's consistency and reliability. While the ICC for individual conversations was slightly lower, it remained substantial, which is impressive given the brevity and unstructured nature of these exchanges. This finding suggests that GPT4, and by extrapolation other large language models more generally, can discern expressive communication abilities from even short snippets of conversation, a testament to large language model sensitivity and depth of language understanding.[001 22] It is noteworthy that GPT4 not only correlated with the expressive communication scores but also provided additional insights beyond the simple linguistic and temporal features that were assessed. The partial correlation analysis indicates that GPT4 captures a broader spectrum of communicative competence, which may includefactors such as pragmatic language skills, conversational coherence, and other subtle linguistic cues that are not easily quantifiable.

[0123] These findings have significant implications for clinical practice and research. In clinical settings, the use of large language models such as GPT4 augments traditional assessments, providing rapid and objective evaluations of communication abilities that are less susceptible to the biases and limitations inherent in human- administered tests. In the context of clinical trials, the methodology presented here enhances the sensitivity of outcome measures, allowing for the detection of subtle treatment effects that might otherwise go unnoticed.

[0124] Any range or device value given herein may be extended or altered without losing the effect sought, as will be apparent to the skilled person.

[0125] Although the subj ect matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0126] It will be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments. The embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benefits and advantages. It will further be understood that reference to 'an' item refers to one or more of those items.

[0127] The operations of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate. Additionally, individual blocks may be deleted from any of the methods without departing from the scope of the subject matter described herein. Aspects of any of the examples described above may be combined with aspects of any of the other examples described to form further examples without losing the effect sought.

[0128] The term 'comprising' is used herein to mean including the method blocks or elements identified, but that such blocks or elements do not comprise an exclusive list and a method or apparatus may contain additional blocks or elements.

[0129] It will be understood that the above description is given by way of example only and that various modifications may be made by those skilled in the art. The above specification, examples and data provide a complete description of the structure and use ofexemplary embodiments. Although various embodiments have been described above with a certain degree of particularity, or with reference to one or more individual embodiments, those skilled in the art could make numerous alterations to the disclosed embodiments without departing from the scope of this specification.

Claims

CLAIMS1. A computer-implemented method of assessing speech ability of a patient in order to diagnose a health condition, the method comprising: receiving an audio signal of a conversation between a patient and a study partner; inputting the audio signal, or data derived from the audio signal, to a model, the model being a large language model; inputting a prompt to the model, the prompt requesting the model to generate an output comprising a numerical score in a specified range, and requesting the numerical score to quantify a health condition of the patient using a specified clinical scale, taking into account the audio signal or data derived from the audio signal; receiving output from the model comprising the numerical score.

2. The method of claim 1 wherein the clinical scale is Vineland Adaptive Behaviour Scale.

3. The method of any preceding claim wherein the prompt requests the model to diagnose the health condition and the output from the model also comprises a diagnosis of whether or not the patient has the health condition.

4. The method of any preceding claim wherein the prompt comprises any of: a direction to the model to respond as: an expert and experienced consultant psychiatrist that specializes in autism spectrum disorder; or an expert and experienced medical doctor that specializes in dementia; no direction to describe autism symptoms in the conversation; a plurality of concatenated conversations between the patient and the study partner.

5. The method of any preceding claim comprising, prior to inputting the prompt to the model, determining that the model has convergent validity with a human doctor expert in diagnosing the medical condition, where convergent validity occurs where a linear relationship between communication ability scores on the clinical scale from the model for a plurality of patients and communication ability scores on the clinical scale for the same patients from the human doctor has a coefficient of correlation, r, with a modulus above 5 and a statistical significance p-value of 0.05 or lower.

6. The method of any preceding claim wherein receiving the audio signal comprises using a microphone in a portable or wearable computing device to capture the audio signal, and wherein the method comprises defining a topic of the conversation via a user interface of the portable or wearable computing device.

7. The method of any preceding claim wherein the model is any of generative pre-trained transformer 4, GPT4, versions 0.2.5-0.2.7, the clinical scale is Vineland Adaptive Behaviour Scale, and the health condition is Autistic Spectrum Disorder.

8. The method of any preceding claim comprising providing the output to a human medical doctor at a user interface, as a candidate communication ability score, together with medical records of the patient, and in response to the medical doctor selecting the candidate communication ability score, storing the candidate communication ability score into a record for the patient.

9. The method of any preceding claim wherein the patient is a member of a clinical trial of a pharmaceutical or other treatment for the medical condition, and wherein the method is repeated during the clinical trial in order to assess whether communication ability of the patient is influenced by the treatment.

10. The method of any of claims 1 to 9 wherein the patient is being treated with a pharmaceutical treatment for the medical condition, and wherein the method is repeated during the pharmaceutical treatment and a dose regime of the pharmaceutical treatment is adjusted according to changes in the numerical score.

11. The method of any preceding claim comprising generating a time limited password to a computer game apparatus, having an intervention for the medical condition, in dependence on the numerical score.

12. The method of any preceding claim comprising storing, as a training example, the numerical score and inputs to the model, the inputs to the model being the audio signal or data derived from the audio signal; and repeating the method to store more training examples.

13. The method of any of claims 3 to 12 comprising storing, as a training example, the diagnosis and inputs to the model, the inputs to the model being the audiosignal or data derived from the audio signal; and repeating the method to store more training examples.

14. The method of any preceding claim wherein the method further comprises identifying participants to be included in a clinical trial using a selection criteria of participants that fall within a specified range of the numerical score.

15. An apparatus for assessing communication ability of a patient in order to diagnose Autistic Spectrum Disorder, the apparatus comprising: a processor; a memory storing instructions that, when executed by the processor, perform operations comprising: receiving an audio signal of a conversation between a patient and a study partner; inputting the audio signal, or data derived from the audio signal, to a model, the model being any of GPT-4 versions 0.2.5-0.2.7; inputting a prompt to the model, the prompt requesting the model to generate an output comprising: a numerical score in a specified range, the numerical score being of an expressive communication ability of the patient on the Vineland Adaptive Behaviour Scale; receiving output from the model comprising the numerical score; and presenting the numerical score to a medical practitioner for use in diagnosis of Autistic Spectrum Disorder.

Citation Information

Cited By

  • Multi-scene self-adaption-based pharmacist clinical ability assessment method and system

    CN122089170A