Methods and systems for using digital voice biomarkers to monitor symptoms of schizophrenia and schizoaffective disorders
A machine learning-based system analyzes voice recordings to determine digital biomarkers for schizophrenia, addressing the limitations of current monitoring methods by providing continuous and objective symptom assessment, enhancing treatment efficacy and adherence.
Patent Information
- Application Number
- PCT/EP2025/061973
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-20
- Filing Date
- 2025-04-30
- Publication Date
- 2025-11-06
AI Technical Summary
Current methods for monitoring schizophrenia symptoms are time-consuming, subjective, and lack continuous, objective monitoring between clinical visits, leading to potential relapses due to infrequent assessments and poor adherence.
A system using machine learning models to analyze acoustic data from voice recordings to determine digital biomarkers for schizophrenia severity, enabling continuous and objective monitoring through portable devices, allowing for adjustments in medication based on symptom severity.
Provides continuous, objective monitoring of schizophrenia symptoms, improving treatment efficacy and adherence by using voice-based digital biomarkers that correlate with clinical assessments, facilitating timely adjustments in medication.
Smart Images

Figure EP2025061973_06112025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR USING DIGITAL VOICE BIOMARKERS TOMONITOR SYMPTOMS OF SCHIZOPHRENIA AND SCHIZOAFFECTIVEDISORDERSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of Provisional U.S. Patent Application No. 63 / 640,559, filed April 30, 2024, Provisional U.S. Patent Application No. 63 / 690,624, filed September 4, 2024, and Provisional U.S. Patent Application No. 63 / 697,069, filed September 20, 2024, the disclosures of which are incorporated herein by reference in their entirety.BACKGROUND
[0002] Schizophrenia is a chronic progressive mental disorder. Symptoms of schizophrenia include cognitive and behavioral disturbances. Schizophrenia is a chronic psychiatric disorder characterized by an array of cognitive and behavioral disturbances. Treatment of schizophrenia often involves meticulous management including frequent visits to the clinic to monitor patient stability and prevent relapses. There is consensus among clinicians, caregivers, and patients that disease management is a key challenge in providing optimal treatment to patients with schizophrenia. Infrequent clinical visits and lack of adherence, combined with the burden and poor reliability of traditional monitoring scales, all lead to deterioration and relapse. As such, there is a need for continuous and objective monitoring, for example that spans the time between clinic visits and / or provides standardized results.SUMMARY
[0003] Systems and methods for assessing schizophrenia, for example severity of schizophrenia, are disclosed herein. The systems and methods may use a machine learning model, for example a trained machine learning model. The model may be an artificial intelligence (AI) / machine learning (ML) model. Audio for a patient may be received, for example by a receiver. The receiver may be a microphone. The audio may be recorded (e.g., in an audio recording), for example in a memory (e.g., of a computing device). The audio recording may be processed, for example at the computing device. The audio recording may be processedto determine acoustic data related to the audio recording. The machine learning model, for example the trained machine learning model, may be (e.g., stored) in the computing device. The system may determine, for example using the trained machine learning model, a digital biomarker. The digital biomarker may correspond to schizophrenia severity of the patient using the acoustic data. The digital biomarker may correspond to a positive and negative syndrome scale (PANSS) score. In some examples, the digital biomarker may be a PANSS score. For example, the digital biomarker may include an indication of one or more symptoms and / or the severity of schizophrenia. The receiver may be comprised in the computing device.
[0004] The audio recording may include a first audio recording. The acoustic data may include a first set of acoustic data. The digital biomarker may include a first digital biomarker. A medication associated with schizophrenia may be administered to a patient, for example at a first time corresponding to the first audio recording. A second audio recording for a patient may be received, for example by the receiver. The audio recording (e.g., second audio recording) may be processed to determine a second set of acoustic data related to the audio recording. A second digital biomarker may be determined using the acoustic data, for example using the (e.g., trained) machine learning model. The second digital biomarker may correspond to schizophrenia severity of the patient. An amount of medication, for example a daily amount, may be determined and / or adjusted. The amount of medication may be determined and / or adjusted based on the second digital biomarker, for example based on a comparison of the second digital biomarker to the first digital biomarker. For example, the amount of medication may be determined and / or adjusted based on the second digital biomarker being one or more of equal to, greater than, or less than the first digital biomarker by a predetermined range.
[0005] The audio recording may be processed into discrete segments, for example discrete segments of a predetermined duration. Additionally, or alternatively, the system may determine the acoustic data corresponding to each of the discrete segments of the audio recording. The predetermined duration may be about 40 ms.
[0006] A model may be trained to produce the trained machine learning model, for example using a plurality of processed audio recordings. The plurality of audio recordings may be obtained from a plurality of patients. The patients may have confirmed schizophrenia diagnoses and / or PANSS scores. Additionally, or alternatively, the model may be trained using PANSS scores, for example corresponding to the audio recording of each of the plurality of patients.
[0007] The acoustic data may include one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, a chroma feature,. A patient may be instructed to produce free-form speech, for example in response to a prompt. The audio recording may include the free-form speech and / or have a predetermined duration. The predetermined duration may be about 1 minute. The prompt may include one or more of text and / or an image.
[0008] In some examples, the medication may include one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone.
[0009] In some examples, multiple machine learning models may be used to assess the severity of schizophrenia in a patient. For instance, a method that uses multiple machine learning models may include receiving an audio recording for a patient, segmenting the audio recording into a plurality of discrete segments of a predetermined duration, and determining acoustic data corresponding to each of the discrete segments of the audio recording. The predetermined duration is in the range of approximately 40-60 ms. The method may include determining, using a first trained machine learning model, a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, where each segment level score is associated with a discrete segment of the plurality of discrete segments. The method may include determining, using a second trained machine learning model, a patient session score based on the plurality of segment level scores, where the patient session score corresponds to schizophrenia severity of the patient based on the acoustic data.
[0010] In some examples, the segment level score may correspond to schizophrenia severity of the patient at the discrete segment level, and / or the patient session score may correspond to a PANSS score. The first trained machine learning model may include a different model than the second trained machine learning model. For example, the first trained machine learning model may include an XGBoost model, and the second trained machine learning model may include a K-means clustering model.
[0011] The method may include adjusting a daily amount of the medication or a type of the medication for the patient based on a comparison between multiple patient session scores for a patient.
[0012] The method may include training a model(s) to produce the first trained machine learning model and the second trained machine learning model using a plurality of processedaudio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
[0013] The acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
[0014] The method may include instructing the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration. The prompt may be text and / or an image, and the predetermined duration may be about 1 minute.
[0015] Systems and methods for assessing schizophrenia, for example severity of schizophrenia, are disclosed herein. For example, a method for assessing the severity of schizophrenia in a patient may include receiving an audio recording for a patient, segmenting the audio recording into a plurality of discrete segments of a predetermined duration, determining acoustic data corresponding to each of the discrete segments of the audio recording, determining, using a first trained machine learning model, a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, wherein each segment level score is associated with a discrete segment of the plurality of discrete segments, and determining, using a second trained machine learning model, a patient session score based on the plurality of segment level scores, wherein the patient session score corresponds to schizophrenia severity of the patient.
[0016] The segment level score may correspond to schizophrenia severity of the patient at the discrete segment level. The patient session score may correspond to a PANSS score. The predetermined duration is in a range of about 40 ms to about 60 ms.
[0017] The first trained machine learning model may include a different model than the second trained machine learning model. The first trained machine learning model may include boosting model, and the second trained machine learning model may include a clustering model. The first trained machine learning model may include an XGBoost model, and the second trained machine learning model may include a K-means clustering model.
[0018] The audio recording may be a first audio recording, acoustic data may be a first set of acoustic data, and the patient session score may be a first patient session score. The method may include administering to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording, receiving a second audio recording for thepatient, segmenting the second audio recording into a second plurality of discrete segments of a predetermined duration, determining second acoustic data corresponding to each of the discrete segments of the second audio recording, determining, using the first trained machine learning model, a second plurality of segment level scores based on the second acoustic data of the second plurality of discrete segments, and determining, using a second trained machine learning model, a second patient session score based on the second plurality of segment level scores, wherein the second patient session score corresponds to schizophrenia severity of the patient based on the second acoustic data. The method may also include adjusting a daily amount of the medication or a type of the medication for the patient based on a comparison between the second patient session score and the first patient session score.
[0019] The medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone
[0020] The method may include training a model to produce the first trained machine learning model and the second trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
[0021] The acoustic data may include one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
[0022] The method may include instructing the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration. The prompt may be text and / or an image. The predetermined duration may be about 1 minute.
[0023] Some examples may define a method for assessing the severity of schizophrenia in a patient using a trained machine learning model that includes receiving an audio recording for a patient, processing the audio recording to determine acoustic data related to the audio recording, and determining, using the trained machine learning model, a digital biomarker corresponding to schizophrenia severity of the patient using the acoustic data.
[0024] The audio recording may be a first audio recording, acoustic data may be a first set of acoustic data, and the digital biomarker may be a first digital biomarker. In such examples, the method may also include administering to the patient a medication associated withschizophrenia at a first time corresponding to the first audio recording, receiving a second audio recording for the patient, processing the second audio recording to determine a second set of acoustic data related to the second audio recording, determining, using the trained machine learning model, a second digital biomarker corresponding to schizophrenia severity of the patient using the second set of acoustic data, and adjusting a daily amount of the medication or a type of the medication for the patient based on a comparison between the second digital biomarker and the first digital biomarker.
[0025] The medication may include one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone.
[0026] The method may include segmenting the audio recording into discrete segments of a predetermined duration, and determining the acoustic data corresponding to each of the discrete segments of the audio recording. The predetermined duration is about 40 ms.
[0027] The digital biomarker may correspond to a PANSS score.
[0028] The method may include training a model to produce the trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
[0029] The acoustic data may include one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
[0030] The method may include instructing the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration. The prompt may be text and / or an image. The predetermined duration may be about 1 minute.
[0031] Although described above in context of methods, the methods may be performed by one or more processors and / or a computer readable storage medium may include instructions stored thereon that, when executed by one or more processors (e.g., of one or more servers), cause the one or more processors to perform any of the methods described above.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG. l is a diagram of an example system for collecting data from various sources and generating a machine learning model capable of determining the progression of schizophrenia.
[0033] FIG. 2 is a block diagram that illustrates an example of a computing device of the example system shown in FIG. 1.
[0034] FIG. 3 A is a schematic illustration of an example system environment that may implement an AI / ML model.
[0035] FIG. 3B illustrates an example of a neural network (NN).
[0036] FIG. 3C is a schematic illustration of an example system environment for training and implementing an AI / ML model that comprises an NN.
[0037] FIG. 4 is an example of sensor data.
[0038] FIGs. 5A-5C are diagrams that illustrate an example of a model that is developed on verbal fluency and free speech to correlate the PANSS scores with acoustic features from short segments and with complete patient visits by aggregating the predictions from the short segments.
[0039] FIG. 6 is a diagram of an example table that provides a summary of the demographics and characteristics of participants of the first and second studies.
[0040] FIG. 7A is a diagram illustrating a correlation between the predicted and the actual PANSS scores for the segment level score and new unseen patient visit score of the first study using one or more of the models described herein.
[0041] FIG. 7B is a diagram illustrating a correlation between the predicted and the actual PANSS scores for the segment level score and new unseen patient visit score of the second study using one or more of the models described herein.
[0042] FIG. 8 is a flow diagram that illustrates a training process for training a machine learning model to determine the progression of schizophrenia of a patient with schizophrenia.
[0043] FIG. 9 is a flow diagram that illustrates an example process for using a trained machine learning model to determine the progression of schizophrenia of a patient with schizophrenia.DETAILED DESCRIPTION
[0044] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, systems and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatus, systems and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the Figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the figures to indicate the same or similar parts.
[0045] Currently, Positive and Negative Syndrome Scale (PANSS) scores are the primary standard for assessing treatment efficacy in patients with schizophrenia. However, obtaining PANSS scores can be a time-consuming, subjective exercise for clinicians requiring in-person patient assessments. As a result, there is a need for monitoring disease progression and / or symptoms that, for example, spans the time between clinic visits and / or provides accurate and objective results. Generally speaking, digital biomarkers may be an objective, quantifiable, physiological, and / or behavioral measure that are collected by means of digital devices that are portable, wearable, implantable, or digestible. A voice-based digital biomarker tool may be used between clinical visits and / or may provide continuous monitoring. Voice digital biomarkers may provide a “beyond the pill” solution to aid in (e.g., continuously) monitoring symptoms’ severity, for example, to help determine treatment efficacy and / or adherence. An acoustic features-based model may be used for assessing schizophrenia symptom severity, for example, considering both patient usability and / or clinical validity. Systems and methods as herein may provide improved usability and / or clinical validity. For instance, systems and methods described herein may train and / or use one or more AL / ML models that are configured to determine a digital biomarker corresponding to a schizophrenia severity of a patient based on acoustic data (e.g., acoustic features of the digital biomarker).
[0046] FIG. 1 is a diagram of an example system 100 for collecting data from various data sources (e.g., sensors and / or survey data) and generating (e.g., training) a machine learning model capable of determining the progression of schizophrenia of a patient (e.g., a patient or an individual). The patient may be a patient who has received a positive diagnosis of schizophrenia. The patient may be on a regimen of medication for schizophrenia, for example, such as those medications described herein. The system 100 may include a computing device110 and one or more remote client devices 130a- 130c.
[0047] As described in more detail herein, the computing device 110 may generate e.g., train) and / or use a machine learning model capable of determining a digital biomarker based on the data received from one or more of the remote client devices 130a- 130c. For example, the computing device 110 may be configured to generate a model based on an acoustic featurebased voice digital biomarker measure in schizophrenia. In one embodiment, the digital biomarkers are indicative of the progression of schizophrenia in a patient. As such, the computing device 110 may train the machine learning model to allow the model to determine progression of schizophrenia on a patient with schizophrenia, for example, based on the data received from the remote client devices 130a-130c.
[0048] The remote client devices 130a- 130c may be configured to capture audio recordings (e.g., voice recordings) of the patient. For instance, the remote client devices 130a-130c may be configured to prompt the patient to perform a task, such as a verbal fluency (e.g., 2 minutes) task and / or a free speech (e.g., 1 minute) task, that results in the patient speaking and the remote client devices 130a- 130c recording the patient’s voice. In some examples, the remote client devices 130a-130c include mobile applications stored in memory that are configured to prompt the patient and record the audio recordings.
[0049] The computing device 110 may be configured to receive the audio recordings from the remote client devices 130a- 130c and determine one or more voice digital biomarkers based on the audio recordings. Alternatively or additionally, the remote client devices 130a-130c may be configured to determine one or more voice digital biomarkers based on the audio recordings. The computing device 110 may also receive PANSS scores (e.g., clinician provided PANSS scores) for training the present model. The computing device 110 may be configured to determine or identify the voice biomarkers based on the audio recordings and the PANSS scores. For example, the computing device 110 may be configured to determine and / or asses any correlations between the voice-based predicted PANSS components and the actual PANSS scores from qualified raters to determine the voice biomarkers.
[0050] Further, in some examples, such as those described in more detail herein, the computing device 110 may be configured to segment an audio recording of a patient into discrete segments of a predetermined duration, and determine acoustic data that corresponds to each of the discrete segments of the audio recording. The predetermined duration can be in the range of roughly 40 - 60 ms. The computing device 110 (e.g., a machine learning modeltrained thereon) may determine a segment level score for each of the plurality of discrete segments based on the acoustic data. As such, the computing device 110 may determine a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, where each segment level score is associated with a discrete segment of the plurality of discrete segments. In some examples, the machine learning model that is configured to determine the segment level score at the segment level may include a boosting model, such as an XGBoost or a gradient boosting model.
[0051] The segment level score may be an example of a digital biomarker. The segment level score may correspond to schizophrenia severity of the patient at the discrete segment level (e.g., based on acoustic data for a discrete segment of the audio data). In some examples, the segment level score may correspond to a PANSS score (e.g., for the discrete segment). The segment level score could include an indication of a probability of relapse and / or a binary indicator. The segment level score may be for diagnostic purposes. In some examples, the segment level score may be a PANSS score (e.g., for the discrete segment).
[0052] Finally, the computing device 110 (e.g., a machine learning model trained thereon) may determine a patient session score based on the segment level scores associated with the plurality of discrete segments of the audio recording associated with a patient. The patient session score may be an example of a digital biomarker. The patient session score may correspond to schizophrenia severity of the patient based on the audio recording (e.g., based on all of the segments of the audio recording of the patient). In some examples, the patient session score may correspond to a PANSS score. The patient session score could include an indication of a probability of relapse and / or a binary indicator. The patient session score may be for diagnostic purposes. In some examples, the patient session score may be a PANSS score.
[0053] In some examples, the machine learning model that is configured to determine the patient session score may be different than the machine learning model that is configured to determine the segment level scores associated with each of the plurality of discrete segments. For example, the machine learning model that is configured to determine the patient session score may include a clustering method, such as a K-means clustering method.
[0054] The computing device 110 may comprise a server or combination of servers that may be configured to receive data from various sources, such as the remote client devices 130a- 130c. The computing device 110 may comprise an operational system or subsystem and ananalytical system or subsystem. The computing device 110 may include a processor (e.g., one or more processors), and the processor may store information in and / or retrieve information from memory of the computing device 110. The memory may include a non-removable memory and / or a removable memory. The non-removable memory may include randomaccess memory (RAM), read-only memory (ROM), a hard disk, or any other type of nonremovable memory storage. The removable memory may include a subscriber identity module (SIM) card, a memory stick, a memory card, or any other type of removable memory. The memory may be local memory or remote memory external to the computing system 110. The memory may store instructions which are executable by the processor. Different information may be stored in different locations in the memory.
[0055] The memory may comprise a computer-readable storage media or machine-readable storage media that maintains computer-executable instructions for performing one or more as described herein. For example, the memory may comprise computer-executable instructions or machine-readable instructions that include one or more portions of the procedures described herein. The processor of the computing device 110 may access the instructions from memory for being executed to cause the processor of the computing device 110 to operate as described herein. The memory may comprise computer-executable instructions for executing configuration software. For example, the computer-executable instructions may be executed to perform, in part and / or in their entirety, one or more procedures as described herein. Further, the memory may have stored thereon one or more settings and / or control parameters associated with the computing device 110.
[0056] The computing device 110 and one or more remote client devices 130a-130c may be connected via a network, such as a public and / or private network 140 (e.g., any combination of the Internet, a cloud network, and / or the like). The network 140 may comprise any suitable network over which computing device 110 and one or more remote client devices 130a- 130c may communicate. The network 140 may include a wired and / or wireless communication network. Example wireless communication networks may be comprised of one or more types of RF communication signals using one or more wireless communication protocols, such as a cellular communication protocol, a WIFI communication protocol, and / or another wireless communication protocol.
[0057] FIG. 2 is a simplified block diagram of an example device 200. The device 200 may be an example of a computing device (e.g., one or more servers), such as the computing deviceexample of a client device, such as the client devices 130a-130c of the system 100 of FIG. 1. In such instances, the device 200 may include a personal computer, such as a laptop or desktop computer, a tablet device, a cellular phone or smartphone, a server, or another type of client device. The device 200 may be configured to develop, train, or use one or more machine learning models described herein. As shown by FIG. 2, the device 200 may comprise a processor 202, a memory 204, a communication device 206, a display 208, one or more input devices 210, and / or one or more output devices 212. It should be appreciated that the device 200 may include fewer or more components than those shown in FIG. 2.
[0058] The processor 202 may include one or more general purpose processors, special purpose processors, conventional processors, digital signal processors (DSPs), microprocessors, integrated circuits, a programmable logic device (PLD), application specific integrated circuits (ASICs), or the like. The processor 202 may perform signal coding, data processing, image processing, power control, input / output processing, and / or any other functionality that enables the device 200 to perform as described herein.
[0059] The processor 202 may store information in and / or retrieve information from the memory 204. The memory 204 may include a non-removable memory and / or a removable memory. The non-removable memory may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of non-removable memory storage. The removable memory may include a subscriber identity module (SIM) card, a memory stick, a memory card, or any other type of removable memory. The memory may be local memory or remote memory external to the device 200. The memory 204 may store instructions which are executable by the processor 202. Different information may be stored in different locations in the memory 204.
[0060] The memory 204 may comprise a computer-readable storage media or machine- readable storage media that maintains computer-executable instructions for performing one or more as described herein. For example, the memory 204 may comprise computer-executable instructions or machine-readable instructions that include one or more portions of the procedures described herein. The processor 202 of the device 200 may access the instructions from memory for being executed to cause the processor 202 of the device 200 to operate as described herein. The memory 204 may comprise computer-executable instructions for executing configuration software. For example, the computer-executable instructions may beexecuted to perform, in part and / or in their entirety, one or more procedures as described herein. Further, the memory 204 may have stored thereon one or more settings and / or control parameters associated with the device 200.
[0061] The processor 202 may communicate with other devices via the communication device 206. The communication device 206 may transmit and / or receive information over a network (e.g., the network 140), which may include one or more other devices. The communication device 206 may perform wireless and / or wired communications. The communication device 206 may include a receiver, transmitter, transceiver, or other device capable of performing wireless communications via an antenna. The communication device 206 may be capable of communicating via one or more protocols, such as a cellular communication protocol, a Wi-Fi communication protocol, Bluetooth®, a near field communication (NFC) protocol, an internet protocol, another proprietary protocol, or any other radio frequency (RF) or communications protocol. The device 200 may include one or more communication devices 206.
[0062] The processor 202 may be in communication with a display 208 for providing information to a patient. The information may be provided via a patient interface on the display 208. The information may be provided as an image generated on the display 208. The display 208 and the processor 202 may be in two-way communication, as the display 208 may include a touch-screen device capable of receiving information from a patient and providing such information to the processor 202. The processor 202 may be configured to generate, on the display 208, an indication of one or more notifications described herein, such as an indication of the progression of schizophrenia of the patient, etc.
[0063] The processor 202 may be in communication with input devices 210 and / or output devices 212. The input devices 210 may include a camera, a microphone, a keyboard or other buttons or keys, a mouse, and / or other types of input devices for sending information to the processor 202. For example, the microphone may be used to capture voice data, such as one or more voice recordings, as described herein. The display 208 may be a type of input device, as the display 208 may include touch-screen sensor capable of sending information to the processor 202, such as the graphical user interfaces (GUIs) described herein.. The output devices 212 may include speakers, indicator lights, or other output devices capable of receiving signals from the processor 202 and providing output from the device 200. The display 208 may be a type of output device, as the display 208 may provide images or other visual display of information received from the processor 202.
[0064] Although not illustrated, the device 200 may include a power supply. In some examples, the power supply may include one or more batteries. In some examples, the power supply may include an AC to DC power converter. The power supply may be configured to power the processor 202 and the other low voltage circuitry of the device 200.
[0065] Although not illustrated, the device 200 may include a GPS circuit (e.g., in instances where the device 200 is a client device). The processor 202 may be in communication with the GPS circuit for receiving geospatial information. The processor 202 may be capable of determining the GPS coordinates of the device 200 based on the geospatial information received from the GPS circuit. The geospatial information may be communicated to one or more other communication devices to identify the location of the device 200.
[0066] Systems, methods, and / or apparatus described herein may train and / or implement an artificial intelligence (Al) and / or machine learning (ML) model. For example, one or more devices in the system 100 and / or device 200 may train and / or implement an AI / ML model, such as those described herein.
[0067] FIG. 3A is a schematic illustration of an example system environment 301 that may implement an AI / ML 309 model. The AI / ML 309 model may include model data and one or more algorithms and / or functions configured to learn from input data 307 that is received to train the AI / ML 309 and / or generate an output 315. The input data 307 may be input in one or more formats, such as an image format, an audio format (e.g., spectrogram or other audio format), a tensor format (e.g., including single-dimensional or multi-dimensional arrays), and / or another data type capable of being input into the AI / ML 309 algorithms. The input data 307 may be the result of pre-processing 305 that may be performed on raw data 303, or the input data 307 may include the raw data 303 itself. The raw data 303 may include image data, text data, audio data, or another sequence of information, such as a sequence of network information related to a communication network, and / or other types of data. The pre-processing 305 may include format changes or other types of processing in order to generate input data 307 in a format for being input into the AI / ML 309 algorithms. For example, image data (e.g., including video data) and / or audio data may be raw data 303 that may be pre-processed during pre-processing 305 to generate the input data 307 in a format configured to be received by the AI / ML 309 algorithm. In a particular embodiment, raw data 303 in the form of audio data may be pre-processed into input data 307 in the form of acoustic data. The output 315 may be generated by the AI / ML 309 algorithm in one or more formats, such as a tensor, a text format (e.g., a word, sentence, or othersequence of text), a numerical format (e.g., a prediction), an audio format, an image format (e.g., including video format), another data sequence format, or / another output format.
[0068] AI / ML may be implemented as described herein using software and / or hardware. The AI / ML may be stored as computer-executable instructions on computer-readable media accessible by one or more processors for performing as described herein. Example AI / ML environments and / or libraries include TENSORFLOW, TORCH, PYTORCH, MATLAB, GOOGLE CLOUD Al and AUTOML, AMAZON SAGEMAKER, AZURE MACHINE LEARNING STUDIO, and / or ORACLE MACHINE LEARNING.
[0069] The AI / ML 309 may include one or more algorithms configured for supervised learning. Supervised learning may be implemented utilizing AI / ML 309 algorithms that are trained during a training process to determine a predictive model using known outcomes. The AI / ML 309 algorithms may be characterized by parameters and / or hyperparameters that may be trained during the training process. The parameters may include values derived during the training process. The parameters may include weights, coefficients, and / or biases. The AI / ML 309 may also include hyperparameters. The hyperparameters may include values used to control the learning process. The hyperparameters may include a learning rate, a number of epochs, a batch size, a number of layers, a number of nodes in each layer, a number of kernels (e.g., CNNs), a size of stride (e.g., CNNs), a size of kernels in a pooling layer (e.g., CNNs), and / or other hyperparameters. Some may use certain parameters and hyperparameters interchangeably.
[0070] The AI / ML 309 may be trained during supervised learning by inputting training data to the AI / ML 209 algorithm and adjusting the parameters and / or hyperparameters toward a known target output 315 while minimizing a loss or error in the output 315 generated by the AI / ML 309 algorithm. The raw data 303 may include or be separated into training data, validation data, and / or test data for training, validation, and / or testing, respectively, the AI / ML 309 algorithms during supervised learning. The training data, validation data, and / or test data may be pre- processed from the raw data 303 for being input into the AI / ML 309 algorithm. During supervised learning, the training data may be labeled prior to being input into the AI / ML 309. For example, in an embodiment according to the present disclosure, the AI / ML 309 may be provided with labels in the form of clinician-provided PANSS scores corresponding to each audio recording for each patient. The training data may be labeled to teach the AI / ML 309 algorithm to learn from the labeled data and to test the accuracy of the AI / ML 309 for being implemented on unlabeled input data 307 during product! on / implementati on of the AI / ML 309 algorithms, or similar AI / ML 309 algorithms utilizing similar parameters and / orhyperparameters. The training data may be used to fit the parameters of the AI / ML 309 model using optimization functions, such as a loss or error function. Often the training data includes pairs of input data 307 and a corresponding target output 315 to which the parameters may be trained to generate (e.g., within a threshold loss or error). The trained or fitted AI / ML 309 model may receive the validation data as input to evaluate the model fit on the training data set, while tuning the hyperparameters of the AI / ML 309 model. The AI / ML 309 model may receive the test data to evaluate a final model fit on the training data set and to assess the performance of the AI / ML 309 model. One or more of the training, validation, and / or testing may be performed during supervised learning for different types of AI / ML 309 models.
[0071] Supervised learning may be implemented for various types of AI / ML 309 algorithms, including algorithms that implement linear regression, logistic regression, neural networks (NNs), decision trees, Bayesian logics, random forests, and / or support vector machines (SVMs). NNs and Deep NNs (DNNs) are popular examples of algorithms utilized in AI / ML models that may be trained using supervised learning. However, the AI / ML 309 models may implement one or more NN and / or non-NN-based algorithms. Various examples of NNs include: perceptrons, multilayer perceptrons (MLPs), feed-forward NNs, fully-connected NNs, convolutional Neural Networks (CNNs), recurrent NNs (RNNs), long-short term memory (LSTM) NNs, and / or residual NNs (ResNets). A perceptron is a NN that includes a function that multiplies its input by a learned weight coefficient to generate an output value. A feed-forward NN is a NN that receives input at one or more nodes of an input layer and moves information in a direction through one or more hidden layers to one or more nodes of an output layer. In a feed-forward NN, one or more nodes of a given layer may be connected to one or more nodes of another layer. A fully connected NN is a NN that includes an input layer, one or more hidden layers, and an output layer. In a fully connected NN, each node in a layer is connected to each node in another layer of the NN. An MLP is a fully connected class of feed-forward NNs. A CNN is a NN having one or more convolutional layers configured to perform a convolution. Various types of NNs may have elements that include one or more CNNs or convolutional layers, such as Generative Adversarial Networks (GANs). GANs may include conditional GANs (CGANs), cycle-consistent GANs (CycleGANs), StyleGANs, DiscoGANs, and / or IsGANs. A GAN may include a generator sub-model and a discriminator sub-model. The generator sub-model may be configured to receive input data and pass true and independently generated data to the discriminator sub-model. The discriminator sub-model may be configured to receive the true and independently generated data from the generator, discriminate the true and independentlygenerated data, and provide feedback to the generator sub-model during training to improve the function of the generator sub-model in independently generating an output based on a received input. The GAN is a popular model for generating data types or data sequences, such as image data, audio data, and / or text, for example. An RNN is a NN that is recurrent in nature, as the nodes include feedback connections and an internal hidden state (e.g., memory) that allows output from nodes in the NN to affect subsequent input to the same nodes. LSTM NNs may be similar to RNNs in that the nodes have feedback connections and an internal hidden state (e.g., memory). However, the LSTM NNs may include additional gates to allow the LSTM NNs to learn longer-term dependencies between sequences of data. A ResNet is a NN that may include skip connections to skip one or more layers of the NN. An autoencoder may be a form of AI / ML 309 that may be implemented for supervised learning, such that parameters and / or hyperparameters may be updated during a training procedure. The parameters and / or hyperparameters may relate to the encoder portion and / or the decoder portion of the autoencoder. Some NNs include one or more attention layers or functions to enhance or focus on some portions of the input data, while diminishing or de-emphasizing other portions.
[0072] Different types of NNs and / or layers may be implemented for processing different types of data and / or producing different types of output. For example, the NN may comprise one or more convolutional layers (e.g., for CNNs or GANs), which may be popular for processing image data and / or audio data (e.g., spectrograms). Each convolutional layer may vary according to various convolutional layer parameters or hyperparameters, such as kernel size (e.g., field of view of the convolution), stride (e.g., step size of the kernel when traversing an image), padding (e.g., for processing image borders), and / or input and output size. The image being processed may include one or more dimensions (e.g, a line of pixels or a two-dimensional array of pixels). The pixels may be represented according to one or more values (e.g, one or more integer values representing color and / or intensity) that may be received by the convolutional layer. The kernel, which may also be referred to as a convolution matrix or mask, may be a matrix used to extract and / or transform features from the input data being received. The kernel may be used for blurring, sharpening, edge detection, and / or the like. An example kernel size may include a 3x3, 5x5, 10x10, etc. matrix (e.g., in pixels for a 2D image). The stride may be the parameter used to identify the amount the kernel is moved over the image data. An example default stride is of a size of 1 or 2 within the matrix (e.g., in pixels for a 2D image). The padding may include the amount of data (e.g., in pixels for a 2D image) that is added to the boundaries of the image data when it is processed by the kernel. The kernel may be moved over the input image data (e.g.,according to the stride length) and perform a dot product with the overlapping input region to obtain an activation value for the region. The output of each convolutional layer may be provided to a next layer of the NN or provided as an output (e.g., image data, feature map, etc.) of the NN itself with the updated features based on the convolution.
[0073] The NN may include layers of a similar type (e.g., convolutional layers, feed-forward layers, fully-connected layers, etc.) and / or having a similar or different configuration (e.g., size, number of nodes, etc.) for each layer. The NN may also, or alternatively, include one or more layers having different types or different subsets of NNs that may be interconnected for training and / or implementation, as described herein. For example, a NN may include both convolutional layers and feed-forward or fully-connected layers.
[0074] FIG. 3B illustrates an example of a neural network 309a. The objective of training may be to apply the input 307a as training data and / or adjust one or more weights, indicated as w and x in FIG. 3B (e.g., which may be referred to as neuron weights and / or link weights), such that the output 315 from the neural network 309a approaches the desired target values which are associated with the input 307a values for the training data. In examples, a neural network may include three layers (e.g., as shown in FIG. 3B). During the training, for given input, the difference between output and desired values may be computed and / or the difference may be used to update the one or more weights in the neural network. If a significant (e.g., large) difference between output and desired value(s) is observed, for example, one or more relatively significant (e.g., large) changes in one or more weights may be expected. A small difference (e.g., between output and desired value(s)) may result in one or more relatively small changes in one or more weights. For example, for schizophrenia assessment, the input 307a may be acoustic data and / or the output 315 may be an assessment of schizophrenia severity. The desired value may be location information acquired by global navigation satellite system (GNSS) with high accuracy.
[0075] Once the neural network 309a completes its training, the difference between the output 315 and desired values may be below a threshold. The neural network 309a may be applied or implemented after training for positioning by feeding input data 307a and / or by estimating or predicting the output 315 as the expected outcome for the associated input 307a. The output 315 may be a severity score corresponding to schizophrenia severity.
[0076] Training a neural network 309a may include identifying one or more of the following information: the input for the neural network; the expected output associated with the input; and / or the actual output from the neural network against which the target values are compared.
[0077] In examples, a neural network model may be characterized by one or more parameters and / or hyperparameters, which may include: the number of weights and / or the number of layers in the neural network.
[0078] As used herein, the term “deep learning” may refer to a class of machine learning algorithms that employ artificial neural networks (e.g., deep neural networks (DNNs)) which were loosely inspired from biological systems and / or include at least one hidden layer. DNNs may be a special class of machine learning models inspired by the human brain where the input is linearly transformed and / or pass through a non-linear activation function one or more (e.g., multiple) times. DNNs may include one or more (e.g., multiple) layers where one or more (e.g., each) layer includes linear transformation and / or a given non-linear activation function(s). The DNNs may be trained using the training data via a back-propagation algorithm. Recently, DNNs have shown state-of-the-art performance in variety of domains, e.g., speech, vision, natural language etc., and / or for various machine learning settings (e.g., supervised, unsupervised, and / or semi-supervised).
[0079] FIG. 3C is a schematic illustration of an example system environment 301a for training and implementing an AI / ML 309 model that comprises an NN 309a. However, other types of AI / ML models (e.g., including NNs and / or non-NN models) may be similarly trained and / or implemented. The NN 309a may be trained and / or implemented on one or more devices to determine and / or update parameters and / or hyperparameters 317 of the NN 309a. Raw data 303a may be generated from one or more sources. For example, the raw data 303 may include image data, text data, audio data, or another sequence of information, such as a sequence of network information related to a communication network, and / or other types of data. The raw data 303 may be preprocessed at 105a to generate training data 207a. The preprocessing may include formatting changes or other types of processing in order to generate the training data 307a in a format for being input into the NN 309a.
[0080] The NN 309a may include one or more layers 311. The configuration of the NN 309a and / or the layers 311 may be based on the parameters and / or hyperparameters 317. As described herein, the parameters may include weights, or coefficients, and / or biases for the nodes or functions in the layers 311. The hyperparameters may include a learning rate, a number of epochs, a batch size, a number of layers, a number of nodes in each layer, a number of kernels (e.g., CNNs), a size of stride (e.g., CNNs), a size of kernels in a pooling layer (e.g., CNNs), and / or other hyperparameters. As described herein, the NN 309a may include a feed forward NN, a fully connected NN a CNN, a GAN, an RNN, a ResNet, and / or one or more other types ofNNs. The NN 309a may be comprised of one or more different types of NNs or different layers for different types of NNs. For example, the NN 309a may include one or more individual layers having one or more configurations.
[0081] During the training process, the training data 307a may be input into the NN 309a and may be used to learn the parameters and / or tune the hyperparameters 317. The training may be performed by initializing parameters and / or hyperparameters of the NN 309a, generating and / or accessing the training data 307a, inputting the training data 307a into the NN 309a, calculating the error or loss from the output of the NN 309a to a target output 315a via a loss function 313 (e.g., utilizing gradient descent and / or associated back propagation), and / or updating the parameters and / or hyperparameters 317.
[0082] The loss function 313 may be implemented using backpropagati on-based gradient updates and / or gradient descent techniques, such as Stochastic Gradient Descent (SGD), synchronous SGD, asynchronous SGD, batch gradient descent, and / or mini -batch gradient descent. Examples of loss or error functions may include functions for determining a squared-error loss, a mean squared error (MSE) loss, a mean absolute error loss, a mean absolute percentage error loss, a mean squared logarithmic error loss, a pixel-based loss, a pixel-wise loss, a cross-entropy loss, a log loss, and / or a fiducial-based loss. The loss functions may be implemented in accordance one or more quality metrics, such as a Signal to Noise Ratio (SNR) metric or another signal or image quality metric.
[0083] An optimizer may be implemented along with the loss function 313. The optimizer may be an algorithm or function that is configured to adapt attributes of the NN 309a, such as a learning rate and / or weights, to improve the accuracy of the NN 309a and / or reduce the loss or error. The optimizer may be implemented to update the parameters and / or hyperparameters 317 ofthe NN 309a.
[0084] The training process may be iterated to update the parameters and / or hyperparameters 317 until an end condition is achieved. The end condition may be achieved when the output of the NN 309a is within a predefined threshold of the target output 315a.
[0085] After the training process is complete, the trained NN 309a, or portions thereof, may be stored for being implemented by one or more devices. The trained NN 309a, or portions thereof, may be implemented in other downstream algorithms or processes, as may be further described herein. The trained NN 309a, or portions thereof, may be implemented on the same device on which the training was performed. The trained NN 309a, or portions thereof, may be transmitted or otherwise provided to another device for being implemented. For example, the NN 309b,309c may include one or more portions of the trained NN 309a. The NN 309b and NN 309c receive respective input data 307b, 307c and to generate respective outputs 315b, 315c. The output 315b, 315c may be generated in one or more formats, such as a tensor, a text format (e.g., a word, sentence, or other sequence of text), a numerical format (e.g., a prediction), an audio format, an image format (e.g., including video format), another data sequence format, and / or another output format. The output 315b, 315c may be aggregated at one or more devices for being further processed and / or implemented in other downstream algorithms or processes, as may be further described herein.
[0086] Alternatively, or additionally, after the training process is complete, the trained parameters and / or tuned hyperparameters 317, or portions thereof, may be stored for being implemented by one or more devices. The trained parameters and / or tuned hyperparameters 317, or portions thereof, may be implemented in other downstream algorithms or processes, as may be further described herein. The trained parameters and / or tuned hyperparameters 317, or portions thereof, may be implemented on the same device on which the training was performed. The trained parameters and / or tuned hyperparameters 317, or portions thereof, may be transmitted or otherwise provided to another device for being implemented. For example, transmitted or otherwise provided to another device or devices that may implement the NN 309b, 309c based on the trained parameters and / or tuned hyperparameters 317. For example, the NN 309b, 309c may be constructed at another device based on the trained parameters and / or tuned hyperparameters 317, or portions thereof. The NN 309b and NN 309c may be configured from the parameters / hyperparameters 317, or portions thereof, to receive respective input data 307b, 307c and to generate respective outputs 315b, 315c. The output 315b, 315c may be generated in one or more formats, such as a tensor, a text format (e.g., a word, sentence, or other sequence of text), a numerical format (e.g., a prediction), an audio format, an image format (e.g., including video format), another data sequence format, and / or another output format. The output 315b, 315c may be aggregated at one or more devices for being further processed and / or implemented in other downstream algorithms or processes, as may be further described herein.
[0087] The AI / ML models and / or algorithms described herein may be implemented on one or more devices. For example, the AI / ML 309 may be implemented in whole or in part on one or more devices, such as one or more devices, and / or one or more other network entities, such as a network server. Example networks in which AI / ML may be distributed may include federated networks. A federated network may include a decentralized group of devices that each include AI / ML. As shown in FIG. 3C, the AI / ML 309b and AI / ML 309c may be distributed acrossseparate devices. Though FIG. 3C shows two models (e.g., AI / ML 309b and AI / ML 309c), any number of models may be implemented across any number of devices. The AI / ML may be implemented for collaborative learning in which the AI / ML is trained across multiple devices. In another example, the AI / ML may be trained at a centralized location or device and one or more portions of the AI / ML, or trained parameters and / or tuned hyperparameters, may be distributed to decentralized locations. For example, updated parameters or hyperparameters may be sent to one or more devices for updating and / or implementing the AI / ML thereon.
[0088] Characterizing neurological disorders based upon analysis of a patient’s voice is a growing field of research. As noted herein, the systems and methods described herein may be configured to train and use an AI / ML model that is configured to determine a score (e.g., which may correspond to a PANSS score) based on one or more audio recordings of the patient. The AI / ML model may be configured to determine voice biomarkers that, for example, are indicative and / or based on a correlation between a score (e.g., a PANSS score) and a parameter of the audio recording (e.g., a voice acoustic parameter).
[0089] FIG. 4 is an example diagram 400 that illustrates both positive and negative correlations between various acoustic features and psychiatric disorders. Positive correlations are illustrated using boxes with lighter fill (e.g., and a positive value), while negative correlations are illustrated using boxes with darker fill (e.g., and a negative value). Examples of such acoustic features include block, jitter, a shimmer, a tremor, Harmonics-to-noise ratio (HNR), FDR, ADR, quasiopen quotient (QOQ), normalized amplitude quotient (NAQ), a peak slope, Fl mean, F2 mean, Fl variability, F2 variability, Fl range, vowel space, LPC, mel -frequency cepstral coefficients (MFCCs), ft) mean, fO variability, fO range, intensity mean, intensity variability, energy variance, energy velocity, maximum phonation time, speech rate, articulation rate, time talking, utterance duration mean, pause duration mean, pause variability, pause rate, and / or pauses total. Although schizophrenia is discussed throughout, data (e.g., sensor data) may be useful for monitoring and / or assessing symptoms and / or progression of other diseases or disorders. For example, data may be used to monitor and / or assess depression, post-traumatic stress disorder (PTSD), obsessive compulsive disorder (OSD), bulimia, anorexia, hypomania, anxiety, and / or schizoaffective disorder. Examples of disorders and / or acoustic features are described in “ Automated Assessment of Psychiatric Disorders Using Speech: A Systematic Review", by Low, Daniel Mark, Bentley, K. EL, & Ghosh, S. S. (2019), https: / / doi.org / 10.1002 / lio2.354, which is hereby incorporated by reference in its entirety.
[0090] Data (e.g., acoustic data or recordings) utilized to train the AI / ML model of the present description has been collected from multiple studies (e.g., trials). For example, two studies (e.g., trials) were independently conducted and produced acoustic data from voice recordings and corresponding clinician provided PANSS scores for training the present model. Patients for both studies (e.g., trials) were aged 18-65 years with a confirmed diagnosis of schizophrenia or schizoaffective disorder. The first study (hereinafter “Study 1”) was an open-label, single- and multiple-dose study. Patients in the first study have been diagnosed with schizophrenia or schizoaffective disorder. Safety, tolerability, and pharmacokinetics of TV-44749 (olanzapine extended release) were assessed. Patients were recorded performing both a verbal fluency (e.g., 4 minutes) task and a free speech (e.g., 2 minutes) task (e.g., in a first sub-study). The verbal fluency task involved prompting the patient to perform tasks such as listing words that begin with a particular prompt letter and listing words that the patient associates with a variety of prompt categories. The free speech task involved prompting the patient to speak for a period of time in response to a provided prompt (e.g., an image). The first study (e.g., a first sub-study of the first study) included 12 patients with 11-19 recordings of each score. Of the 12 patients, 9 were men. The mean age of the 12 patients was 51.5 years. There were 7 PANSS scores for each patient. The mean (SD) of total PANSS within each patient were 48-70 (e.g., standard deviation of 0.9- 8.4). FIG. 6 depicts a summary of the demographics and characteristics of participants of the first study.
[0091] The second study aimed to assist in the development of algorithms to support the clinical assessment of positive and negative symptoms of schizophrenia and other diseaserelevant information in patients with schizophrenia. Voice recordings through a verbal fluency (e.g., 2 minutes) task and a free speech (e.g., 1 minute) task were collected. The recordings were collected in a mobile application. For example, the recordings may be collected in an application on a smartphone. For both studies (e.g., the first study and the second study), the correlations between the voice-based predicted PANSS components using the AI / ML model of the present description and the actual PANSS scores from qualified raters were assessed and voice markers from the free speech tasks were identified.
[0092] The second study (hereinafter “Study 2”) included 51 patients. Of the 51 patients, 43 were men. The mean age of the 53 patients was 41.1 years. Each patient had 2-7 visits with PANSS scores. The mean (SD) of total PANSS within each patient was 42-95 (e.g., standard deviation of 0-26.9). In both studies (e.g., the first study and the second study), most patients did not have significant fluctuations in symptoms severity as reflected in the PANSS scores. For bothstudies (e.g., the first study and the second study), the range for correlations between the model- predicted PANSS scores and the observed PANSS scores were r=0.85-0.92 (e.g., for the models described herein). FIG. 6 depicts a summary of the demographics and characteristics of participants of the second study (Study 2).
[0093] Analysis of the findings from the studies (e.g., the first study and the second study) using the AI / ML systems and methods described herein produced trained machine learning models demonstrating a strong correlation between PANSS scores and voice parameters (e.g., voice digital biomarkers) in patients with schizophrenia and / or schizoaffective disorder. Voice digital biomarkers are useful for providing consistent and objective monitoring and / or results. Biomarkers may be useful for continuous remote monitoring and / or predictive analytics, for example for relapses and / or treatment adherence. Biomarkers may be used in addition to, or alternatively to, support and / or validation, for example for human subjective assessments in clinical trials.
[0094] In particular, separate machine learning models, such as those described herein, were developed on data related to the verbal fluency task and data related to the free speech task obtained from the above-described studies to correlate the PANSS scores with acoustic features (e.g., digital biomarkers) from short segments (e.g., frames) of acoustic data and with complete patient visits by aggregating the predictions from the short segments of acoustic data. FIGs. 5A-5C are diagrams illustrating an example of a model developed on verbal fluency or free speech to correlate PANSS scores with acoustic features from short segments and with complete patient visits by aggregating the predictions from the short segments. FIGs. 5A-5C includes a collection and analysis of voice recording data from a mobile application. For example, FIG. 5 A illustrates examples of mobile application interfaces that may be generated by a client device or computer system and used to collect voice recordings of a patient. FIG. 5B illustrates examples of various acoustic features that can be determined and / or used by the model (e.g., a computing system or a client device) based on voice recordings that are captured using, for example, the mobile application interfaces shown in FIG. 5 A. As illustrated in FIG. 5B, the acoustic features of the voice recording could include any combination of a waveform over time, an amplification (AMP) envelope over time, an indication of gaps in speech over time, an indication of zero-crossing of the voice recording over time, an indication of mel- frequency cepstral coefficients (MFCC) over time, an indication of Fundamental frequency (FO) pitch, an indication of a Delta of the voice recording over time, and / or an indication of the linear frequency (e.g., SPECT linear frequency) of the voice recording over time. A model(e.g., the client device and / or the computer system) may use voice recordings and the acoustic features defined by those voice recordings to generate a model that is capable of determining a PANSS score for a patient and / or a severity score corresponding to a PANSS score for a patient. FIG. 5C illustrates a graph (e.g., a plot graph) that illustrates an example comparison between a predicted PANSS score that was determined using the models described herein and the observed PANSS score (e.g., as determined by a trained clinician).
[0095] Upon comparison, there was an insubstantial statistical difference between the respective outputs of the model trained on verbal fluency data and the model trained on free speech data. As such, an AI / ML model according to the present disclosure may require only a timed response to a free speech task, as opposed to both a free speech task and a verbal fluency task. Further, the acoustic data extracted from the patient voice recordings of the patients of particular importance to the trained machine learning model based upon the training data from the above-described trials includes amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma features.
[0096] A delta may be calculated by determining the difference between a value at a first time and a value at a second time. For example, the value may be an amplitude (e.g., dB) value. Delta-delta may be calculated by determining the difference between two deltas. For example, an amplitude value at a first time minus an amplitude value at a second time may be a first delta. An amplitude value at a third time minus an amplitude value at a fourth time may be a second delta value. The delta-delta may be a difference between the first delta value and the second delta value.
[0097] Amplitude envelope may include an indication of how sound volume (e.g., measured in local, dB) changes over time. For example, the amplitude envelope may include an indication of a maximum amplitude and a minimum amplitude. A root-mean square (RMS) may include a squaring of an amplitude (e.g., instantaneous amplitude) for a (e.g., each) point in time. The points in time may be equally spaced.
[0098] The RMS may be calculated by determining the average of each amplitude value (e.g., for each point in time) and taking the square root of the average. RMS deviation may be the difference between the RMS (e.g., total RMS) and the RMS for any one amplitude value (e.g., for one point in time).
[0099] Zero crossing may include a point where a value crosses a zero level axis. Forexample, the amplitude (e.g., measured in local, dB) value may cross the zero level axis at the zero crossing.
[0100] A fundamental frequency may include the lowest frequency (e.g., measured in Hz) of a periodic waveform.
[0101] The mel-frequency cepstral coefficients (MFCCs) may include an indication of a short-term power spectrum of a sound, for example a coefficient of the short-term power spectrum of a sound.
[0102] A chroma feature may include an indication of tonal content of a sound. For example, the chroma feature may include a profile of one or more pitch classes (e.g., of twelve pitch classes). The chroma feature may include harmonic and / or melodic characteristics of a sound.
[0103] FIG. 7A illustrates a first diagram 700 that shows a correlation between the segment level scores (ie., PANSS scores assessed according to a trained model for a discrete segment of a patient session as described herein, also referred to as predicted PANSS scores) and the clinician-determined PANSS scores for the patients (e.g., also referred to as actual PANSS scores) for Study 1 as described above (e.g., the top graph 700), and second diagram 710 that shows a correlation between patient session scores (i.e., PANSS score assessed according to a trained model for an entire patient session as described herein, also referred to as predicted PANSS scores) and the clinician-determined PANSS scores for the patients (e.g., also referred to as actual PANSS scores) of Study 1 (e.g., bottom graph 710).
[0104] FIG. 7B illustrates a first diagram 750 that shows a correlation between the segment level scores (i.e., PANSS scores assessed according to a trained model for a discrete segment of a patient session as described herein, also referred to as predicted PANSS scores) and the clinician-determined PANSS scores for the patient (also referred to as actual PANSS scores) for Study 2 as described above (e.g., top graph 750), and a second diagram 760 that shows a correlation between patient session scores (i.e., PANSS score assessed according to a trained model for an entire patient session as described herein, also referred to as predicted PANSS scores) and the clinician-determined PANSS scores for the patients (also referred to as actual PANSS scores) of Study 2 (e.g., the bottom graph 760). In both examples, the analyses presented were conducted on the free speech task (e.g., and not on the verbal fluency task) because, for example, a free speech task is a good example of a real-world application suitable for a low-burden monitoring tool.
[0105] Referring to FIG. 7A, the graph 700 illustrates a correlation between the segment level scores and the actual PANSS scores for a plurality of patients of Study 1. As described herein, an audio recording of a patient can be processed to segment the audio recording into discrete segments of a predetermined duration, and acoustic data can be determined that corresponds to each of the discrete segments of the audio recording. The predetermined duration can be in the range of roughly 40 - 60 ms. A machine learning model may be trained to determine a segment level score for each of the plurality of discrete segments based on the acoustic data. As such, the machine learning model may determine a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, where each segment level score is associated with a discrete segment of the plurality of discrete segments. In some examples, the machine learning model that is configured to determine the segment level score at the segment level may include a boosting model, such as an XGBoost or a gradient boosting model.
[0106] The segment level score may be an example of a digital biomarker. The segment level score may correspond to schizophrenia severity of the patient at the discrete segment level (e.g., based on acoustic data for a discrete segment of the audio data). In some examples, the segment level score may correspond to a PANSS score (e.g., for the discrete segment). The segment level score could include an indication of a probability of relapse and / or a binary indicator. The segment level score may be for diagnostic purposes. In some examples, the segment level score may be a PANSS score (e.g., for the discrete segment).
[0107] As such, the graph 700 may illustrate an example of a plurality of segment level scores, such as segment level scores 702a, 702b, associated with different patients. For example, the segment level scores 702a may comprise a plurality of segment level scores determined by the machine learning model for a first patient. For example, each segment level score of the plurality of segment level scores 702a may be associated with a discrete segment of the plurality of discrete segments of the audio recording of the first patient. Similarly, the segment level scores 702b may comprise a plurality of segment level scores determined by the machine learning model of a second patient. For example, each segment level score of the plurality of segment level scores 702b may be associated with a discrete segment of the plurality of discrete segments of the audio recording of the second patient.
[0108] The graph 710 illustrates a correlation between the patient session scores (e.g., PANSS scores) and the actual PANSS scores for the plurality of patients of the first study. Asdescribed herein, a machine learning model may be trained to determine a patient session score based on the segment level scores associated with the plurality of discrete segments of the audio recording associated with a patient. The patient session score may be an example of a digital biomarker. The patient session score may correspond to schizophrenia severity of the patient based on the audio recording (e.g., based on all of the segments of the audio recording of the patient). In some examples, the patient session score may correspond to a PANSS score. The patient session score could include an indication of a probability of relapse and / or a binary indicator. The patient session score may be for diagnostic purposes. In some examples, the patient session score may be a PANSS score.
[0109] In some examples, the machine learning model that is configured to determine the patient session score may be different than the machine learning model that is configured to determine the segment level scores associated with each of the plurality of discrete segments. For example, the machine learning model that is configured to determine the patient session score may include a clustering method, such as a K-means clustering method.
[0110] Referring now to FIG. 7B, the graph 750 illustrates a correlation between the segment level scores and the actual PANSS scores determined by a trained clinician for a plurality of patients of Study 2. As described herein, a machine learning model may be trained to determine a segment level score for each of the plurality of discrete segments of an audio recording of a patient. The machine learning model that is configured to determine the segment level score at the segment level may include a boosting model, such as an XGBoost or a gradient boosting model. As noted above, the segment level score may be an example of a digital biomarker. For example, the segment level score may correspond to schizophrenia severity of the patient at the discrete segment level (e.g., based on acoustic data for a discrete segment of the audio data).[OHl] As such, the graph 750 may illustrate an example of a plurality of segment level scores, such as segment level scores 752a, 752b, associated with different patients. For example, the segment level scores 752a may comprise a plurality of segment level scores by the machine learning model of a first patient. For example, each segment level score of the plurality of segment level scores 752a may be associated with a discrete segment of the plurality of discrete segments of the audio recording of the first patient. Similarly, the segment level scores 752b may comprise a plurality of segment level scores by the machine learning model of a second patient, where each segment level score of the plurality of segment levelscores 752b may be associated with a discrete segment of the plurality of discrete segments of the audio recording of the second patient.
[0112] The graph 760 illustrates a correlation between the patient session scores and the actual PANSS scores determined by a trained clinician for the plurality of patients of the second study. As described herein, a machine learning model may be trained to determine a patient session score based on the segment level scores associated with the plurality of discrete segments of the audio recording associated with a patient. As noted above, the patient session score may be an example of a digital biomarker. As also noted above, in some examples, the machine learning model that is configured to determine the patient session score may be different than the machine learning model that is configured to determine the segment level scores associated with each of the plurality of discrete segments. For example, the machine learning model that is configured to determine the patient session score may include a clustering method, such as a K-means clustering method.
[0113] As illustrated in FIG. 7A and 7B, a strong correlation was observed between PANSS scores and voice parameters using the models described herein during the first and second studies. For both studies, the range for correlations between the model -determined and the actual PANSS scores was r = 0.92 - 0.94 for individual segments and 0.71 - 0.80 for new patient visits (e.g., as shown in FIG. 7A and 7B).
[0114] Accordingly, the AI / ML models trained using the data of Studies 1 and 2 as described herein showed a strong correlation between PANSS scores and voice acoustic parameters (e.g., digital biomarkers, such as segment level scores and / or patient session scores), suggesting that a program or application that assesses voice-based digital biomarkers can be a supportive tool for PANSS evaluation using one or more of the models described herein. A voice-based digital biomarker tool, such as the models described herein, could address unmet needs and provide continuous remote monitoring for better disease management by early detection of clinical changes, therefore identifying adherence issues and predicting relapses. Human subjective assessments in clinical trials including patients with schizophrenia could be supplemented and validated by voice digital biomarkers and the models described herein.
[0115] FIG. 8 is a flow diagram that illustrates a training procedure 800 for training a machine learning model to determine the progression of schizophrenia of a patient with schizophrenia. The training process may be performed by a device, such as the computing device 110 of the system 100 of FIG. 1. A processor of the computing device may perform thetraining procedure 800 to train a machine learning model to determine progression of schizophrenia of a patient over a period of time. The processor may be configured to perform the procedure 800 periodically, for example, in response to an alarm or schedule that may be configurable. Additionally, or alternatively, the processor may be configured to perform the procedure 800 on command. At 810, the training procedure 800 may start.
[0116] At 812, the system may receive patient data from one or more devices (e.g., the one or more remote client devices 130a-130c, and / or the device 200). For example, the patient data may be audio data, such as one or more audio recordings of a patient. The data may be captured by a sensor, such as a microphone. The audio may be recorded (e.g., in an audio recording), for example in a memory of the device. The patient data may be recorded from free speech and / or recorded in response to one or more prompts provided to the patient. In some examples, the audio recording (e.g., data) may include the free-form speech and / or have a predetermined duration. The predetermined duration may be less than about 5 minutes, less than about 4 minutes, less than about 3 minutes, less than about 2 minutes, or less than about 1 minute. In some aspects, the predetermined duration may be about 30 seconds, about 1 minute, about 1.5 minutes, about 2 minutes, about 2.5 minutes, about 3 minutes, about 3.5 minutes, about 4 minutes, about 4.5 minutes, or about 5 minutes. The prompt may include one or more of text and / or an image that are generated by the device. In some examples, the audio recordings may be segmented into discrete segments, for example discrete segments of a predetermined duration.
[0117] In some aspects, the predetermined duration may be less than about 100 ms, less than about 90 ms, less than about 80 ms, less than about 70 ms, less than about 60 ms, less than about 50 ms, less than about 40 ms, less than about 30 ms, or less than about 20 ms. For example, the predetermined duration may be about 10-90 ms, about 20-80 ms, about 30-70 ms, or about 40-60 ms. In some aspects, the predetermined duration may be about 10 ms, about 15 ms, about ms, about 25 ms, about 30 ms, about 35 ms, about 40 ms, about 45 ms, about 50 ms, about 55 ms, about 60 ms, about 65 ms, about 70 ms, about 75 ms, about 80 ms, about 85 ms, about 85 ms, or about 90 ms..
[0118] At 814, acoustic data may be determined based on the patient data. For example, acoustic data may be determined by processing the audio recording, for example at the computing device. As noted above, in some examples, the audio recordings may be segmented into discrete segments, for example discrete segments of a predetermined duration. Theacoustic data may include one or more acoustic features, such as amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, or chroma feature.
[0119] At 816, a machine learning model may be trained. The model may be trained to produce the trained machine learning model, for example using a plurality of processed audio recordings. The plurality of audio recordings may be obtained from a plurality of patients, such as the patient data described above that was captured during one or more studies. The patients may have confirmed schizophrenia diagnoses and / or PANSS scores, and an indication of the schizophrenia diagnoses and / or PANSS scores may be provided to the model with the acoustic data. The model may be trained using PANSS scores, for example corresponding to the audio recording of each of the plurality of patients.
[0120] In some examples, multiple machine learning models may be trained at 816. For example, a first machine learning model may be trained to determine a segment level score for each of a plurality of discrete segments associated with a patient. As such, the machine learning model may be trained to determine a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, where each segment level score is associated with a discrete segment of the plurality of discrete segments. The segment level score may be an example of a digital biomarker. For example, the segment level score may correspond to schizophrenia severity of the patient at the discrete segment level (e.g., based on acoustic data for a discrete segment of the audio data). As noted above, in some examples, the machine learning model that is configured to determine the segment level score at the segment level may include a boosting model, such as an XGBoost or a gradient boosting model.
[0121] A second machine learning model may be trained to determine a patient session score based on the segment level scores associated with the plurality of discrete segments of the audio recording associated with the patient. The patient session score may be an example of a digital biomarker. For example, the patient session score may correspond to schizophrenia severity of the patient based on the audio recording (e.g., based on all of the segments of the audio recording of the patient). The second machine learning model that is configured to determine the patient session score may be different than the first machine learning model that is configured to determine the segment level scores associated with each of the plurality of discrete segments. For example, the second machine learning model that is configured todetermine the patient session score may include a clustering method, such as a K-means clustering method.
[0122] Once the training of the machine learning model(s) is complete, at step 816, the computing device may apply the model to the test data, at step 818, to examine model performance (e.g., reproducibility, error, etc.), for example including sums of squares error, and / or "random seed" reproducibility (e.g., the ability of model to produce the same answer). The computing device may examine convergent validity by comparison to PANSS scores. As noted herein, the R value of the model may be determined at 818. After training the machine learning model, the computing device may exit the procedure 800 at 820.
[0123] FIG. 9 is a flow diagram that illustrates an example procedure 900 for using a trained machine learning model to determine the progression of schizophrenia of a patient with schizophrenia. The procedure 900 may be performed by a processor of the computing device 110, for example to use a machine learning model to determine progression of schizophrenia of one or more patients. The machine learning model may be trained according to procedure 800, as described above. The processor may be configured to perform the procedure 900 periodically, for example, in response to an alarm or schedule that may be configurable. Additionally, or alternatively, the processor may be configured to perform the procedure 900 on command. The procedure 900 may start at 910.
[0124] At 911, a schizophrenia treatment regimen may be started. For example, a medication associated with schizophrenia may be administered to the patient. The medication may include one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone. The step 911 may be performed by a physician and not the computing device. In some examples, the computing device may receive an indication of the treatment regimen of the patient e.g., the initial treatment regimen of the patient) at 911.
[0125] In some examples, at 911, the treatment regimen may include administering to the patient olanzapine or the pharmaceutical composition disclosed herein, wherein the olanzapine or the pharmaceutical composition provides a therapeutically effective plasma concentration of olanzapine for a period of at least about 14 days or at least 21 days or for about 14 days, about 21 days, about 28 days, about 30 days, about 42 days or about 56 days following administration to a patient.
[0126] Examples of pharmaceutical compositions of the medication and methods for providing or delivering the medication of the present disclosure are provided in International Application No. PCT / IB2023 / 050494 (filed on January 20, 2023 and entitled “Olanzapine, compositions thereof and methods of use thereof’) and U.S. Patent No. 10,646,443 (filed on March 20, 2018 and entitled “Sustained release olanzapine formulations”), both applications are expressly incorporated herein by reference in their entirety.
[0127] The term “olanzapine” refers to the compound having the chemical name 2-methyl-4-(4- methyl-l-piperazinyl)-10H-thieno[2,3-b] [l,5]benzodiazepine and the chemical structure:
[0128]
[0129] In some examples, the medication may be delivered through a vial and, for instance, the vial may contain from about 318 mg to about 950 mg of olanzapine or a pharmaceutically acceptable salt of olanzapine. In some embodiments, the vial may contain from about 550 mg to about 950 mg of olanzapine or a pharmaceutically acceptable salt of olanzapine, or from about 450 to about 800 mg of olanzapine or a pharmaceutically acceptable salt of olanzapine.
[0130] In some aspects, the medication may be provided as part of a kit, and the kit may comprise a container containing 318 mg - 950 mg, optionally 550 - 950 mg, optionally 450 mg - 800 mg of olanzapine, such as, for example, 318 mg, 320 mg, 325 mg, 330 mg, 335 mg, 340 mg, 345 mg, 350 mg, 355 mg, 360 mg, 365 mg, 370 mg, 375 mg, 380 mg, 385 mg, 390 mg, 395 mg,400 mg, 405 mg, 410 mg, 415 mg, 420 mg, 425 mg, 430 mg, 435 mg, 440 mg, 445 mg, 450 mg,455 mg, 460 mg, 465 mg, 470 mg, 475 mg, 480 mg, 485 mg, 490 mg, 495 mg, 500 mg, 505 mg,510 mg, 515 mg, 520 mg, 525 mg, 530 mg, 531 mg, 535 mg, 540 mg, 545 mg, 550 mg, 555 mg,560 mg, 565 mg, 570 mg, 575 mg, 580 mg, 585 mg, 590 mg, 595 mg, 600 mg, 605 mg, 610 mg,615 mg, 620 mg, 625 mg, 630 mg, 635 mg, 637 mg, 640 mg, 645 mg, 650 mg, 655 mg, 660 mg,665 mg, 670 mg, 675 mg, 680 mg, 685 mg, 690 mg, 695 mg, 700 mg, 705 mg, 710 mg, 715 mg,720 mg, 725 mg, 730 mg, 735 mg, 740 mg, 743 mg, 745 mg, 750 mg, 755 mg, 760 mg, 765 mg,770 mg, 775 mg, 780 mg, 785 mg, 790 mg, 795 mg, 800 mg, 805 mg, 810 mg, 815 mg, 820 mg,825 mg, 830 mg, 835 mg, 840 mg, 845 mg, 850 mg, 855 mg, 860 mg, 865 mg, 870 mg, 875 mg,880 mg, 885 mg, 890 mg, 895 mg, 900 mg, 905 mg, 910 mg, 915 mg, 920 mg, 925 mg, 930 mg, 935 mg, 940 mg, 945 mg, or 950 mg of olanzapine.
[0131] In some aspects, the kits may comprise a container containing about 318 mg - about 950 mg, optionally about 550 - 950 mg, optionally 450 mg - 800 mg of olanzapine, such as, for example, about 318 mg, about 320 mg, about 325 mg, about 330 mg, about 335 mg, about 340 mg, about 345 mg, about 350 mg, about 355 mg, about 360 mg, about 365 mg, about 370 mg, about 375 mg, about 380 mg, about 385 mg, about 390 mg, about 395 mg, about 400 mg, about 405 mg, about 410 mg, about 415 mg, about 420 mg, about 425 mg, about 430 mg, about 435 mg, about 440 mg, about 445 mg, about 450 mg, about 455 mg, about 460 mg, about 465 mg, about 470 mg, about 475 mg, about 480 mg, about 485 mg, about 490 mg, about 495 mg, about 500 mg, about 505 mg, about 510 mg, about 515 mg, about 520 mg, about 525 mg, about 530 mg, about 531 mg, about 535 mg, about 540 mg, about 545 mg, about 550 mg, about 555 mg, about 560 mg, about 565 mg, about 570 mg, about 575 mg, about 580 mg, about 585 mg, about 590 mg, about 595 mg, about 600 mg, about 605 mg, about 610 mg, about 615 mg, about 620 mg, about 625 mg, about 630 mg, about 635 mg, about 637 mg, about 640 mg, about 645 mg, about 650 mg, about 655 mg, about 660 mg, about 665 mg, about 670 mg, about 675 mg, about 680 mg, about 685 mg, about 690 mg, about 695 mg, about 700 mg, about 705 mg, about 710 mg, about 715 mg, about 720 mg, about 725 mg, about 730 mg, about 735 mg, about 740 mg, about 743 mg, about 745 mg, about 750 mg, about 755 mg, about 760 mg, about 765 mg, about 770 mg, about 775 mg, about 780 mg, about 785 mg, about 790 mg, about 795 mg, about 800 mg, about 805 mg, about 810 mg, about 815 mg, about 820 mg, about 825 mg, about 830 mg, about 835 mg, about 840 mg, about 845 mg, about 850 mg, about 855 mg, about 860 mg, about 865 mg, about 870 mg, about 875 mg, about 880 mg, about 885 mg, about 890 mg, about 895 mg, about 900 mg, about 905 mg, about 910 mg, about 915 mg, about 920 mg, about 925 mg, about 930 mg, about 935 mg, about 940 mg, about 945 mg, or about 950 mg of olanzapine.
[0132] Further, in some aspects, the disclosure provides pharmaceutical dosage forms for administration by subcutaneous injection wherein the dosage forms comprise olanzapine and at least one biodegradable polymer.
[0133] In some aspects, the at least one biodegradable polymer is a poly(lactide), poly (glycolide), poly(lactide-co-glycolide), poly-l-lactic acid, poly-d-lactic acid, copolymers of the foregoing, poly(aliphatic carboxylic acids), copolyoxalates, poly caprolactone, polydioxanone, poly(ortho carbonates), poly (acetals), poly(lactic acid-caprolactone), polyorthoesters, poly(glycolic acid-caprolactone), poly(amino acid), polyesteramide,polyanhydrides, polyphosphazines, poly(alkylene alkylate), biodegradable polyurethane, polyvinylpyrrolidone, polyalkanoic acid, polyethylene glycol, copolymer of polyethylene glycol and polyorthoester, albumin, chitosan, casein, waxes or blends or copolymers thereof.
[0134] At 912, a patient’s data, such as an audio recording of the patient, may be received by the computing device. For example, the data may include one or more of in-clinic data, self-test (e.g., weekly) data, and / or continuous (e.g., passively) data. A patient may be instructed to produce data by free-form speech, for example in response to a prompt. The audio recording (e.g., data) may include the free-form speech and / or have a predetermined duration. The predetermined duration may be less than about 5 minutes, less than about 4 minutes, less than about 3 minutes, less than about 2 minutes, or less than about 1 minute. In some aspects, the predetermined duration may be about 30 seconds, about 1 minute, about 1.5 minutes, about 2 minutes, about 2.5 minutes, about 3 minutes, about 3.5 minutes, about 4 minutes, about 4.5 minutes, or about 5 minutes. The prompt may include one or more of text and / or an image. The first and / or second audio recording for a patient may be received. The audio recordings may be segmented into discrete segments, for example discrete segments of a predetermined duration. In some aspects, the predetermined duration may be less than about 100 ms, less than about 90 ms, less than about 80 ms, less than about 70 ms, less than about 60 ms, less than about 50 ms, less than about 40 ms, less than about 30 ms, or less than about 20 ms. For example, the predetermined duration may be about 10 to about 90 ms, about 20 to about 80 ms, about 30 to about 70 ms, or about 40 to about 60 ms. In some aspects, the predetermined duration may be about 10 ms, about 15 ms, about ms, about 25 ms, about 30 ms, about 35 ms, about 40 ms, about 45 ms, about 50 ms, about 55 ms, about 60 ms, about 65 ms, about 70 ms, about 75 ms, about 80 ms, about 85 ms, about 85 ms, or about 90 ms..
[0135] At 914, the computing device may determine acoustic data based on the audio recordings e.g., which may or may not have been segmented). The acoustic data may include one or more acoustic features, such as amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, a chroma feature.
[0136] In some examples, as described herein, the computing device may determine a plurality of sets of acoustic data based on one or more audio recordings of the patient. For example, the acoustic data may include a first set of acoustic data and / or a second set of acoustic data, for example corresponding respectively to the first and second audio recording ofthe patient (e.g., at 912). The acoustic data may be determined acoustic data corresponding to each of the discrete segments of the audio recording. Additionally, or alternatively, the acoustic data (e.g., first and / or second acoustic data) related to the audio recording (e.g., at 912) may be determined by the computing device, for example by processing a first and / or second set of acoustic data.
[0137] At 916, the (e.g., trained at 800) machine learning model may be used to determine progression of schizophrenia. The model may determine a digital biomarker. The digital biomarker may correspond to schizophrenia severity of the patient using the acoustic data. In some examples, the digital biomarker may correspond to a positive and negative syndrome scale (PANSS) score. The digital biomarker could include an indication of a probability of relapse and / or a binary indicator. The digital biomarker may be for diagnostic purposes. In some examples, the digital biomarker may be a PANSS score.
[0138] The digital biomarker may include a first digital biomarker and / or a second digital biomarker, for example corresponding to the first and / or second acoustic data respectively (e.g., at 914). The (e.g., first) digital biomarker may include an indication of one or more symptoms and / or the severity of schizophrenia. The (e.g., second) digital biomarker may correspond to schizophrenia severity of the patient. An amount of medication, for example a daily amount, may be determined and / or adjusted by a clinician, for example based on a comparison between the first and second digital biomarker provided by the model. In another example, the amount of medication may be determined and / or adjusted based on the second digital biomarker being one or more of equal to, greater than, or less than the first digital biomarker. Alternatively, a type of medication may be determined and / or adjusted by a clinician based on a comparison between the first and second digital biomarker provided by the model. Whereas a clinician may alternatively have to perform multiple time-intensive and subjective PANSS assessments, the digital biomarker provided by the model can provide quick and objective measures of schizophrenia progression to aid a clinician in determining patient treatments. The computing device may exit the procedure 900 at 918.
[0139] In some examples, multiple machine learning models may be used to determine progression of schizophrenia of a patient at 916. For example, a first machine learning model may determine a segment level score for each of a plurality of discrete segments associated with a patient. As such, the machine learning model may be trained to determine a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, whereeach segment level score is associated with a discrete segment of the plurality of discrete segments. The segment level score may be an example of a digital biomarker. For example, the segment level score may correspond to schizophrenia severity of the patient at the discrete segment level (e.g., based on acoustic data for a discrete segment of the audio data). As noted above, in some examples, the machine learning model that is configured to determine the segment level score at the segment level may include a boosting model, such as an XGBoost or a gradient boosting model.
[0140] A second machine learning model may determine a patient session score based on the segment level scores associated with the plurality of discrete segments of the audio recording associated with the patient. The patient session score may be an example of a digital biomarker. For example, the patient session score may correspond to schizophrenia severity of the patient based on the audio recording e.g., based on all of the segments of the audio recording of the patient). The second machine learning model that is configured to determine the patient session score may be different than the first machine learning model that is configured to determine the segment level scores associated with each of the plurality of discrete segments. For example, the second machine learning model that is configured to determine the patient session score may include a clustering method, such as a K-means clustering method.
[0141] The patient session score may include a first patient session score and / or a second patient session score, for example corresponding to the first and / or second acoustic data respectively (e.g., at 914). The (e.g., first) patient session score may include an indication of one or more symptoms and / or the severity of schizophrenia. The (e.g., second) patient session score may correspond to schizophrenia severity of the patient. An amount of medication, for example a daily amount, may be determined and / or adjusted by a clinician, for example based on a comparison between the first and second patient session score provided by the models. In another example, the amount of medication may be determined and / or adjusted based on the second patient session score being one or more of equal to, greater than, or less than the first patient session score. Alternatively, a type of medication may be determined and / or adjusted by a clinician based on a comparison between the first and second patient session score provided by the models. Whereas a clinician may alternatively have to perform multiple timeintensive and subjective PANSS assessments, the patient session score provided by the modelscan provide quick and objective measures of schizophrenia progression to aid a clinician in determining patient treatments.
[0142] In addition to what has been described herein, the methods and systems may also be implemented in a computer program(s), software, or firmware incorporated in one or more computer-readable media for execution by a computer(s) or processor(s), for example. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and tangible / non-transitory computer-readable storage media. Examples of tangible / non-transitory computer-readable storage media include, but are not limited to, a read only memory (ROM), a random-access memory (RAM), removable disks, and optical media such as CD-ROM disks, and digital versatile disks (DVDs).
[0143] While this disclosure has been described in terms of certain embodiments and generally associated methods, alterations and permutations of the embodiments and methods will be apparent to those skilled in the art. Accordingly, the above description of example embodiments does not constrain this disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of this disclosure.
Claims
CLAIMSWhat is claimed is:
1. A method for assessing the severity of schizophrenia in a patient, the method comprising: receiving an audio recording for a patient; segmenting the audio recording into a plurality of discrete segments of a predetermined duration; determining acoustic data corresponding to each of the discrete segments of the audio recording; determining, using a first trained machine learning model, a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, wherein each segment level score is associated with a discrete segment of the plurality of discrete segments; and determining, using a second trained machine learning model, a patient session score based on the plurality of segment level scores, wherein the patient session score corresponds to schizophrenia severity of the patient.
2. The method of claim 1, wherein the segment level score corresponds to schizophrenia severity of the patient at the discrete segment level.
3. The method of any one of claims 1-2, wherein the patient session score corresponds to a PANSS score.
4. The method of any one of claims 1-3, wherein the predetermined duration is in a range of about 40 ms to about 60 ms.
5. The method of any one of claims 1-4, wherein the first trained machine learning model comprises a different model than the second trained machine learning model.
6. The method of any one of claims 1-5, wherein the first trained machine learning model comprises boosting model, and wherein the second trained machine learning model comprises a clustering model.
7. The method of any one of claims 1-6, wherein the first trained machine learning model comprises an XGBoost model, and wherein the second trained machine learning model comprises a K-means clustering model.
8. The method of any one of claims 1-7, wherein the audio recording is a first audio recording, acoustic data is a first set of acoustic data, and the patient session score is a first patient session score, further comprising: administering to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording; receiving a second audio recording for the patient; segmenting the second audio recording into a second plurality of discrete segments of a predetermined duration; determining second acoustic data corresponding to each of the discrete segments of the second audio recording; determining, using the first trained machine learning model, a second plurality of segment level scores based on the second acoustic data of the second plurality of discrete segments; determining, using a second trained machine learning model, a second patient session score based on the second plurality of segment level scores, wherein the second patient session score corresponds to schizophrenia severity of the patient based on the second acoustic data; and adjusting a daily amount of the medication or a type of the medication for the patient based on a comparison between the second patient session score and the first patient session score.
9. The method of claim 8, wherein the medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone10. The method of any one of claims 1-9, further comprising: training a model to produce the first trained machine learning model and the second trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
11. The method of any one of claims 1-10, wherein the acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
12. The method of any one of claims 1-11, further comprising: instructing the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration.
13. The method of claim 12, wherein the prompt is text and / or an image.
14. The method of claim 12, wherein the predetermined duration is about 1 minute.
15. A method for assessing the severity of schizophrenia in a patient using a trained machine learning model, the method comprising: receiving an audio recording for a patient; processing the audio recording to determine acoustic data related to the audio recording; and determining, using the trained machine learning model, a digital biomarker corresponding to schizophrenia severity of the patient using the acoustic data.
16. The method of claim 15, wherein the audio recording is a first audio recording, acoustic data is a first set of acoustic data, and the digital biomarker is a first digital biomarker, further comprising: administering to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording; receiving a second audio recording for the patient; processing the second audio recording to determine a second set of acoustic data related to the second audio recording; determining, using the trained machine learning model, a second digital biomarker corresponding to schizophrenia severity of the patient using the second set of acoustic data; and adjusting a daily amount of the medication or a type of the medication for the patient based on a comparison between the second digital biomarker and the first digital biomarker.
17. The method of claim 16, wherein the medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone.
18. The method of any one of claims 16-17, wherein the processing step comprises: segmenting the audio recording into discrete segments of a predetermined duration; and determining the acoustic data corresponding to each of the discrete segments of the audio recording.
19. The method of claim 18, wherein the predetermined duration is about 40 ms.
20. The method of any one of claims 15-19, wherein the digital biomarker corresponds to a PANSS score.
21. The method of any one of claims 15-20, further comprising: training a model to produce the trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
22. The method of any one of claims 15-21, wherein the acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
23. The method of any one of claims 15-22, further comprising: instructing the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration.
24. The method of claim 23, wherein the prompt is text and / or an image.
25. The method of claim 23, wherein the predetermined duration is about 1 minute.
26. A system for assessing the severity of schizophrenia in a patient, the system comprising: one or more processors configured to: receive an audio recording for a patient; segment the audio recording into a plurality of discrete segments of a predetermined duration; determine acoustic data corresponding to each of the discrete segments of the audio recording; determine, using a first trained machine learning model, a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, wherein each segment level score is associated with a discrete segment of the plurality of discrete segments; and determine, using a second trained machine learning model, a patient session score based on the plurality of segment level scores, wherein the patient session score corresponds to schizophrenia severity of the patient.
27. The system of claim 26, wherein the segment level score corresponds to schizophrenia severity of the patient at the discrete segment level.
28. The system of any one of claims 26-27, wherein the patient session score corresponds to a PANSS score.
29. The system of any one of claims 26-28, wherein the predetermined duration is in a range of about 40 ms to about 60 ms.
30. The system of any one of claims 26-29, wherein the first trained machine learning model comprises a different model than the second trained machine learning model.
31. The system of any one of claims 26-30, wherein the first trained machine learning model comprises boosting model, and wherein the second trained machine learning model comprises a clustering model.
32. The system of any one of claims 26-31, wherein the first trained machine learning model comprises an XGBoost model, and wherein the second trained machine learning model comprises a K-means clustering model.
33. The system of any one of claims 26-32, wherein the audio recording is a first audio recording, acoustic data is a first set of acoustic data, and the patient session score is a first patient session score; and wherein the one or more processors are configured to: administer to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording; receive a second audio recording for the patient; segment the second audio recording into a second plurality of discrete segments of a predetermined duration; determine second acoustic data corresponding to each of the discrete segments of the second audio recording; determine, using the first trained machine learning model, a second plurality of segment level scores based on the second acoustic data of the second plurality of discrete segments; determine, using a second trained machine learning model, a second patient session score based on the second plurality of segment level scores, wherein the second patient session score corresponds to schizophrenia severity of the patient based on the second acoustic data; and adjust a daily amount of the medication or a type of the medication for the patient based on a comparison between the second patient session score and the first patient session score.
34. The system of claim 33, wherein the medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone35. The system of any one of claims 26-34, wherein the one or more processors are configured to: train a model to produce the first trained machine learning model and the second trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
36. The system of any one of claims 26-35, wherein the acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
37. The system of any one of claims 26-36, wherein the one or more processors are configured to: instruct the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration.
38. The system of claim 37, wherein the prompt is text and / or an image.
39. The system of claim 38, wherein the predetermined duration is about 1 minute.
40. A system for assessing the severity of schizophrenia in a patient using a trained machine learning model, the system comprising: one or more processors configured to: receive an audio recording for a patient; process the audio recording to determine acoustic data related to the audio recording; and determine, using the trained machine learning model, a digital biomarker corresponding to schizophrenia severity of the patient using the acoustic data.
41. The system of claim 40, wherein the audio recording is a first audio recording, acoustic data is a first set of acoustic data, and the digital biomarker is a first digital biomarker; and wherein the one or more processors are configured to: administer to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording; receive a second audio recording for the patient; process the second audio recording to determine a second set of acoustic data related to the second audio recording; determine, using the trained machine learning model, a second digital biomarker corresponding to schizophrenia severity of the patient using the second set of acoustic data; and adjust a daily amount of the medication or a type of the medication for the patient based on a comparison between the second digital biomarker and the first digital biomarker.
42. The system of claim 41, wherein the medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone43. The system of any one of claims 40-42, wherein the one or more processors are configured to: segment the audio recording into discrete segments of a predetermined duration; and determine the acoustic data corresponding to each of the discrete segments of the audio recording.
44. The system of claim 43, wherein the predetermined duration is about 40 ms.
45. The system of any one of claims 40-44, wherein the digital biomarker corresponds to a PANSS score.
46. The system of any one of claims 40-45, wherein the one or more processors are configured to: train a model to produce the trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
47. The system of any one of claims 40-46, wherein the acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
48. The system of any one of claims 40-47, wherein the one or more processors are configured to: instruct the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration.
49. The system of claim 48, wherein the prompt is text and / or an image.
50. The system of claim 48, wherein the predetermined duration is about 1 minute.
51. A computer-readable storage medium comprising instructions stored thereon that, when executed by one or more processors of one or more servers, cause the one or more processors to: receive an audio recording for a patient; segment the audio recording into a plurality of discrete segments of a predetermined duration; determine acoustic data corresponding to each of the discrete segments of the audio recording; determine, using a first trained machine learning model, a plurality of segment level scores based on the acoustic data of the plurality of discrete segments, wherein each segment level score is associated with a discrete segment of the plurality of discrete segments; and determine, using a second trained machine learning model, a patient session score based on the plurality of segment level scores, wherein the patient session score corresponds to schizophrenia severity of the patient.
52. The computer-readable storage medium of claim 51, wherein the segment level score corresponds to schizophrenia severity of the patient at the discrete segment level.
53. The computer-readable storage medium of any one of claims 51-52, wherein the patient session score corresponds to a PANSS score.
54. The computer-readable storage medium of any one of claims 51-53, wherein the predetermined duration is in a range of about 40 ms to about 60 ms.
55. The computer-readable storage medium of any one of claims 51-54, wherein the first trained machine learning model comprises a different model than the second trained machine learning model.
56. The computer-readable storage medium of any one of claims 51-55, wherein the first trained machine learning model comprises boosting model, and wherein the second trained machine learning model comprises a clustering model.
57. The computer-readable storage medium of any one of claims 51-56, wherein the first trained machine learning model comprises an XGBoost model, and wherein the second trained machine learning model comprises a K-means clustering model.
58. The computer-readable storage medium of any one of claims 51-57, wherein the audio recording is a first audio recording, acoustic data is a first set of acoustic data, and the patient session score is a first patient session score; and wherein the computer-readable storage medium comprises instructions stored thereon, that, when executed by one or more processors of one or more servers, cause the one or more processors to: administer to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording; receive a second audio recording for the patient; segment the second audio recording into a second plurality of discrete segments of a predetermined duration; determine second acoustic data corresponding to each of the discrete segments of the second audio recording; determine, using the first trained machine learning model, a second plurality of segment level scores based on the second acoustic data of the second plurality of discrete segments; determine, using a second trained machine learning model, a second patient session score based on the second plurality of segment level scores, wherein the second patient session score corresponds to schizophrenia severity of the patient based on the second acoustic data; and adjust a daily amount of the medication or a type of the medication for the patient based on a comparison between the second patient session score and the first patient session score.
59. The computer-readable storage medium of claim 58, wherein the medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone60. The computer-readable storage medium of any one of claims 51-59, wherein the computer-readable storage medium comprises instructions stored thereon, that, when executed by one or more processors of one or more servers, cause the one or more processors to: train a model to produce the first trained machine learning model and the second trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
61. The computer-readable storage medium of any one of claims 51-60, wherein the acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
62. The computer-readable storage medium of any one of claims 51-61, wherein the computer-readable storage medium comprises instructions stored thereon, that, when executed by one or more processors of one or more servers, cause the one or more processors to: instruct the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration.
63. The computer-readable storage medium of claim 62, wherein the prompt is text and / or an image.
64. The computer-readable storage medium of claim 63, wherein the predetermined duration is about 1 minute.
65. A computer-readable storage medium comprising instructions stored thereon that, when executed by one or more processors of one or more servers, cause the one or more processors to: receive an audio recording for a patient; process the audio recording to determine acoustic data related to the audio recording; and determine, using a trained machine learning model, a digital biomarker corresponding to schizophrenia severity of the patient using the acoustic data.
66. The computer-readable storage medium of claim 65, wherein the audio recording is a first audio recording, acoustic data is a first set of acoustic data, and the digital biomarker is a first digital biomarker; and wherein the computer-readable storage medium comprises instructions stored thereon that, when executed by one or more processors of one or more servers, cause the one or more processors to: administer to the patient a medication associated with schizophrenia at a first time corresponding to the first audio recording; receive a second audio recording for the patient; process the second audio recording to determine a second set of acoustic data related to the second audio recording; determine, using the trained machine learning model, a second digital biomarker corresponding to schizophrenia severity of the patient using the second set of acoustic data; and adjust a daily amount of the medication or a type of the medication for the patient based on a comparison between the second digital biomarker and the first digital biomarker.
67. The computer-readable storage medium of claim 66, wherein the medication comprises one or more of chlorpromazine, fluphenazine, haloperidol, perphenazine, thioridazine, thiothixene, aripiprazole, asenapine, clozapine, iloperidone, lurasidone, olanzapine, paliperidone, quentiapine, risperidone, and / or ziprasidone68. The computer-readable storage medium of any one of claims 65-67, wherein the computer-readable storage medium comprises instructions stored thereon that, when executed by one or more processors of one or more servers, cause the one or more processors to: segment the audio recording into discrete segments of a predetermined duration; and determine the acoustic data corresponding to each of the discrete segments of the audio recording.
69. The computer-readable storage medium of claim 68, wherein the predetermined duration is about 40 ms.
70. The computer-readable storage medium of any one of claims 65-69, wherein the digital biomarker corresponds to a PANSS score.
71. The computer-readable storage medium of any one of claims 65-70, wherein the computer-readable storage medium comprises instructions stored thereon that, when executed by one or more processors of one or more servers, cause the one or more processors to: train a model to produce the trained machine learning model using a plurality of processed audio recordings obtained from a plurality of patients with confirmed schizophrenia diagnoses and PANSS scores corresponding to the audio recording of each of the plurality of patients.
72. The computer-readable storage medium of any one of claims 65-71, wherein the acoustic data comprises one or more of amplitude envelope, root mean square deviation, zero crossing count, fundamental frequency, mel-frequency cepstral coefficients (MFCCs), delta, delta-delta, and chroma feature.
73. The computer-readable storage medium of any one of claims 65-72, wherein the computer-readable storage medium comprises instructions stored thereon that, when executed by one or more processors of one or more servers, cause the one or more processors to: instruct the patient to produce free-form speech in response to a prompt, wherein the audio recording comprises the free-form speech and has a predetermined duration.
74. The computer-readable storage medium of claim 73, wherein the prompt is text and / or an image.
75. The computer-readable storage medium of claim 73, wherein the predetermined duration is about 1 minute.
Citation Information
Patent Citations
Sustained release olanzapine formulations
US10646443B2
Driving chip and display panel
WO2023050494A1