User analysis and predictive techniques for digital therapeutic systems

A digital therapeutic system using machine learning on speech and video biometrics addresses the challenge of real-time mental state assessment, providing accurate predictions and timely interventions for postpartum depression.

JP2025534111APending Publication Date: 2025-10-09CURIO DIGITAL THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522852
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2023-10-20
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current systems fail to accurately assess a patient's mental state in real-time, particularly for conditions like postpartum depression, due to reliance on patient honesty in self-reporting and the need for expert setup of advanced sensors, leading to potential misdiagnosis and untreated mental health issues.

Method used

A digital therapeutic system using machine learning to analyze user interactions, including speech and video biometrics, to predict mental states and provide personalized content or alerts to healthcare providers.

Benefits of technology

Enables accurate, real-time mental state assessment and early intervention for conditions like postpartum depression, improving patient support and treatment outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025534111000001_ABST
    Figure 2025534111000001_ABST
Patent Text Reader

Abstract

Disclosed are systems and methods for analyzing a user interacting with or utilizing a digital therapeutic system, the systems and methods including receiving audio and video of a user in connection with the user interacting with or receiving digital therapeutic content, determining speech data or other user data from the audio and video, determining a mental state of the user based on the determined data, and taking action based on the determined mental state, such as providing the determined mental state to at least one of the user or a healthcare provider associated with the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 380,312, entitled "USER ANALYSIS AND PREDICTIVE TECHNIQUES FOR DIGITAL THERAPEUTIC SYSTEMS," filed October 20, 2022, which is incorporated herein by reference in its entirety.

[0002] Mental state refers to mood and mental state, and can include a variety of things such as emotions, desires, and pain experiences. A person's mental state is detected by analyzing several vital signs, such as body temperature, heart rate, blood pressure, and respiratory rate, which are indicators of internal physical health. Fluctuations in these vital signs indicate changes in a person's physical and mental health, and are detected through psychological evaluations, in which mental health professionals communicate with patients to learn about their thinking, behavioral patterns, and more. A person's mental state is detected by analyzing vital sign measurements and the results of psychological evaluations.

[0003] Currently, patients must visit a clinic or laboratory, where their vital signs are measured and a physical psychological evaluation is performed by a mental health professional. However, there are situations in which patients cannot physically visit or meet with a mental health professional. In such situations, patients receive online consultations in which a mental health professional conducts a psychological evaluation. However, in such situations, mental health professionals cannot accurately assess the patient's mental state in real time. For example, if a patient suffers from depression or postpartum depression, the patient may tell their doctor that they are feeling good or energetic, but may actually be suffering from depression inside. In such cases, health professionals cannot accurately assess the patient's real-time mental state, which can lead to serious problems such as increased risk of engaging in risky behavior, interfering with work and relationships, and more. Mild depression can worsen if left untreated, making it more difficult to overcome.

[0004] To solve this problem, several systems and devices have been introduced to detect a person's mental state. Patients are asked to answer a series of questions honestly. The answers are then compared with data stored in a database to predict their mental state. However, these systems and devices cannot accurately predict a patient's mental state because patients may give false answers. Advanced systems for detecting a person's mental state, consisting of sweat sensors and electrocardiogram sensors, are also available on the market, but these advanced systems are difficult to use and require experts, such as medical professionals, to set up and operate the system.

[0005] To overcome the aforementioned drawbacks, machine learning techniques have been used to detect and monitor mental states through facial expressions or by analyzing a patient's speech, emotions, and behavior. However, accurate results have not yet been achieved with such techniques. Machine learning-based systems process input images or videos to predict a patient's mental state. However, a patient may appear physically normal, but be depressed internally. Therefore, currently available systems, methods, or devices are not reliable.

[0006] Mood detection systems are also used to detect a person's mental state. However, currently available mood detectors can only detect basic aspects of a patient's mood and cannot detect complex moods, which can lead to incorrect emotional predictions and false positives. Similarly, stress detectors built into fitness bands are commercially available and process a patient's heart rate and / or respiratory rate to predict the amount of stress. However, wearing fitness bands all day is not feasible. Furthermore, radiation from these devices can cause serious illness in patients suffering from depression or neurological disorders.

[0007] In particular, pregnant women experience anxiety during the postpartum period, which, even at subclinical levels, can have extremely harmful and long-term effects on mothers and their infants. Despite a prevalence of around 20%, postpartum depression (PPD) is rarely diagnosed or treated. The ready availability of large medical claims databases provides the opportunity and means to develop AI tools to identify and quantify disease risk earlier, thereby enabling earlier intervention and better outcomes with significantly reduced associated costs.

[0008] Therefore, a system that can assess an individual's mental state and risk of PPD is needed to provide women with appropriate support, both during and after pregnancy, which is currently lacking in the medical system. Summary of the Invention

[0009] The present disclosure is directed to systems and methods for analyzing and making predictions about a user based on their interaction with a digital therapeutic system. In some embodiments, the systems and methods are configured to predict a user's mental state based on their interaction with the digital therapeutic system and / or the content provided thereby. In some embodiments, the systems and methods are configured to identify potentially at-risk individuals. By analyzing a user in this manner, a digital therapeutic system can tailor the content provided to the user, provide notifications to the user and / or healthcare providers alerting them when the user may need additional support, and take other beneficial actions.

[0010] In one embodiment, the present disclosure is directed to a computer system for providing digital therapeutic content to a user via a user device, the user device having a camera and a microphone for recording audio and video of the user, the computer system having a processor; and memory, coupled to the processor, storing instructions that, when executed by the processor, cause the computer system to provide the digital therapeutic content to the user device, receive the audio and video of the user from the user device in association with the digital therapeutic content, determine speech-based biometric indicators associated with the user from the audio, determine visual-based biometric indicators associated with the user from the video via remote photoplethysmography, determine a mental state of the user based on at least one of the determined speech-based biometric indicators of the user or the determined visual-based biometric indicators, and provide the determined mental state to at least one of the user or a healthcare provider associated with the user.

[0011] In one embodiment, the present disclosure provides a computer-implemented method for providing digital therapeutic content to a user via a user device, the user device having a camera and a microphone for recording audio and video of the user; The method is directed to a computer-implemented method comprising: providing, by the computer system, the digital therapeutic content to the user device; receiving, by the computer system, audio and video of the user from the user device in association with the digital therapeutic content; determining, by the computer system, speech-based biometric indicators associated with the user from the audio; and determining, by the computer system, visual-based biometric indicators associated with the user from the video via remote photoplethysmography; determining, by the computer system, a mental state of the user based on at least one of the determined speech-based biometric indicators or the determined visual-based biometric indicators; and providing, by the computer system, the determined mental state to at least one of the user or a healthcare provider associated with the user.

[0012] In some embodiments, the determined mental state is provided via a user interface of a digital therapeutic app executed by the user device, the digital therapeutic app being communicatively coupled to the computer system.

[0013] In some embodiments, the digital therapeutic content is configured for the treatment of postpartum depression.

[0014] In some embodiments, the memory further stores a machine learning model trained to identify a distress state based on audio input data, and the speech indicators are determined based on the machine learning model.

[0015] In some embodiments, the mental state is determined based on both the determined speech biometrics and the determined vision biometrics.

[0016] In some embodiments, the psychiatric condition comprises at least one of anxiety, depression, or post-traumatic stress disorder.

[0017] In some embodiments, the method further comprises adjusting, by the computer system, the digital therapeutic content provided to the user based on the determined mental state. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a diagram of a digital therapeutic system and systems for interacting therewith according to an embodiment of the present disclosure. [Figure 2A] FIG. 2A is a flow diagram of a first process for analyzing a user's mental state according to one embodiment of the present disclosure. [Figure 2B] FIG. 2B is a flow diagram of a second process for analyzing a user's mental state according to an embodiment of the present disclosure. [Figure 2C] FIG. 2C is a flow diagram of a third process for analyzing a user's mental state according to an embodiment of the present disclosure. [Figure 3] FIG. 3 illustrates various ML-based approaches for analyzing audio data, according to one embodiment of the present disclosure. [Figure 4] FIG. 4 illustrates a process for analyzing a user's heart rate from video relative to baseline data according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0019] This disclosure is not limited to the particular systems, devices, and methods described, as these may vary, and the terminology used herein is for the purpose of describing particular versions or embodiments only and is not intended to limit the scope of the disclosure.

[0020] The following terms, for purposes of this application, shall have the respective meanings defined below. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Nothing in this disclosure is to be construed as an admission that the embodiments described in this disclosure are not entitled to antedate such disclosure by virtue of prior invention.

[0021] As used herein, the singular forms "a," "an," and "thee" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to a "medicament" is a reference to one or more pharmaceuticals and equivalents thereof known to those skilled in the art, and so forth.

[0022] As used herein, the term "about" means plus or minus 10% of the number with which it is used. Thus, about 50 days means a range of 45 to 55 days.

[0023] As used herein, the terms "consisting of" or "consisting of" mean that an apparatus or method includes only those elements, steps, or components that are specifically recited in a particular claimed embodiment or claim.

[0024] In embodiments or claims where the term "comprising" is used as a transitional phrase, such embodiments can also be envisioned by replacing the term "comprising" with the term "consisting of" or "essentially consisting of."

[0025] As used herein, the term "module" refers to hardware, firmware, software, or any combination thereof operable to provide a specified functionality.

[0026] Digital Therapeutic System This application is generally directed to providing digital therapeutic services to users. In one embodiment, a digital therapeutic system 100 may be accessed via a user device 120 over a network 130 (e.g., the Internet or another telecommunications network). The digital therapeutic system 100 can provide digital therapeutic content 106 to a user via the device 120. The digital therapeutic system 100 may include a computer system, such as a server or server system, configured to provide the digital therapeutic content 106 to a user through the user device 120. The digital therapeutic system 100 may further include a memory 102 and a processor 104 adapted to execute instructions stored in the memory 102 to provide the digital therapeutic content 106 to the user device 120 and perform other tasks described herein. The user device 120 may include a mobile device (e.g., a smartphone), a tablet, a laptop, a desktop computer, or any other device capable of accessing and / or displaying the digital therapeutic content 106. In one embodiment, the user may download a smartphone app 122 to the user device 120 that can access the digital therapeutic system 100 and provide the digital therapeutic content 106. In other embodiments, the digital therapeutic system 100 may be accessed, for example, as a website, a web application, or a software-as-a-service (SaaS) model. The digital therapeutic app 122 may provide a user interface through which the user can access, view, and / or interact with the digital therapeutic content 106 provided via the digital therapeutic system 100.

[0027] In one potential implementation, the digital therapeutic system 100 can be configured to provide pregnancy-related or otherwise pregnancy-related digital therapeutic services. In one embodiment, the digital therapeutic content 106 may be designed to manage symptoms of depression and anxiety during pregnancy or after birth. In particular, the digital therapeutic content 106 may be designed to develop skills helpful in managing symptoms of anxiety and depression, promote social support and relationship quality, and encourage help-seeking behaviors in users. In one embodiment, the digital therapeutic content 106 provided via the digital therapy app 122 may prompt users to upload audio and / or video content, which is received by the digital therapeutic system 100. For example, the digital therapeutic content 106 may request that the user upload video (e.g., video journal entries) and / or audio of themselves. The digital therapeutic content 106 may request that the user upload video and / or audio on a regular (e.g., daily) basis, aperiodically, or in response to various user inputs or other parameters. To facilitate recording of the video and / or audio content, the user equipment 120 may include a camera 124, a microphone 126, and / or other recording devices. The uploaded user video and / or audio may be stored, for example, in a database 112 associated with the digital therapeutic system 100. In some embodiments described in more detail below, the digital therapeutic system 100 may be configured to analyze the user video and / or audio (as well as other user data) for predictive analysis to adjust the digital therapeutic content 106 delivered to a user, notify the user of detected trends, and / or notify a medical professional accordingly. In particular, the digital therapeutic system 100 may include a video analysis module 108, an audio analysis module 110, or a combination thereof.The video analysis module 108 and / or the audio analysis module 110 may be embodied as instructions stored in the memory 102 that are executable by the processor 104 to perform the described tasks. The video analysis module 108 may be configured to analyze video content uploaded by the user for predictive analysis as described in further detail below. Similarly, the audio analysis module 110 may be configured to analyze the audio data (e.g., audio recordings) uploaded by the user for predictive analysis, for example, as described in U.S. Patent Application No. 17 / 725,145, filed April 20, 2022, entitled "A SYSTEM FOR REAL TIME DETECTION AND ANALYSIS OF SPECIFIC SPEECH BIOMARKERS," described below and incorporated by reference in its entirety.

[0028] In one embodiment, the digital therapeutic system 100 may further be communicatively connected to a healthcare provider 140 associated with the user. For example, the digital therapeutic system 100 may be configured to upload data to an electronic medical record (EMR) associated with the user, send a message (e.g., email) to the user's healthcare provider, or update a user profile provided by the digital therapeutic system 100 that is accessible by the healthcare provider 140. In one embodiment, the digital therapeutic system 100 may notify the healthcare provider 140 only in response to appropriate permission granted by the user.

[0029] The digital therapeutic content 106 may be designed to provide personalized self-help tools for women, such as those trying to manage symptoms of depression or anxiety during pregnancy or after childbirth. The digital therapeutic content 106 can guide expectant and new mothers through their journey, easing the transition to parenthood and providing helpful tips, self-guided strategies, and reminders along the way. The digital therapeutic content 106 may be designed to be completed over a specific period of time (e.g., eight weeks). The digital therapeutic content 106 may include a series of modules focused on developing skills to help manage symptoms of anxiety and depression, promoting social support and relationship quality, and encouraging help-seeking behaviors. Digital tools such as those provided by the digital therapeutic content 106 are useful for providing self-guided self-improvement and treatment strategies due to their flexibility, privacy, personalization, and ease of use. Additionally, the digital therapeutic content 106 can include interactive exercises that provide personalized feedback to support learning, and built-in trackers allow users to easily track their progress through the digital therapeutic content 106.

[0030] In some embodiments, the digital therapeutic content 106 may be designed to provide cognitive behavioral therapy (CBT) to the user. Accordingly, the digital therapeutic content 106 may include one or more modules providing CBT content that the user can interact with or view to receive CBT. The digital therapeutic content 106 may be developed by or in collaboration with clinical psychologists or other professionals, for example, using CBT-based principles to encourage the development of skills that help manage symptoms of depression and anxiety. The MamaLift program addresses minimizing risk factors for postpartum depression, including lack of social support, while promoting psychological processes and self-regulation skills, such as emotion regulation, psychological flexibility, and self-understanding.

[0031] Mental state analysis As described above, a user interacts with the digital therapeutic system 100 to, for example, receive the digital therapeutic content 106 therefrom. As part of their interaction with the digital therapeutic system 100, users may periodically (e.g., daily) upload video and / or audio recordings of themselves. Advantageously, the digital therapeutic system 100 may utilize the uploaded video and / or audio generated from the user's interaction with the system 100 to monitor the user's mental state. In particular, the digital therapeutic system 100 may identify one or more biometric indicators associated with the user based on audio and / or video data and, accordingly, determine the user's mental state using machine learning and algorithmic techniques. In some embodiments, the digital therapeutic system 100 may also take a variety of different actions based on the user's detected mental state, such as adjusting the digital therapeutic content provided to the user or providing notifications to the user and / or healthcare provider 140.

[0032] One embodiment of a process 200 for analyzing a user's mental state is shown in FIG. 2. In one embodiment, the process 200 may be embodied as instructions stored in a memory (e.g., memory 102) that, when executed by a processor (e.g., processor 104), causes the digital therapeutic system 100 to perform the process 200. In various embodiments, the process 200 may be embodied as software, hardware, firmware, and various combinations thereof. In various embodiments, the process 200 may be performed by and / or among various different devices or systems. For example, various combinations of the steps of the process 200 may be performed by the digital therapeutic system 100, the network 130, and / or the user device 120 (e.g., a computer, laptop, or smartphone). In various embodiments, systems performing the process 200 may utilize distributed processing, parallel processing, cloud processing, and / or edge computing technologies. Although the process 200 is described below as being performed by the digital therapeutic system 100, it should be understood that functions may be performed accordingly by one or more devices or subsystems associated with the digital therapeutic system 100, individually or collectively.

[0033] In particular, the digital therapeutic system 100 executing the process 200 can receive audio and / or video 202 recorded by the user (by the user or a third party such as a family member). For example, the received audio and / or video 202 can include video journal entries that the user is prompted to create and upload via the digital therapeutic app 122, and can include both audio and video content. In one embodiment, the digital therapeutic app 122 can prompt the user to upload video and / or audio of themselves describing how they are feeling, either independently or in conjunction with the provision of the digital therapeutic content 106 via the digital therapeutic app 122.

[0034] Additionally, the digital therapeutic system 100 may analyze uploaded video and / or audio content (e.g., via the audio analysis module 110) 204, 206 for one or more biometric indicators associated with the user. In one embodiment, the digital therapeutic system 100 may analyze only the audio data 204 for speech-based biometric indicators. In another embodiment, the digital therapeutic system 100 may analyze only the video data 206 for vision-based biometric indicators. In yet another embodiment, the digital therapeutic system 100 may analyze the audio and video data 204, 206 in combination with each other for a variety of different biometric indicators.

[0035] In various embodiments, the digital therapeutic system 100 can analyze the audio data 204 for user-related speech biometrics using a variety of different machine learning (ML)-based and / or algorithmic techniques. The speech-based biometrics may include vocal changes exhibited by the user (e.g., due to increased muscle tension caused by stress or anxiety), speech content, etc. In one embodiment, the digital therapeutic system 100 can store and execute trained ML models for feature extraction and classification of digital signal processing of speech data. In one embodiment, the digital therapeutic system 100 may analyze speech content 206 using natural language processing techniques to identify specific words spoken by the user. In one embodiment, the digital therapeutic system 100 may analyze the user's speech signal and content 204, 206 using techniques described in U.S. Patent Application No. 17 / 725,145, incorporated herein by reference. As shown in FIG. 3, the digital therapeutic system 100 may analyze the user's speech signal and / or speech content using a shallow ML-based approach, a deep ML-based approach, or a combination thereof.

[0036] In one embodiment, the audio analysis module 110 may implement or otherwise include an audio classification model for analyzing the audio data 204. In this embodiment, the audio classification model may be trained or programmed to identify distress conditions (e.g., anxiety, depression, or post-traumatic stress disorder) within an audio sample. As generally described above, the user may be encouraged to record audio and / or video of themselves as part of a diary or self-assessment feature via the digital therapeutic app 122. Accordingly, the digital therapeutic system 100 may be configured, via the audio analysis module 110, to analyze audio recorded from the user to identify the presence of such a stress condition. If the user is determined to be exhibiting signs or symptoms of a stress condition, the digital therapeutic system 100 may take a variety of different actions, including adjusting the distal therapeutic content 106 provided to the user via the digital therapeutic app 122 or notifying a healthcare provider 140.

[0037] In one implementation, the audio classification model executable by the audio analysis module 110 to analyze the audio data 204 is constructed using a residual neural network. A residual neural network is a neural network with skip connections that connect activations in one layer to additional layers by skipping several intervening layers. The skip connections form residual blocks, and the residual neural network is built by stacking residual blocks. In this implementation, raw audio files are converted into spectrograms, which are used as input to the residual network. The raw audio data may be chunked before being converted into spectrograms. In one exemplary implementation, the audio data is separated into windows of a defined length, a fast Fourier transform is calculated for each window, the data is converted from the time domain to the frequency domain, a mel scale is generated, the frequency spectrum from the audio data is separated into a defined number of equally spaced frequencies, and a spectrogram is calculated for each window corresponding to the frequencies in the mel scale. The spectrogram for each window is then used as input to train the residual network.

[0038] In one exemplary embodiment in which the input data was not split, a sample size of 189 audio files was used, split into a training data set of 101 files and a validation data set of 88 files. In another exemplary embodiment in which the input data was split, a sample size of 5,757 audio files was used, split into a training data set of 3,054 files and a validation data set of 2,703 files. In both cases, as is common in the field of machine learning, a residual neural network was trained on the training data set and validated on the validation data set. In various implementations, the trained residual neural network demonstrated 66-75% accuracy on the validation data set. Therefore, it was determined that the trained audio classification model could accurately and consistently identify whether a user was exhibiting stress or distress based on audio recordings of the user themselves.

[0039] In various embodiments, the digital therapeutic system 100 can analyze 206 video data for user-related speech biometrics using a variety of different ML-based and / or algorithmic techniques. Video-based biometrics may include heart rate (e.g., beats per minute) or heart rate variability (HRV). Determining heart rate or HRV is useful because such biometrics are related to stress and are therefore useful for identifying whether a user is suffering from distress or anxiety (e.g., due to PPD). In one embodiment, the digital therapeutic system 100 can analyze 206 video data for user-related visual biometrics using remote photoplethysmography (rPPG) technology. In an exemplary embodiment shown in FIG. 2, the digital therapeutic system 100 may extract 208 the user's face from received user video content. In one particular embodiment, the user's face may be extracted 208 frame-by-frame. Accordingly, the digital therapeutic system 100 may identify a region of interest (ROI) on the extracted user's face image 208. In one embodiment, the ROI is classified as a skin or non-skin portion of the user's face. For the identified skin ROI, the digital therapeutic system 100 may determine the user's heart rate using rPPG. In particular, the digital therapeutic system 100 may use RGB-based statistical analysis on the identified skin cells of the corresponding ROI, which can be used to calculate the user's heart rate spectrum frame by frame. However, in other embodiments, the digital therapeutic system 100 may utilize other ML-based and / or algorithmic techniques to identify biometric indicators associated with the user.

[0040] Based on the speech-based and / or video-based biometric indicators (e.g., the user's heart rate), the digital therapeutic system 100 may determine the user's mental state 216 based on demographic measures and clinically validated measures, such as the Edinburgh Postnatal Depression Scale (EPDS). In some embodiments, the aforementioned functions may be repeated for one or more iterations (e.g., 3-5 days) to develop a baseline score for the user. Thus, the digital therapeutic system 100 may track variations in parameters calculated from audio and / or video content uploaded by the user. Based on variations in the calculated parameters over time, the digital therapeutic system 100 may predict the user's mental state.

[0041] Referring now to FIG. 4, one specific implementation of the rPPG technique for longitudinal user video data is shown. As described above, the digital therapeutic system receives video content uploaded by a user, extracts the user's face from each video frame, and processes the extracted image to identify the ROI. Once the ROI is identified, the digital therapeutic system calculates RGB values ​​for the skin ROI for each frame, from which a blood volume pulse (BVP) spectrum can be determined. This spectrum can be used to determine the user's heart rate across frames of video content. In some embodiments, the extracted face data may undergo preprocessing before calculating the BVP spectrum. In particular, the preprocessing can include detrending and filtering. In some embodiments, the BVP spectrum may be calculated using a variety of different methods 223, including independent component analysis (ICA), principal component analysis (PCA), point-of-sale (POS), singe-scale retinex (SSR), local group invariance (LGI), GREEN, CHROM, local group invariance (LGI), and other techniques. Additionally, the digital therapeutic system 100 may obtain 224 previously determined ground truth data or baseline values ​​for the user (e.g., from database 112), analyze 226 the pre-characterized baseline heart rate values, and compare 228 the user's heart rate (i.e., beats per minute) for a particular instance of uploaded video content to the user's baseline value. If the user's heart rate for the particular uploaded video content deviates from the user's baseline value by at least a threshold amount, this may indicate a problem with the user's mental state.

[0042] As described above, the digital therapeutic system 100 can implement one or more machine learning models and / or algorithms to perform the functions of the process 200 described above, including analyzing audio data 204 and analyzing video data 206. In various embodiments, the machine learning models and / or algorithms may include neural networks, decision trees (e.g., random forests), support vector machines, regression, hidden Markov models, and other types of machine learning techniques known in the art. Furthermore, the neural networks may include any general category of neural networks, including deep neural networks, convolutional neural networks, autoencoders, recurrent neural networks, etc. Furthermore, the machine learning models described herein can be trained using supervised or unsupervised learning techniques.

[0043] In some embodiments, the process 200 performed by the digital therapeutic system 100 may be executed when audio and / or video content is uploaded by a user. Accordingly, the digital therapeutic system 100 may provide 218 the user's predicted mental state. In various embodiments, the user's predicted mental state may be provided 218 to the user (e.g., as a push notification or via the UI of the digital therapeutic app 122) or to a healthcare provider 140 (e.g., as an email or as a message delivered via a healthcare provider web portal for the digital therapeutic system 100).

[0044] Additionally, while the functions and / or steps of the process 200 are depicted in a particular order or arrangement, it should be noted that the depicted order and / or arrangement of steps and / or functions is provided for illustrative purposes only. Unless expressly stated to the contrary herein, various steps and / or functions of the process 200 may be performed in different orders, in parallel with one another, or in an interleaved manner.

[0045] While various illustrative embodiments incorporating principles of the present teachings have been disclosed, the present teachings are not limited to the disclosed embodiments. Instead, this application is intended to cover any variations, uses, or adaptations of the present teachings using their general principles. Further, this application is intended to cover such departures from the present disclosure as come within known or customary practice in the art to which the present teachings pertain.

[0046] In the above detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, like numerals generally identify like elements unless the context dictates otherwise. The exemplary embodiments described in this disclosure are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the various features of the present disclosure, as generally described herein and illustrated in the figures, may be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein.

[0047] The present disclosure is not limited in terms of the specific embodiments described in this application, which are intended as illustrations of various features. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from the spirit and scope of the present disclosure. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing description. It is to be understood that the present disclosure is not limited to particular methods, reagents, compounds, compositions, or biological systems. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0048] With respect to the use of virtually any plural and / or singular term herein, one of ordinary skill in the art can translate from plural to singular and / or from singular to plural as appropriate to the context and / or application. Various singular / plural permutations may be expressly provided herein for clarity.

[0049] In general, those skilled in the art will understand that the terms used herein are generally intended as "open" terms (e.g., the term "comprising" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," the term "comprising" should be interpreted as "including, but not limited to," etc.). Although various compositions, methods, and devices are described in terms "consisting of" various components or steps (which should be interpreted to mean "including, but not limited to"), the compositions, methods, and devices can also "consist essentially of" or "consist of" the various components and steps, and such terms should be interpreted to define an essentially closed group of members.

[0050] Furthermore, even when a particular number is explicitly recited, one of ordinary skill in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the plain recitation "two occurrences" without other modifiers means at least two occurrences, or two or more occurrences). Furthermore, when phrases similar to "at least one of A, B, and C, etc." are used, such syntax is generally intended in the sense that one of ordinary skill in the art would understand the idiom (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). When phrases similar to "at least one of A, B, or C, etc." are used, generally, such configuration is intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those of ordinary skill in the art will further understand that virtually any conjunction word and / or phrase presenting two or more alternative terms, whether in this specification, sample embodiments, or drawings, should be understood to contemplate the possibility of including either term, either term, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B," or "A and B."

[0051] Additionally, where features of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is thereby also described in terms of any individual member or subgroup of members of the Markush group.

[0052] As will be understood by those skilled in the art, all ranges disclosed herein encompass all possible subranges and combinations of subranges for all purposes, including in terms of providing a written description. Any recited range can be readily recognized as fully descriptive and allowing for the same range to be broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third, upper third, etc. As will also be understood by those skilled in the art, all terms such as "up to," "at least," etc., are inclusive of the recited number and refer to ranges that can be subsequently broken down into subranges, as described above. Finally, as will be understood by those skilled in the art, ranges include individual members. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, 5 cells, etc.

[0053] The term "about" as used herein refers to variations in numerical quantities that may occur, for example, due to real-world measurement or handling procedures, inadvertent errors in these procedures, differences in the manufacture, source, or purity of compositions or reagents, etc. Generally, the term "about" as used herein means greater or less than the stated value or range of values ​​by 1 / 10 of the stated value, e.g., ±10%. The term "about" also refers to variations that would be recognized as equivalents by those of ordinary skill in the art, provided that the variations do not encompass known values ​​practiced by the prior art. Each value or range of values ​​preceded by the term "about" is also intended to encompass embodiments of the stated absolute value or range of values. Whether modified by the term "about," quantitative values ​​referred to in this disclosure include equivalents to the stated value, e.g., variations in the numerical quantities of such values ​​that may occur but are recognized as equivalents by those of ordinary skill in the art.

[0054] The various features and functions disclosed above, or alternatives thereof, may be combined into many other different systems or applications. Various alternatives, modifications, variations, or improvements not presently foreseen or anticipated by those skilled in the art may subsequently be made, each of which is also intended to be encompassed by the disclosed embodiments.

[0055] The functions and process steps herein may be performed automatically or in whole or in part in response to user instructions. An activity (including a step) that is performed automatically is performed in response to one or more executable instructions or device operations without a user directly initiating the activity.

Claims

1. 1. A computer system for providing digital therapeutic content to a user via a user device, the user device having a camera and a microphone for recording audio and video of the user, the computer system comprising: a processor; a memory coupled to the processor, the memory being configured to, when executed by the processor, cause the computer system to: causing the digital therapeutic content to be provided to the user device; receiving the audio and video of the user from the user device in association with the digital therapeutic content; determining a speech-based biometric associated with the user from the audio; determining a visual-based biometric indicator associated with the user from the image by remote photoplethysmography; determining a mental state of the user based on at least one of the determined speech-based biometric indicators or the determined vision-based biometric indicators; and storing instructions for providing the determined mental state to at least one of the user or a healthcare provider associated with the user. The memory; A computer system comprising:

2. 10. The computer system of claim 1, wherein the determined mental state is provided via a user interface of a digital therapeutic app executed by the user device, the digital therapeutic app being communicatively coupled to the computer system.

3. 3. The computer system of claim 1 or claim 2, wherein the digital therapeutic content is configured for the treatment of postpartum depression.

4. 4. The computer system according to claim 1, wherein the memory comprises: The computer system further stores a machine learning model trained to identify a distress state based on speech input data, and the speech indicators are determined based on the machine learning model.

5. 5. The computer system of claim 1, wherein the mental state is determined based on both the determined speech biometrics and the determined vision biometrics.

6. 6. The computer system of claim 1, wherein the mental condition comprises at least one of anxiety, depression, and post-traumatic stress disorder.

7. 7. The computer system according to claim 1, wherein the memory further comprises: A computer system storing instructions that, when executed by the processor, adjust the digital therapeutic content provided to the user based on the determined mental state.

8. 1. A computer-implemented method for providing digital therapeutic content to a user via a user device, the user device having a camera and a microphone for recording audio and video of the user; providing, by a computer system, the digital therapeutic content to the user device; receiving, by the computer system, the audio and video of the user associated with the digital therapeutic content from the user device; determining, by the computer system, a speech-based biometric indicator associated with the user from the audio; determining, by the computer system, a visual-based biometric indicator associated with the user from the video via remote photoplethysmography; determining, by the computer system, a mental state of the user based on at least one of the determined speech-based biometric indicators or the determined vision-based biometric indicators; providing, by the computer system, the determined mental state to at least one of the user or a healthcare provider associated with the user; A method comprising:

9. 10. The method of claim 8, wherein the determined mental state is provided via a user interface of a digital therapeutic app executed by the user device, the digital therapeutic app being communicatively coupled to the computer system.

10. 10. The method of claim 8 or claim 9, wherein the digital therapeutic content is configured for the treatment of postpartum depression.

11. 11. The method of claim 8, wherein the computer system executes a machine learning model trained to identify distress states based on audio input data, and the speech indicators are determined based on the machine learning model.

12. The method of any of claims 8 to 11, wherein the mental state is determined based on both the determined speech-based biometric indicators and the determined vision-based biometric indicators.

13. 13. The method of any one of claims 8 to 12, wherein the psychiatric condition comprises at least one of anxiety, depression, and post-traumatic stress disorder.

14. The method according to any one of claims 8 to 13, further comprising: and adjusting, by the computer system, the digital therapeutic content provided to the user based on the determined mental state.