Voice-based medical assessment
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-03-03
AI Technical Summary
Manual assessments of neurological disorders and diseases are often inaccurate and inconsistent, and medical professionals may not be available when conditions arise.
A voice-based medical assessment system that includes a query module to audibly ask questions, a response module to receive verbal responses, and a detection module to analyze these responses for medical condition assessment, independent of language or dialect, using acoustic features.
Provides accurate and consistent medical condition assessments, including concussion evaluation, through automated analysis of voice patterns, offering non-invasive and cost-effective alternatives to traditional methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to speech analysis, and more particularly to the automated assessment and diagnosis of one or more medical conditions based on collected speech samples. [Background technology]
[0002] Assessments of neurological disorders and diseases and other medical conditions are often performed manually by medical professionals and may be based on forms completed by hand using pencil and paper. Manual assessments can be inaccurate and / or inconsistent, and medical professionals are not always available when a disorder or other medical condition arises. Summary of the Invention [Means for solving the problem]
[0003] An apparatus for voice-based medical assessment is described. In one embodiment, a query module is configured to audibly ask a question to a user through a speaker of a mobile computing device. In particular embodiments, a response module is configured to receive a user's verbal response from a microphone of the mobile computing device. In certain embodiments, a detection module is configured to present a medical condition assessment to the user based on an analysis of the received user's verbal response.
[0004] In other embodiments, the apparatus includes means for audibly querying a user from the mobile computing device. In particular embodiments, the apparatus includes means for receiving a verbal response from the user on the mobile computing device. In some embodiments, the apparatus includes means for assessing a medical condition of the user based on the received verbal response from the user.
[0005] A voice-based medical evaluation system is described. In certain embodiments, multiple distributed voice modules are disposed on a computing device for multiple users. In one embodiment, the multiple distributed voice modules are configured to question the multiple users and / or record verbal responses from the multiple users on the computing device. In various embodiments, a backend server device is configured to store at least baseline recorded verbal responses from the multiple users, test case recorded verbal responses from the multiple users, and / or medical condition evaluations for at least the test case recorded verbal responses. In one embodiment, the backend server is configured to serve the stored baseline recorded verbal responses, test case recorded verbal responses, and / or evaluations for at least a subset of the multiple users via the multiple distributed voice modules on the computing device.
[0006] A method for audio-based medical assessment is described. In one embodiment, the method includes asking a user one or more questions using a user interface of a computing device. In other embodiments, the method includes recording, on the computing device, one or more baseline verbal responses of the user to the one or more questions. In particular embodiments, the method includes re-asking the user one or more questions using a user interface of the computing device in response to a potential concussion event. In certain embodiments, the method includes recording, on the computing device, one or more test case verbal responses of the user to the one or more re-asked questions. In one embodiment, the method includes assessing, on the computing device, the likelihood of the user suffering from a concussion based on audio analysis of the one or more recorded baseline verbal responses and the one or more recorded test case verbal responses.
[0007] A computer program product is described that includes a computer-readable storage medium. In certain embodiments, the computer-readable storage medium stores computer-usable program code executable to perform operations for audio-based medical assessment. In certain embodiments, one or more of these operations may be substantially similar to one or more steps described above with respect to the disclosed devices, systems, and / or methods. The present invention provides, for example, the following items. (Item 1) 1. An apparatus comprising: a query module configured to ask a question audibly to a user through a speaker of the mobile computing device; a response module configured to receive a user's verbal response from a microphone of the mobile computing device; a detection module configured to provide a medical condition assessment to the user based on an analysis of the verbal responses received from the user; An apparatus comprising: (Item 2) 2. The apparatus of claim 1, wherein the detection module is configured to determine the assessment based on one or more acoustic features of the received verbal response without taking into account linguistic features of the received verbal response, such that the assessment is independent of one or more of a language and dialect of the received verbal response. (Item 3) 10. The device of claim 1, wherein the user comprises a clinical trial participant and the evaluation comprises an evaluation of the efficacy of a medical treatment for the medical condition. (Item 4) 4. The apparatus of claim 3, wherein the detection module includes at least one of a plurality of distributed detection modules disposed on mobile computing devices for a plurality of clinical trial participants, including a placebo group not receiving the medical treatment and a group receiving the medical treatment, and the plurality of distributed detection modules are configured to perform a blinded assessment of the medical condition for both the placebo group and the group receiving the medical treatment. (Item 5) 4. The device of claim 3, wherein the assessment of the efficacy of the medical treatment is based, at least in part, on one or more biomarkers indicative of the user's quality of life from the received verbal responses. (Item 6) 6. The device of claim 5, wherein the one or more biomarkers are indicative of one or more of physical fatigue, tiredness, mental fatigue, stress, anxiety, and depression. (Item 7) 10. The device of claim 1, wherein the user comprises a prospective clinical trial participant and the assessment comprises the user's eligibility for a clinical trial for a medical condition. (Item 8) 2. The apparatus of claim 1, wherein the evaluation includes a first score determined for the user on the mobile computing device. (Item 9) 9. The apparatus of claim 8, wherein the evaluation further includes a second score determined for the user on a backend server in communication with the mobile computing device over a network. (Item 10) 2. The device of claim 1, wherein the medical condition comprises a concussion. (Item 11) 2. The device of claim 1, wherein the medical condition comprises one or more of depression, stroke, Alzheimer's disease, and Parkinson's disease. (Item 12) 10. The apparatus of claim 1, further comprising an interface module configured to play back the recorded verbal responses received from the user and the recorded verbal responses received from a plurality of other users to different users based on hierarchical access control permissions for the different users. (Item 13) 10. The apparatus of claim 1, wherein the response module is further configured to receive data from one or more sensors of the mobile computing device, and the detection module is further configured to perform the analysis based at least in part on the received data. (Item 14) Item 14. The device of item 13, wherein the one or more sensors include an image sensor and the received data includes one or more images of the user. (Item 15) Item 14. The device of item 13, wherein the one or more sensors include a touch screen, and the received data includes touch input received from the user during an interactive video game configured to screen the user for one or more symptoms of the medical condition. (Item 16) Item 14. The apparatus of item 13, wherein the one or more sensors include one or more of an accelerometer and a gyroscope, and the received data includes information about movement of the mobile computing device by the user. (Item 17) 1. A system comprising: a plurality of distributed voice modules disposed on a computing device for a plurality of users, the plurality of distributed voice modules configured to ask questions of the plurality of users and record verbal responses from the plurality of users on the computing device; a back-end server device configured to store at least baseline recorded verbal responses from the plurality of users, test case recorded verbal responses from the plurality of users, and medical condition evaluations for at least the test case recorded verbal responses, and to provide the stored baseline recorded verbal responses, test case recorded verbal responses, and evaluations to at least a subset of the plurality of users on the computing devices via the plurality of distributed voice modules; A system comprising: (Item 18) 18. The system of claim 17, wherein the plurality of users includes participants in a clinical trial for the medical condition, and a subset of the plurality of users includes one or more administrators of the clinical trial who have hierarchical access control permissions to access the stored baseline recorded verbal responses, test case recorded verbal responses, and assessments. (Item 19) 18. The system of claim 17, wherein the plurality of distributed speech modules are configured to make a determination about the evaluation on the computing device based on the reference recorded verbal responses and the test case recorded verbal responses. (Item 20) 18. The system of claim 17, wherein the backend server device is configured to make a determination about the evaluation based on the reference recorded verbal responses and the test case recorded verbal responses. (Item 21) 1. An apparatus comprising: means for audibly querying a user from the mobile computing device; means for receiving a verbal response from the user on the mobile computing device; means for assessing the user for a medical condition based on verbal responses received from the user; An apparatus comprising: (Item 22) Item 21. The device according to item 21, further comprising: means for authenticating different users in a hierarchy of users; means for granting access to different records and different ratings to different users based on tier access control permissions for the user's tier; An apparatus comprising: (Item 23) 1. A method comprising: using a user interface of the computing device to ask the user one or more questions; recording, on a computing device, one or more baseline verbal responses of the user to the one or more questions; responsive to a potential concussion event, re-asking the user the one or more questions using a user interface of the computing device; recording, on a computing device, one or more test case verbal responses of the user to the re-asked one or more questions; assessing, on a computing device, a likelihood that the user has a concussion based on an audio analysis of the one or more recorded reference verbal responses and the one or more recorded test case verbal responses; A method comprising: (Item 24) 24. The method of claim 23, wherein the speech analysis is based on one or more acoustic features without taking into account linguistic features, such that the evaluation is independent of one or more of the language and dialect of the one or more recorded reference oral responses and the one or more recorded test case oral responses. (Item 25) 24. The method of claim 23, further comprising receiving data related to the user from one or more of an image sensor, a touch screen, an accelerometer, and a gyroscope, and wherein the evaluation is further based at least in part on the received data. [Brief explanation of the drawings]
[0008] In order that the advantages of the present invention may be readily understood, the invention, briefly described above, will now be more particularly described by reference to specific embodiments thereof which are illustrated in the accompanying drawings, through which the invention will be described and explained with more specificity and detail, it being understood that these drawings illustrate only typical embodiments of the invention and therefore should not be considered as limiting its scope. [Figure 1A] FIG. 1 is a schematic block diagram illustrating one embodiment of an audio-based medical evaluation system. [Figure 1B] FIG. 1 is a schematic block diagram illustrating another embodiment of an audio-based medical evaluation system. [Figure 2] FIG. 1 is a schematic block diagram illustrating an embodiment of a system for processing audio data with a mathematical model to perform a medical diagnosis. [Figure 3] FIG. 1 is a schematic block diagram illustrating one embodiment of a training corpus of speech data. [Figure 4] FIG. 1 is a schematic block diagram illustrating one embodiment of a list of prompts for use in diagnosing a medical condition. [Figure 5] FIG. 1 is a schematic block diagram illustrating one embodiment of a system for selecting features to train a mathematical model for diagnosing a medical condition. [Figure 6A] FIG. 1 is a schematic block diagram illustrating one embodiment of a graphical representation of pairs of feature values and diagnostic values. [Figure 6B] FIG. 10 is a schematic block diagram illustrating another embodiment of a graphical representation of pairs of feature and diagnostic values. [Figure 7] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for selecting features to train a mathematical model for diagnosing a medical condition. [Figure 8]FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for selecting a prompt for use with a mathematical model for diagnosing a medical condition. [Figure 9] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for training a mathematical model for diagnosing a medical condition appropriate to a set of selected prompts. [Figure 10] FIG. 1 is a schematic block diagram illustrating one embodiment of a computing device that can be used to train and deploy mathematical models to diagnose medical conditions. [Figure 11] FIG. 2 is a schematic block diagram illustrating one embodiment of an audio module. [Figure 12] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for audio-based medical assessment. [Figure 13] FIG. 10 is a schematic flow chart diagram illustrating another embodiment of an audio-based medical assessment method. DETAILED DESCRIPTION OF THE INVENTION
[0009] References throughout this specification to "one embodiment," "an embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, throughout this specification, appearances of "in one embodiment," "in an embodiment," and similar language may all refer to the same embodiment, but may, but do not necessarily, mean "one or more but not all embodiments," unless expressly specified otherwise. Words such as "including," "comprising," "having," and variations thereof mean "including but not limited to," unless expressly specified otherwise. An enumerated list of items does not imply that any or all of those items are mutually exclusive and / or mutually inclusive, unless expressly specified otherwise. Additionally, the terms "a," "an," and "the" are intended to mean "one or more" unless expressly specified otherwise.
[0010] Furthermore, the features, advantages, and characteristics of the described embodiments may be combined in any suitable manner. However, those skilled in the art will recognize that embodiments may be practiced without one or more of the specific features or advantages of a particular embodiment. In other cases, additional features and advantages may be recognized in certain embodiments, but may not be present in all embodiments.
[0011] These features and advantages of the embodiments will become more fully apparent from the following description and appended claims, or may be learned by the practice of the embodiments as set forth below. As will be appreciated by one skilled in the art, aspects of the present invention may be embodied as a system, method, and / or computer program product. Accordingly, aspects of the present invention may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, all of which may be collectively referred to herein as a "circuit," "module," or "system." Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable medium(s) having program code embodied therein.
[0012] Many of the functional units described herein are labeled modules (or components) to more specifically emphasize their implementation independence. For example, a module may be implemented as a hardware circuit comprising custom VLSI circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc.
[0013] Modules may also be implemented in software for execution by various types of processors. For illustrative purposes, an identified module of program code may comprise one or more physical or logical blocks of computer instructions, which may, for illustrative purposes, be organized as objects, procedures, or functions. However, the executables of an identified module need not be physically located together and may comprise entirely different instructions stored in different locations that, when logically coupled together, constitute a module and perform the module's stated purpose.
[0014] Indeed, a module of program code may be one instruction, or many instructions, and may be distributed across various different code segments, among different programs and across various memory devices. Similarly, operational data may be identified and illustrated within a module herein and may be embodied in any suitable form and organized within any suitable type of data structure. The operational data may be collected as a single data collection or distributed across different locations, including across different storage devices, or may exist, at least in part, simply as electronic signals over a system or network. When a module or portion of a module is implemented in software, the program code may also be stored on and / or propagated within one or more computer-readable media.
[0015] A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to carry out aspects of the present invention.
[0016] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, or semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer diskettes, hard disks, random access memory ("RAM"), read-only memory ("ROM"), erasable programmable read-only memory ("EPROM" or flash memory), static random access memory ("SRAM"), portable compact disk read-only memory ("CD-ROM"), digital versatile disk ("DVD"), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as ridge structures in grooves in which instructions are recorded, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as being, per se, radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or transitory signals such as electrical signals transmitted down a wire.
[0017] The computer-readable program instructions described herein may be downloaded to each computing / processing device from a computer-readable storage medium or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. In each computing / processing device, a network adapter card or network interface receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0018] The computer-readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and the like, and traditional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection to the external computer may be made (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may utilize the state information of the computer readable program instructions to execute the computer readable program instructions and personalize the electronic circuitry to perform aspects of the present invention.
[0019] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0020] These computer-readable program instructions, when executed by a processor of a computer or other programmable data processing apparatus, can provide these instructions to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored on a computer-readable storage medium and can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein constitutes an article of manufacture containing instructions for performing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0021] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device and cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0022] The schematic flowcharts and / or block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of apparatuses, systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the schematic flowcharts and / or block diagrams may represent a module, segment, or portion of code that constitutes one or more executable instructions of program code for implementing the specified logical function(s).
[0023] It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. Other steps and methods may be conceived that are equivalent in function, logic, or effect to one or more blocks, or portions thereof, of the illustrated figures.
[0024] It will be understood that various types of arrows and lines may be employed in the flowcharts and / or block diagrams, but these do not limit the scope of the corresponding embodiments. Indeed, some arrows and other connectors may be used only to indicate the logical flow of the illustrated embodiments. By way of illustration, arrows may indicate wait or monitoring periods of unspecified length between recited steps in the illustrated embodiments. It will also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and program code.
[0025] 1A illustrates one embodiment of a system 100 for audio collection and / or audio-based medical assessment. In one embodiment, the system 100 includes one or more hardware devices 102, one or more audio modules 104 (e.g., one or more audio modules 104a, one or more back-end audio modules 104b, etc., located on one or more hardware devices 102), one or more data networks 106 or other communication channels, and / or one or more back-end servers 108. While a specific number of hardware devices 102, audio modules 104, data networks 106, and / or back-end servers 108 are shown in FIG. 1 in certain embodiments, those skilled in the art will recognize, in light of this disclosure, that any number of hardware devices 102, audio modules 104, data networks 106, and / or back-end servers 108 may be included in the system 100 for audio collection and / or audio-based medical assessment.
[0026] In general, in various embodiments, the voice module 104 is configured to receive and / or record voice audio data from a user (e.g., a patient, an athlete, another user, etc.) and / or assess and / or diagnose the presence and / or severity of one or more medical conditions (e.g., a disability, an illness, a disease, etc.) based on the collected voice audio data. The voice module 104 asks or queries the user (e.g., audibly through speakers, headphones, etc., visible by written text on a hardware display device, and / or otherwise using one or more user interface elements of the hardware device 102) and prompts a verbal response from the user, which the voice module 104 receives and / or records. The voice module 104 can assign a rating to the user and provide the rating and / or the voice audio data to a back-end voice module 104b, etc.
[0027] The voice module 104 can interact with the user by asking questions, recording the user's voice responses, determining whether the responses are accurate, etc. For a particular protocol, the voice module 104 can ask one or more questions multiple times (e.g., two, three, etc.) before moving on to a subsequent question, etc. Based on the voice audio data, the voice module 104 can assess and / or diagnose one or more diseases or other medical conditions (e.g., concussion, depression, stress, stroke, cognitive well-being, mood, honesty, Alzheimer's disease, Parkinson's disease, cancer, etc.). For example, after capturing the audio, the voice module 104 may score the response (e.g., by the device voice module 104a on the hardware device 102) and present one or more initial scores to the user, and may further analyze the audio (e.g., by the backend voice module 104b on the server device 108) and present a secondary score to the user regarding conclusions and / or other specific diseases or medical conditions. The voice module 104 may extract one or more word sequences and / or features and pass the extracted word sequences and / or features to one or more machine learning models trained for specific diseases and / or other medical conditions.
[0028] The audio module 104 can compare the user's answers to previous answers from when the user was healthy (e.g., to a baseline answer). The audio module 104 can normalize the results based on the user's demographic data (e.g., age, gender, etc.). The audio module 104 can determine the normalization data as part of the training process, such as a range of expected scores for each demographic.
[0029] In certain embodiments, instead of presenting an assessment, in addition to presenting an assessment, as part of an assessment, etc., the audio module 104 can assess the efficacy and / or success of a clinical trial, drug approval process, etc. For example, instead of or in addition to administering a questionnaire to a clinical trial participant, which may be subjective, the audio module 104 can objectively assess and / or model changes in the audio of a clinical trial participant over the course of a clinical trial. For example, the audio module 104 may collect audio from a clinical trial and / or research participant (e.g., at a doctor's visit, at home, etc.) and further create one or more models for a placebo group and / or a test group. In some embodiments, the audio module 104 may compare the results of the audio assessment and / or modeling for a medical trial and / or research participant with the results of a questionnaire or other test, present a similar score and / or on the same scale as the questionnaire or other test, etc. In certain embodiments, the audio module 104 may perform audio-based medical evaluations to determine the effectiveness of treatment protocols (e.g., medications and / or other treatment procedures) where chemistry is unknown, to verify and / or determine the effectiveness of chemistry, etc.
[0030] In certain embodiments, the voice module 104 may also screen participants in preparation for clinical trials and / or medical studies. For example, the voice module 104 may qualify an individual for a depression study, etc., if the individual clearly exhibits biomarkers in their voice that are consistent with individuals suffering from depression. Screening participants for clinical trials and / or medical studies using the voice module 104 to identify biomarkers in their voices is arguably more objective and / or accurate than subjectively identifying trial participants using written questionnaires or similar tools. Other methods that can achieve objectivity and / or accuracy, such as blood tests, magnetic resonance imaging (MRI) scans, etc., may, in certain embodiments, be more expensive and invasive than voice analysis by the voice module 104. The voice module 104, in one embodiment, can provide the same objectivity and / or accuracy as other tests, but in a non-invasive and cost-effective manner. In one embodiment, clinical trial participant identification using the voice module 104 provides an objective, biomarker-data-driven tool.
[0031] In some embodiments, the voice module 104 may also differentiate and / or qualify one or more new medications (e.g., drugs) using behavioral parameters (e.g., behavioral parameters that are objectively measured and determined to contribute to quality of life, rather than simply approving a medication based on its effective disease prevention, etc.). In one embodiment, the voice module 104 uses voice biomarkers to identify a person's condition (e.g., physical fatigue, tiredness, mental fatigue, stress, anxiety, depression, cognitive brain disorders, etc.) and measure quality of life and / or one or more other behavioral parameters. The specific condition and / or behavioral parameters indicative of quality of life will likely vary based on medical treatment, concomitant medical conditions, etc. For example, a tumor patient may experience "chemo brain" as a side effect of cancer treatment, and the voice module 104 may detect the affected patient's cognitive thinking skills based on an analysis of the patient's voice, which indicates the presence of "chemo brain," which reduces the patient's quality of life.
[0032] For example, an anti-cancer drug therapy may be effective but may be detrimental to the quality of life of the individual receiving the anti-cancer drug therapy. A patient receiving anti-cancer drug therapy may survive, for example, five years after initial diagnosis, but the five years the patient is receiving treatment may be miserable due to changes in quality of life caused by the anti-cancer drug therapy. This anti-cancer drug therapy would not have been identified or treated as a result had it not been qualified by voice module 104 or otherwise. In this example, a new drug therapy being tested may have similar or slightly lower efficacy but a much higher quality of life, but would not have been approved or selected for use if quality of life had not been measured by voice module 104 or considered as a factor in drug trials and / or medical research.
[0033] Instead of subjectively measuring quality of life and / or behavioral outcomes from receiving a treatment or medication using a questionnaire or similar tool, in certain embodiments, the voice module 104 can objectively identify one or more quality of life changes in the patient using biomarkers or other indicators in the voice data from the patient. In one embodiment, using the voice module 104 to identify quality of life and / or other behavioral parameters related to a medication or cancer treatment is a biomarker data-driven, objective tool. As described in more detail below, the voice module 104 can assess quality of life, medical condition, etc. based on an analysis of the user's responses to one or more prompts. For example, in the "chemo brain" example described above, the voice module 104 can present the user with a series of prompts selected to assess the current state of one or more symptoms associated with "chemo brain," such as memory loss. To monitor for memory loss, in certain embodiments, the audio module 104 may audibly list words and / or numbers to the user and ask the user to repeat them, may display a series of photographs to the user and ask the user to repeat a description of the series of photographs, etc., and changes in the accuracy of the user's responses over time can be monitored to indicate memory loss and a decline in quality of life.
[0034] In one embodiment, system 100 includes one or more hardware devices 102. Hardware device 102 and / or one or more backend servers 108 (e.g., computing devices, information processing devices, etc.) may include one or more of a desktop computer, a laptop computer, a mobile device, a tablet computer, a smart phone, a set-top box, a gaming console, a smart TV, a smart watch, a fitness band, a head-mounted optical display (e.g., a virtual reality headset, smart glasses, etc.), an HDMI or other electronic display dongle, a personal digital assistant, and / or other computing device comprising a processor (e.g., a central processing unit (CPU), a processor core, a field programmable gate array (FPGA) or other programmable logic, an application specific integrated circuit (ASIC), a controller, a microcontroller, and / or other semiconductor integrated circuit device), volatile memory, and / or a non-volatile storage medium. In particular embodiments, hardware device 102 communicates with one or more backend servers 108 over data network 106, described below. In yet other embodiments, hardware device 102 may execute various programs, program code, applications, instructions, functions, etc.
[0035] In various embodiments, the audio module 104 may be embodied as hardware, software, or some combination of hardware and software. In one embodiment, the audio module 104 may comprise executable program code stored on a non-transitory computer-readable storage medium for execution on a processor of the hardware device 102, the backend server 108, etc. For example, the audio module 104 may be embodied as executable program code executing on one or more of the hardware device 102, the backend server 108, a combination of one or more of the foregoing, etc. In such an embodiment, various modules that perform the operations of the audio module 104, as described below, may be located on the hardware device 102, the backend server 108, a combination of the two, etc.
[0036] In various embodiments, the audio module 104 may be embodied as a hardware appliance that can be installed or deployed on the backend server 108, on the user's hardware device 102 (e.g., a dongle, a protective case for the phone 102 or tablet 102 containing one or more semiconductor integrated circuit devices that communicate with the phone 102 or tablet 102 wirelessly and / or through a data port such as a USB or proprietary communications port, or other peripheral device), or anywhere on the data network 106 and / or elsewhere co-located with the user's hardware device 102. In particular embodiments, the audio module 104 may comprise a hardware device such as a secure hardware dongle or other hardware appliance device (e.g., a set-top box, a network appliance, etc.). The hardware device may be attached to other hardware devices 102, such as laptop computers, servers, tablet computers, smartphones, etc., by either a wired connection (e.g., USB connection) or a wireless connection (e.g., Bluetooth, Wi-Fi, Near Field Communications (NFC), etc.). The hardware device may be attached to an electronic display device (e.g., a television or monitor using an HDMI port, a DisplayPort port, a Mini DisplayPort port, a VGA port, a DVI port, etc.), operate substantially independently on a data network 106, or the like. The hardware appliance of the audio module 104 may include a power interface, a wired and / or wireless network interface, a graphical interface (e.g., a graphics card and / or a GPU with one or more display ports) that outputs to a display device, and / or a semiconductor integrated circuit device configured to perform the functions described herein with respect to the audio module 104, as described below.
[0037] In such embodiments, the audio module 104 may comprise a semiconductor integrated circuit device (e.g., one or more chips, dies, or other discrete logic hardware), such as a field programmable gate array (FPGA) or other programmable logic, firmware for the FPGA or other programmable logic, microcode for execution on a microcontroller, an application specific integrated circuit (ASIC), a processor, a processor core, etc. In one embodiment, the audio module 104 may be implemented on a printed circuit board with one or more electrical wiring or connections (e.g., electrical wiring or connections to volatile memory, non-volatile storage media, network interfaces, peripheral devices, graphical / display interfaces, etc.). The hardware appliance may also include one or more pins, pads, or other electrical connections (e.g., in communication with one or more electrical wires on a printed circuit board, etc.) configured to send and receive data, as well as one or more hardware circuits and / or other electrical circuits configured to perform various functions of the audio module 104.
[0038] The semiconductor integrated circuit device or other hardware appliance of the audio module 104, in certain embodiments, includes and / or is communicatively coupled to one or more volatile memory media, which may include, but are not limited to, random access memory (RAM), dynamic RAM (DRAM), cache, etc. In one embodiment, the semiconductor integrated circuit device or other hardware appliance of the audio module 104 includes and / or is communicatively coupled to one or more non-volatile memory media. Non-volatile memory media may include, but are not limited to, NAND flash memory, NOR flash memory, nano random access memory (nanoRAM or NRAM), nanocrystal wire-based memory, silicon-oxide based sub-10 nanometer process memory, graphene memory, silicon-oxide-nitride-oxide-silicon (SONOS), resistive RAM (RRAM®), programmable metallization cell (PMC), conductive-bridging RAM (CBRAM), magnetoresistive RAM (MRAM), dynamic RAM (DRAM), phase-change RAM (PRAM or PCM), magnetic storage media (e.g., hard disk, tape), optical storage media, etc. In one embodiment, the data network 106 includes a digital communications network that transmits digital communications. The data network 106 may include a wireless network such as a wireless cellular network, a local wireless network such as a Wi-Fi network, a Bluetooth® network, a near field communication (NFC) network, an ad hoc network, or the like.The data network 106 may include a wide area network (WAN), a storage area network (SAN), a local area network (LAN), a fiber optic network, the Internet, or other digital communications network. The data network 106 may also include one or more servers, routers, switches, and / or other networking equipment. The data network 106 may also include one or more computer-readable storage media, such as hard disk drives, optical drives, non-volatile memory, RAM, etc.
[0039] In one embodiment, the one or more backend servers 108 may include one or more network-accessible computing systems, such as one or more web servers hosting one or more web sites, corporate intranet systems, application servers, application programming interface (API) servers, authentication servers, etc. The backend servers 108 may include one or more servers located remotely from the hardware device 102. The backend servers 108 may include at least a portion of the audio module 104, may configure the hardware of the audio module 104, may store executable program code for the audio module 104 on one or more non-transitory computer-readable storage media, and / or may otherwise perform one or more of the various operations of the audio module 104 described herein for shared content tracking and attribution.
[0040] Figure 1B illustrates an example system 109 that uses a person's voice to diagnose a medical condition. Figure 1B includes a medical condition diagnosis service 140 that receives the person's voice data and processes the voice data to determine whether the person has a medical condition. For example, the medical condition diagnosis service 140 may process the voice data to compute a "yes" or "no" determination regarding whether the person has a medical condition, or a score indicating the probability or likelihood that the person has a medical condition and / or the severity of the condition.
[0041] As used herein, a diagnosis relates to any determination as to whether a person may have a medical condition, or any determination as to the possible severity of a medical condition. A diagnosis can include any form of assessment, conclusion, opinion, or judgment regarding a medical condition. In some cases, a diagnosis may be inaccurate, and a person diagnosed with a medical condition may not actually have the condition.
[0042] The medical condition diagnostic service 140 can receive the person's voice data using any suitable technique. For example, the person may speak into the mobile device 110, which may record the voice and transmit the recorded voice data to the medical condition diagnostic service 140 over the network 130. Any suitable technique and any suitable network can be used to transmit the voice data recorded by the mobile device 110 to the medical condition diagnostic service 140. For example, an application or "app" may be installed on the mobile device 110 and may use REST (representational state transfer) API (application programming interface) calls to transmit the voice data over the Internet or a mobile telephone network. In another example, a medical provider may have a medical provider computer 120 that can be used to record the person's voice and transmit the voice data to the medical condition diagnostic service 140.
[0043] In some implementations, medical condition diagnostic service 140 may be installed on mobile device 110 or healthcare provider computer 120, eliminating the need to transmit audio data over a network. The example of Figure 1B is not limiting, and any suitable technique may be used to transmit audio data for processing by the mathematical model.
[0044] The output of the medical condition diagnostic service 140 can then be used for any suitable purpose, for example, to present information to the person who provided the audio data or to a medical professional treating this person.
[0045] 2 is an example system 200 for processing audio data with a mathematical model to perform medical diagnosis. When processing the audio data, features can be calculated from the audio data, and these features can then be processed with the mathematical model. Any suitable type of feature can be used.
[0046] The features can include acoustic features, which are any features computed from the audio data without involving or relying on performing speech recognition on the audio data (e.g., the acoustic features do not use information about the speech data uttered in the audio data). For example, the acoustic features may include mel-frequency cepstral coefficients, perceptual linear prediction features, jitter, or shimmer.
[0047] The features can include linguistic features, where the linguistic features are computed using the results of speech recognition. For example, the linguistic features may include speaking rate (e.g., number of vowels or syllables per second), number of pause fillers (e.g., "ums" and "eres"), word difficulty (e.g., less commonly used words), or the portion of the phonetic sequence of the word following the filler.
[0048] 2, speech data is processed by an acoustic feature computation component 210 and a speech recognition component 220. The acoustic feature computation component 210 can compute acoustic features, such as any of the acoustic features described herein, from the speech data. The speech recognition component 220 can perform automatic speech recognition on the speech data using any suitable technique (e.g., Gaussian mixture models, acoustic modeling, language modeling, and neural networks).
[0049] Because the speech recognition component 220 may use acoustic features when performing speech recognition, some of the processing of these two components may overlap, and other configurations are possible. For example, the acoustic features component 210 could calculate the acoustic features required by the speech recognition component 220, thus eliminating the need for the speech recognition component 220 to calculate acoustic features altogether.
[0050] The linguistic feature computation component 230 can receive speech recognition results from the speech recognition component 220 and process the speech recognition results to determine linguistic features, such as any of the linguistic features described herein. The speech recognition features can be in any suitable format and can include any suitable information. For example, the speech recognition results can include a word lattice that includes multiple possible word sequences, information about fillers, and the timing of words, syllables, vowels, fillers, or any other units of speech.
[0051] The medical condition classifier 240 may process the acoustic and linguistic features through a mathematical model and output one or more diagnostic scores indicating whether the person has a medical condition, such as a score indicating the probability or likelihood that the person has the medical condition and / or a score indicating the severity of the medical condition. The medical condition classifier 240 may use any suitable technique, such as a classifier implemented with a support vector machine or a neural network such as a multilayer perceptron.
[0052] The performance of the condition classifier 240 may depend on the features computed by the acoustic feature computation component 210 and the linguistic feature computation component 230. Furthermore, one set of features that provides correct processing for one condition may not provide correct processing for another condition. For example, speech difficulties may be an important feature for diagnosing Alzheimer's disease, but may not be useful for determining whether a person has a concussion. As another example, features related to the pronunciation of vowels, syllables, or words may be important for Parkinson's disease, but may be less important for other conditions. Therefore, a technique is needed to determine a first set of features that provides correct processing for a first condition, and the process may need to be repeated to determine a second set of features that provides correct processing for a second condition.
[0053] In some implementations, medical condition classifier 240 may use other features, which may be referred to as non-speech features, in addition to acoustic and linguistic features. For example, features may be derived from or calculated from a person's demographic information (e.g., gender, age, location), information from medical history (e.g., weight, most recent blood pressure reading, or previous diagnoses), or any other suitable information.
[0054] The selection of features for diagnosing a medical condition may become even more important in situations where the amount of training data for training a mathematical model is relatively small. For example, training a mathematical model for diagnosing concussion may require training data that includes speech data from a large number of individuals immediately after experiencing a concussion. Such data may exist in small amounts, and obtaining additional examples of such data may require significant periods of time.
[0055] When training a mathematical model, the smaller the amount of training data, the greater the risk of overfitting. In this case, the mathematical model may adapt to specific training data, but the small amount of training data may cause the model to be unable to process new data correctly. For example, a model may be able to detect all concussions in the training data, but have a high error rate when processing production data of people at risk of concussion.
[0056] One technique for preventing overfitting when training a mathematical model is to reduce the number of features used to train the mathematical model. The amount of training data required to train the model without overfitting increases as the number of features increases. Therefore, by using fewer features, it becomes possible to build a model using a smaller amount of training data.
[0057] When a model needs to be trained with a small number of features, it becomes increasingly important to select features that will enable the model to perform correctly. For example, when a large amount of training data is available, the model can be trained using hundreds of features, and the likelihood that the right features will be used is greater. Conversely, when only a small amount of training data is available, the model may be trained using as few as 10 features, and it becomes increasingly important to select the features that are most important for diagnosing the medical condition.
[0058] We now present examples of features that can be used to diagnose a medical condition. Acoustic features can be computed using short-time segment features. When processing audio data, the duration of this audio data may vary. For example, some audio may be one or two seconds long, while other audio may be several minutes or longer. For consistency when processing audio data, it is useful to process it in short-time segments (sometimes called frames). For example, each short-time segment may be 25 milliseconds long, with segments progressing in 10-millisecond increments and allowing a 15-millisecond overlap between two consecutive segments.
[0059] The following are non-limiting examples of short-term segmental features: spectral features (such as mel-frequency cepstral coefficients or perceptual linear prediction), prosodic features (features like utterance tone, energy, probability), speech quality features (features like jitter, jitter of jitter, fluctuation, or harmonic to noise ratio), entropy (where entropy can be calculated posteriorly from acoustic models trained on natural speech data, e.g., to capture how accurately an utterance was pronounced).
[0060] Short-term segment features can be combined to compute acoustic features for speech. For example, a two-second speech sample can generate 200 short-term segment features for pitch, which can be combined to compute one or more acoustic features for pitch.
[0061] Any suitable technique can be used to combine short-term segment features to compute acoustic features for an audio sample. In some embodiments, the acoustic features can be calculated using statistics of short-time segment features (e.g., arithmetic mean, standard deviation, skewness, kurtosis, 1st quartile, 2nd quartile, 3rd quartile, 2nd quartile minus 1st quartile, 3rd quartile minus 2nd quartile, 0.01 percentile, 0.99th percentile, 0.99th percentile minus 0.01 percentile), the percentage of short-time segments whose values are higher than a threshold (e.g., the threshold is 75% of the range plus the minimum), the percentage of segments whose values are higher than a threshold (e.g., the threshold is 90% of the range plus the minimum), the slope of a linear approximation of the value, the offset of a linear approximation of the value, a linear error calculated as the difference between the linear approximation and the actual value, or a quadratic error calculated as the difference between the linear approximation and the actual value. In some implementations, acoustic features can be calculated as i-vectors or identity vectors of short-term segmental features. The identity vectors can be calculated using any suitable technique, such as performing an identity-to-vector transformation using factorial analysis techniques and Gaussian mixture models.
[0062] The following are non-limiting examples of linguistic features: speaking rate, such as calculated by dividing the duration of all spoken words by the number of vowels, or any other suitable measure of speaking rate; the number of fillers, which may indicate hesitation in speech, calculated by (1) dividing the number of fillers by the duration of spoken words, or (2) dividing the number of fillers by the number of spoken words; and a measure of word difficulty or unusual word use. For example, word difficulty can be calculated using statistics of 1-gram probabilities of spoken words, such as by classifying words according to word frequency percentiles (e.g., 5%, 10%, 15%, 20%, 30%, or 40%). The speech parts of words following filler words, such as (1) the number of each part-of-speech class divided by the number of words spoken, or (2) the number of each part-of-speech class divided by the sum of the number of all parts-of-speech.
[0063] In some implementations, the linguistic features may also include determining whether a person answered a question correctly. For example, a person may be asked what year it is or who the president of the United States is. The person's speech can be processed to determine what the person said in response to the question and further to determine whether the person answered the question correctly.
[0064] To train a model to diagnose medical conditions, a corpus of training data can be collected that includes speech examples that indicate a person's diagnosis, such as whether they have no concussion, a mild, moderate, or severe concussion.
[0065] Figure 3 shows an example of a training corpus containing audio data for training a model to diagnose concussion. For example, in the table of Figure 3, rows may correspond to entries in a database. In this example, each entry includes an identifier for a person, a known diagnosis for that person (e.g., no concussion, mild, moderate, or severe concussion), an identifier for a prompt or question presented to the person (e.g., "How are you feeling today?"), and a filename for a file containing the audio data. The training data may be stored in any suitable format using any suitable storage technology.
[0066] The training corpus can store representations of human speech using any suitable format. For example, the speech data items in the training corpus may include digital samples of an audio signal received at a microphone, or may include processed versions of the audio signal, such as Mel-frequency cepstral coefficients.
[0067] A single training corpus may contain speech data for multiple medical conditions, or a separate training corpus may be used for each medical condition (e.g., a first training corpus for concussion and a second training corpus for Alzheimer's disease). A separate training corpus may be used to store speech data from people with unknown or undiagnosed medical conditions, since this training corpus can be used to train models for multiple medical conditions.
[0068] Figure 4 shows an example of stored prompts that can be used to diagnose a medical condition. Each prompt can be presented to a person, either by a person (e.g., a medical professional) or a computer, to obtain the person's speech in response to the prompt. Each prompt can have a prompt identifier so that it can be cross-referenced with prompt identifiers in a training corpus. The prompts in Figure 4 may be stored using any suitable storage technique, such as a database.
[0069] 5 illustrates an example system 500 that can be used to select features for training a mathematical model to diagnose a medical condition, and then use the selected features to train the mathematical model. System 500 can be used multiple times to select features for different medical conditions. For example, a first use of system 500 can select features for diagnosing a concussion, and a second use of system 500 can select features for diagnosing Alzheimer's disease.
[0070] 5 includes a training corpus 510 of speech data items for training a mathematical model for diagnosing a medical condition. The training corpus 510 may include any suitable information, such as speech data of a plurality of people with and without a medical condition, labels indicating whether a person has a medical condition, and any other information described herein.
[0071] The acoustic feature computation component 210, speech recognition component 220, and linguistic feature computation component 230 can be implemented as described above to compute acoustic and linguistic features for the speech data in the training corpus. The acoustic feature computation component 210 and linguistic feature computation component 230 can compute multiple features so that best performing features can be determined. This is in contrast to Figure 2, where these components are used in a production system and therefore only need to compute previously selected features.
[0072] The feature selection score calculation component 520 may calculate a selection score for each feature (which may be an acoustic feature, a linguistic feature, or any other feature described herein). To calculate a selection score for a feature, a pair of numbers may be created for each speech data item in the training corpus. The first number in the pair is the value of the feature, and the second number in the pair is an indicator of a medical condition diagnosis. The value of the indicator of a medical condition diagnosis may have two values (e.g., 0 if the person does not have the medical condition and 1 if the person has the medical condition) or may have more than one number (e.g., a real number between 0 and 1, or multiple integers indicating the likelihood or severity of the medical condition).
[0073] Thus, for each feature, a pair of numerical values can be obtained for each speech data item in the training corpus. Figures 6A and 6B show two conceptual plots of the numerical value pairs for the first feature and the second feature. For Figure 6A, there does not appear to be a pattern or correlation between the values of the first feature and the corresponding diagnostic values, while for Figure 6B, there appears to be a pattern or correlation between the values of the second feature and the diagnostic values. Therefore, it can be concluded that the second feature is likely to be a useful feature for determining whether a person has a medical condition, while the first feature is not.
[0074] The feature selection score calculation component 520 can calculate a selection score for the feature using the feature value and diagnostic value pairs. The feature selection score calculation component 520 can calculate any suitable score that indicates a pattern or correlation between the feature values and the diagnostic values. For example, the feature selection score calculation component 520 can calculate a Rand index, an adjusted Rand index, mutual information, adjusted mutual information, a Pearson correlation, an absolute Pearson correlation, a Spearman correlation, or an absolute Spearman correlation.
[0075] The selection score can indicate the usefulness of the feature in detecting a pathology, for example, a high selection score may indicate that a feature should be used when training a mathematical model, and a low selection score may indicate that the feature should not be used when training a mathematical model.
[0076] The feature stability determination component 530 can determine whether a feature (which may be an acoustic feature, a linguistic feature, or any other feature described herein) is stable or unstable. To perform the stability determination, the audio data items can be divided into groups, sometimes referred to as folds. For example, the audio data items may be divided into five folds. In one implementation, the audio data items may be divided into folds such that each fold has an approximately equal number of audio data items for different gender and age groups.
[0077] Statistics for each fold can be compared to statistics for other folds. For example, for the first fold, the median (or mean, or any other statistical value relating to the center or middle of a distribution) feature value (denoted M1) can be determined. Statistics can also be calculated for combinations of other folds. For example, the median (denoted M0) of the feature values and a statistical measure of the variability of the feature values, such as the interquartile range, variance, or standard deviation (denoted V0), can be calculated for combinations of several other folds. If the median for the first fold is too different from the median for the second fold, the feature can be determined to be unstable. For example,
[0078]
number
[0079] If , the feature can be determined to be unstable. where C is the scaling factor. This process can then be repeated for each of the other folds. For example, as described above, the median of the second fold can be compared to the median and variability of the other folds.
[0080] In one embodiment, after comparing each fold with the other folds, if the median of each fold is not too far from the medians of the other folds, the feature can be determined to be stable. Conversely, if the median of any fold is too far from the medians of the other folds, the feature can be determined to be unstable.
[0081] In some implementations, the feature stability determination component 530 may output a Boolean value for each feature to indicate whether the feature is stable or not. In some implementations, the stability determination component 530 may also output a stability score for each feature. For example, the stability score may be calculated as the largest distance (e.g., Mahalanobis distance) between the median of one fold and the median of another fold.
[0082] The feature selection computation component 540 can receive the selection scores from the feature selection score computation component 520 and the stability determination from the feature stability determination component 530 and select a subset of features to be used to train the mathematical model. The feature selection component 540 can select the features that have the highest selection scores and are sufficiently stable.
[0083] In some implementations, the number of features to be selected (or the maximum number of features to be selected) may be preset. For example, the number N may be determined based on the amount of training data, and N features may be selected. The feature selection may be determined by removing unstable features (e.g., features determined to be unstable or features with stability scores below a threshold), and then selecting the N features with the highest selection scores.
[0084] In some implementations, the number of features selected may be based on the selection score and a stability determination, for example, feature selection may be determined by removing unstable features and then selecting all features with a selection score above a threshold.
[0085] In some embodiments, the selection score and stability score may be combined when selecting features. For example, a combined score may be calculated for each feature (such as by adding or multiplying the selection score and stability score for the feature), and this combined score may be used to select features.
[0086] The selected features can then be used by model training component 550 to train a mathematical model. For example, model training component 550 can iterate through the speech data items of the training corpus to obtain selected features for the speech data items, and then use the selected features to train the mathematical model. In some implementations, dimensionality reduction techniques, such as principal component analysis or linear discriminant analysis, may be applied to the selected features as part of model training. Any suitable mathematical model can be trained, such as any of the mathematical models described herein.
[0087] In some embodiments, other techniques, such as wrapper methods, may be used for feature selection or may be used in combination with the feature selection techniques described above. Wrapper methods may select a set of features, train a mathematical model using this selected set of features, and then use the trained model to evaluate the performance of the set of features. When the number of possible features is relatively small and / or training time is relatively short, all possible sets of features may be evaluated and the best performing set may be selected. When the number of possible features is relatively large and / or training time is a significant factor, optimization techniques may be used to iteratively find a set of features that performs well. In some embodiments, system 500 may be used to select a set of features, and then a wrapper method may be used to select a subset of these features as the final set of features.
[0088] Figure 7 is a flowchart of an example embodiment for selecting features for training a mathematical model for diagnosing a medical condition. In Figure 7 and other flowcharts herein, the order of steps is exemplary; other orders are possible, not all steps are required, steps may be combined (in whole or in part) or subdivided, and some steps may be omitted or other steps may be added in some embodiments. Any of the methods described by the flowcharts described herein may be implemented, for example, by any of the computers or systems described herein.
[0089] In step 710, a training corpus of speech data items is obtained. The training corpus may include any other suitable information, such as a representation of an audio signal of a person's speech, a medical diagnostic indication of the person from whom the speech was obtained, and any of the information described herein.
[0090] In step 720, speech recognition results are obtained for each speech data item in the training corpus. The speech recognition results may be pre-computed and stored with the training corpus, or may be stored elsewhere. The speech recognition results may include any suitable information, such as transcripts, a list of the highest scoring transcripts (e.g., an N-best list), a lattice of possible transcriptions, and timing information, such as start and end times of words, fillers, or other speech units.
[0091] In step 730, acoustic features are computed for each speech data item in the training corpus. The acoustic features may include any features computed without using speech recognition results for the speech data item, such as any of the acoustic features described herein. The acoustic features may include or be computed from data used in the speech recognition process (e.g., mel-frequency cepstral coefficients or perceptual linear predictors), but the acoustic features do not use speech recognition results, such as information about words or fillers present in the speech data item.
[0092] Linguistic features are computed for each speech data item in the training corpus in step 740. The linguistic features may include any features computed using speech recognition results, such as any of the linguistic features described herein.
[0093] In step 750, a feature selection score is calculated for each acoustic and linguistic feature. To calculate the feature selection score for a feature, the value of the feature for each speech data item in the training corpus may be used along with other information, such as known diagnostic values corresponding to the speech data item. The feature selection score may be calculated using any of the techniques described herein, such as by calculating absolute Pearson correlation. In some implementations, feature selection scores may be calculated for other features as well, such as features related to a person's demographic information.
[0094] In step 760, the feature selection scores are used to select a number of features. For example, a number of features with the highest selection scores may be selected. In some implementations, a stability measure may be calculated for each feature, and both the feature selection scores and the stability measure may be used to select a number of features, such as by using any of the techniques described herein.
[0095] In step 770, a mathematical model is trained using the selected features. Any suitable mathematical model may be trained, such as a neural network or a support vector machine. After the mathematical model is trained, it may be deployed in a production system, such as speech module 104, system 109, etc., of FIG. 1B, to perform diagnosis of a medical condition.
[0096] The steps of Figure 7 can be performed in a variety of ways. For example, in one embodiment, steps 730 and 740 may be performed in a loop, repeatedly performed for each speech data item in the training corpus. In a first iteration, acoustic and linguistic features may be calculated for a first speech data item, in a second iteration, acoustic and linguistic features may be calculated for a second speech data item, and so on.
[0097] When using the deployed model to diagnose a medical condition, a series of prompts or questions can be uttered to the person to obtain speech from the person being diagnosed. Any suitable prompts can be used, such as any of the prompts in Figure 4. After features are selected as described above, the prompts can be selected such that they provide useful information about the selected features.
[0098] For example, suppose the selected feature is pitch. Pitch has been determined to be a useful feature for diagnosing medical conditions, but some prompts may be better than others at yielding a useful pitch feature. Very short utterances (e.g., yes / no answers) may not provide enough data to accurately calculate pitch, so prompts that generate longer responses may be more useful in obtaining information about pitch.
[0099] As another example, suppose the selected feature is word difficulty. Word difficulty has been determined to be a useful feature for diagnosing medical conditions, but some prompts may be better than others at yielding a useful word difficulty feature. A prompt asking a user to read a presented passage generally results in the words in the passage being spoken, and therefore the word difficulty feature will have the same value each time the prompt is presented. This means that the prompt is not useful for obtaining information about word difficulty. In contrast, an open-ended question such as "Tell me about your day" may result in greater vocabulary diversity in responses and therefore provide more useful information about word difficulty.
[0100] Selecting a set of prompts can also improve the performance of a system for diagnosing a medical condition and provide a better experience for the person being assessed. Using the same set of prompts for each person can enable a system for diagnosing a medical condition to achieve more accurate results because data collected from multiple people is easier to compare than if different prompts were used for each person. Furthermore, using a fixed set of prompts makes it easier to predict a person's ratings and the desired duration of the rating appropriate for assessing a medical condition. For example, to assess whether a person has Alzheimer's disease, it is acceptable to use more prompts to collect a greater amount of data. However, to assess whether a person suffered a concussion at a sporting event, it may be necessary to use fewer prompts to obtain results more quickly.
[0101] In some implementations, prompts may be selected by calculating a prompt selection score. The training corpus may have multiple speech data items for a prompt, or even many speech data items. For example, the training corpus may include examples of prompts used by different people, or the same prompt may be used multiple times by the same person.
[0102] FIG. 8 is a flowchart of an example embodiment of selecting a prompt for use with a deployed model to diagnose a medical condition. Steps 810 through 840 may be performed for each prompt (or subset of prompts) in the training corpus to calculate a prompt selection score for each prompt.
[0103] In step 810, a prompt is obtained, and in step 820, a speech data item corresponding to the prompt is obtained from the training corpus. A medical diagnostic score is calculated for each audio data item corresponding to the prompt, in step 830. For example, the medical diagnostic score for an audio data item may be a numerical value output by a mathematical model (e.g., the mathematical model trained in FIG. 7) that indicates the likelihood that a person has a medical condition and / or the severity of that condition.
[0104] In step 840, the calculated medical diagnostic score is used to calculate a prompt selection score for the prompt. Calculating the prompt selection score may be similar to calculating the feature selection score, as described above. For each audio data item corresponding to a prompt, a pair of numerical values may be obtained. For each pair, the first numerical value of the pair may be the medical diagnostic score calculated from the audio data item, and the second numerical value of the pair may be a known medical condition diagnosis for the person (e.g., known that the person has a medical condition or is indicative of the severity of the condition). Plotting these pairs of numerical values results in a plot similar to FIG. 6A or 6B, where, depending on the prompt, there may or may not be a pattern or correlation between the pairs of numerical values.
[0105] The prompt selection score for a prompt may include any score that indicates a pattern or correlation between the calculated medical diagnostic score and a known medical condition diagnosis. For example, the prompt selection score may include a Rand index, an adjusted Rand index, mutual information, adjusted mutual information, a Pearson correlation, an absolute Pearson correlation, a Spearman correlation, or an absolute Spearman correlation.
[0106] In step 850, it is determined whether any more prompts remain to be processed. If so, processing can proceed to step 810 where additional prompts can be processed. If all prompts have been processed, processing can proceed to step 860.
[0107] In step 860, the prompt selection scores are used to select multiple prompts. For example, a number of prompts with the highest prompt selection scores may be selected. In some implementations, a stability determination may be calculated for each prompt, and multiple prompts may be selected using both the prompt selection score and the prompt stability score, such as by using any of the techniques described herein.
[0108] The selected prompts are used with the deployed medical condition diagnosis service in step 870. For example, when diagnosing a person, the selected prompts can be presented to the person and the person's voice can be obtained in response to each of the prompts.
[0109] In some implementations, other techniques, such as wrapper methods, may be used for prompt selection or may be used in combination with the prompt selection techniques presented above. In some implementations, a set of prompts may be selected using the process of Figure 8, and then a subset of these prompts may be selected using wrapper methods as the final set of features.
[0110] In some embodiments, a person involved in creating the medical condition diagnostic service may assist in selecting the prompts. This person may use their knowledge or experience to select prompts based on the selected characteristic. For example, if the selected characteristic is word difficulty, this person may review the prompts and select those that are likely to provide useful information regarding word difficulty. This person may select one or more prompts that are likely to provide useful information for each of the selected characteristics.
[0111] In one embodiment, the person can review the prompts selected by the process of Figure 8 and add or remove prompts to improve the performance of the medical condition diagnosis system. For example, two prompts may each provide useful information about the difficulty of a word, but the information provided by these two prompts may be so redundant that using both prompts may not provide a significant benefit over using only one of them.
[0112] In some embodiments, after prompt selection, a second mathematical model appropriate for the selected prompt may be trained. The mathematical model trained in FIG. 7 may process a single utterance (in response to a prompt) to generate a medical diagnostic score. The process of performing a diagnosis may include processing multiple utterances corresponding to multiple prompts, and then processing each of the utterances through the mathematical model of FIG. 7 to generate multiple medical diagnostic scores. It may be necessary to combine multiple medical diagnostic scores in some way to determine an overall medical diagnosis. Thus, the mathematical model trained in FIG. 7 may not be appropriate for a selected set of prompts.
[0113] When the selected prompts are used in a session to diagnose a person, each of the prompts can be presented to the person to obtain an utterance corresponding to each of the prompts. Instead of processing the utterances separately, the utterances can be processed simultaneously by the model to generate a medical diagnostic score. Thus, the model can adapt to the selected prompts because it is trained to simultaneously process utterances corresponding to each of the selected prompts.
[0114] Figure 9 is a flow chart of an example embodiment for training a mathematical model appropriate for a set of selected prompts. In step 910, a first mathematical model is obtained, such as by using the process of Figure 7. In step 920, a plurality of prompts are selected using the first mathematical model, such as by the process of Figure 8.
[0115] In step 930, a second mathematical model is trained to simultaneously process multiple speech data items corresponding to multiple selected prompts to generate a medical diagnosis score. When training the second mathematical model, a training corpus including a session with speech data items corresponding to each of the multiple selected prompts can be used. When training this mathematical model, inputs to the mathematical model may be fixed to the speech data items from the session and corresponding to each of the selected prompts. The output of the mathematical model may be fixed to a known medical diagnosis.
[0116] The parameters of this model can then be trained to optimally process audio data items to simultaneously produce a medical diagnostic score. Any suitable training technique can be used, such as stochastic gradient descent.
[0117] The second mathematical model can then be deployed as part of a medical condition diagnosis service, such as the speech module 104, the service of Figure 1. Because the second mathematical model is trained to process the utterances simultaneously, rather than individually, the second mathematical model can outperform the first mathematical model, i.e., the training can combine information from all the utterances to produce a more accurate medical condition diagnosis score.
[0118] Figure 10 illustrates components of one embodiment of a computing device 1000 for implementing any of the techniques described above. While the components are shown in Figure 10 as being on one computing device, the components may be distributed across multiple computing devices, such as in a system of computing devices that includes end-user computing devices (e.g., smart phones or tablets) and / or server computing devices (e.g., cloud computing).
[0119] The computing device 1000 may include any components typical of a computing device, such as volatile or non-volatile memory 1010, one or more processors 1011, and one or more network interfaces 1012. The computing device 1000 may also include any input and output components, such as a display, a keyboard, and a touch screen. The computing device 1000 may also include various components or modules that provide specific functionality, and these components or modules may be implemented in software, hardware, or a combination thereof. Various example components are described below as an example implementation; however, other implementations may include additional components or may exclude some of the components described below.
[0120] The computing device 1000 may have an acoustic feature computation component 1021 that can compute acoustic features for an audio data item as described above. The computing device 1000 may have a linguistic feature computation component 1022 that can compute linguistic features for an audio data item as described above. The computing device 1000 may have a speech recognition component 1023 that can generate speech recognition results for an audio data item as described above. The computing device 1000 may have a feature selection score computation component 1031 that can compute selection scores for features as described above. The computing device 1000 may have a feature stability score computation component 1032 that can perform or compute stability scores as described above. The computing device 1000 may have a feature selection component 1033 that can select features using the selection scores and / or stability determinations as described above. The computing device 1000 may have a prompt selection score computation component 1041 that can compute selection scores for prompts as described above. The computing device 1000 can have a prompt stability score calculation component 1042 that can make a stability determination or calculate a stability score as described above. The computing device 1000 can have a prompt selection component 1043 that can select a prompt using the selection score and / or the stability determination as described above. The computing device 1000 can have a model training component 1050 that can train a mathematical model as described above. The computing device 1000 can have a medical condition diagnosis component 1060 that can process the audio data items to determine a medical diagnosis score as described above.
[0121] Computing device 1000 may include or have access to various data stores, such as training corpus data store 1070. The data stores may use any well-known storage technology, such as files, relational or non-relational databases, or any non-transitory computer-readable medium.
[0122] 11 illustrates one embodiment of the voice module 104. In particular embodiments, the voice module 104 may be substantially similar to one or more of the device voice module 104a and / or the back-end voice module 104b, as described above with respect to FIG. 1A. In the illustrated embodiment, the voice module 104 includes a query module 1102, a response module 1104, a detection module 1106, and an interface module 1108.
[0123] In one embodiment, the query module 1102 asks and / or queries the user with one or more questions, prompts, requests, etc. In particular embodiments, the query module 1102 may audibly and / or verbally ask the user questions (e.g., using a speaker on the computing device 102, such as an integrated speaker, headphones, Bluetooth speakers, headphones, etc.). For example, certain underlying medical conditions, such as a concussion, may make it difficult for the user to read the questions and / or prompts, and asking the user audibly may simplify and / or expedite diagnosis. In yet other embodiments, the query module 1102 may display one or more questions and / or other prompts to the user (e.g., on an electronic display screen of the computing device 102), or another user (e.g., a coach, parent, medical professional, administrator, etc.) may read the one or more questions and / or other prompts to the user. In various embodiments, one or more questions or prompts may be selected, as described above with respect to prompt selection component 1043, to facilitate diagnosis of one or more medical conditions.
[0124] In particular embodiments, multiple query modules 1102 located on multiple different computing devices 102 may query and / or ask multiple different users. For example, multiple distributed query modules 1102 may collect voice samples for a clinical trial, train machine learning models to diagnose medical conditions, collect lab data to facilitate prompt selection, etc.
[0125] In one embodiment, the query module 1102 queries and / or otherwise queries a user with a predetermined health condition, such as a known health condition, a predetermined stage of a medical condition, etc., to collect one or more baseline audio recordings, training data, or other data. In particular embodiments, the query module 1102 queries and / or otherwise queries a user in response to a potential medical event or other trigger. The query module 1102 may query a user and collect test case audio recordings or other test case data in response to a user requesting a medical evaluation based on data from sensors of the computing device 102, such as a wearable device or mobile device, and / or based on receiving other triggers indicating that an injury has likely occurred, one or more symptoms of a disease have been detected, etc. For example, in response to a crash, fall, accident, and / or other potential concussive event (e.g., in a sporting event or other activity), a user (injured athlete or other person, coach, parent, medical professional, administrator, etc.) may request a medical evaluation (e.g., using the graphical user interface of interface module 1108 to trigger one or more questions from query module 1102, collection of audio data and / or other data from response module 1104, and / or a medical evaluation from detection module 1106, etc.). In certain embodiments, a concussion may include a disruption of brain function resulting from direct or indirect force to the user's head. A concussion may result in headaches, unsteadiness, confusion or other brain disorders, abnormal behavior and / or personality, etc.
[0126] For example, in embodiments in which the medical condition includes a concussion, the query module 1102 may audibly question the user and / or collect sensor data related to the user to detect whether the user's eyes are open, whether the user's eyes are open in response to pain, whether the user's eyes are open in response to audio, whether the user's eyes are open voluntarily, whether the user is able to provide verbal responses, whether the user is making incomprehensible sounds, whether the user is responding to questions or other prompts with inappropriate words, whether the user is confused, whether the user is disoriented, whether the user has little or no motor response, whether the user experiences pain with extension (e.g., arm abduction, forearm supination, etc.), whether the user experiences pain with abnormal flexion (e.g., forearm pronation, flexor posturing, etc.), whether the user's pain is pulling, whether the user can identify the location of pain (e.g., intentionally inducing pain), whether the user complies with verbal / audible commands from the query module 1102, etc. In one embodiment, the query module 1102 may pose some questions to the user being evaluated and / or diagnosed, and pose other questions to an administrator (e.g., a medical professional, coach, parent, trainer, etc.). For example, the query module 1102 may ask the administrator about one or more signs that the administrator may have observed in the user being evaluated and / or diagnosed, such as imbalance, ataxia, disorientation, confusion, memory loss, a blank or vacant expression, visible facial or other injuries, observed physical examination results (e.g., range of motion, flexibility, sensation, strength, balance tests, coordination tests, etc.), and / or other observations.
[0127] In particular embodiments, the query module 1102 may ask and / or prompt the user about the sporting event the user was attending when the potential medical phenomenon occurred, the user's team, the date and / or time, memory test questions, etc. For example, the query module 1102 may ask the user audibly and / or in writing questions such as "Where are we today?", "Is it first half or second half now?", "Who was the last scorer in this game?", "Which team did you play for last week or the last game?", "Did your team win the last game?", "What month is it now?", "What date is today?", "What day of the week is it today?", "What year is it?", "What time is it now?", audibly list words and / or numbers and prompt the user to repeat them, show the user a series of photographs and prompt the user to repeat descriptions of the series of photographs, etc. The one or more questions and / or prompts of the query module 1102 may enable the detection module 1106 to determine a Standardized Concussion Assessment Tool (SCAT) score, a SCAT2 score, a SCAT3 score, a SCAT5 score, a Glasgow Coma Score (GCS), a Maddocks Score, a Concussion Recognition Tool (CRT) score, and / or other concussion scores.
[0128] In one embodiment, the response module 1104 is configured to receive response data (e.g., audio data of verbal responses, text data of typed responses, sensor data, image and / or video data from a camera or other image sensor, touch input from a touch screen and / or touchpad, movement information from an accelerometer and / or gyroscope, etc.) in response to one or more questions and / or other queries from the query module 1102. For example, in particular embodiments, the response module 1104 may use a microphone on the computing device 102 (e.g., a mobile computing device 102 taken to a soccer field, other sporting event, etc.) to record the user's verbal responses (e.g., answers) to one or more questions or other prompts from the query module 1102.
[0129] In one embodiment, the response module 1104 stores the received response data, such as audio recordings, sensor data, etc., on a computer-readable storage medium of the computing device 102, 110, and the interface module 1108 provides the received response data to one or more authorized users and / or makes the received response data accessible for use in other ways so that the detection module 1106 can access and / or process the received response data to diagnose and / or assess a medical condition, train a model to diagnose and / or assess a medical condition, etc. In other embodiments, the response module 1104 can provide the received response data directly to the detection module 1106 to diagnose and / or assess a medical condition (e.g., without otherwise storing the data, without temporarily storing and / or caching the data, etc.).
[0130] The response module 1104 may also separately receive and / or store reference response data (e.g., in response to one or more reference questions or prompts from the query module 1102) and test case response data (e.g., in response to one or more test case questions or prompts from the query module based on a potential medical event, etc.). In particular embodiments, the response module 1104 may receive only the test case response data, and the detection module 1106 may perform an evaluation or other diagnosis of a medical condition based on analysis of the test case data and data from different users (e.g., other users known to have the medical condition, etc.). The response module 1104 may also store and / or organize the received response data in a database and / or other predefined data structure accessible by the detection module 1106, the interface module 1108, etc.
[0131] By storing a user's response history (e.g., baseline response data, test case response data, ratings, scores, etc.), in certain embodiments, the response module 1104 may enable the detection module 1106 to dynamically assess a medical condition for a user in response to a medical event. For example, the response module 1104 may store the user's response data on the mobile computing device 102, on a backend server 108 in communication with the mobile computing device 102 over the data networks 106, 130, etc., enabling the detection module 1106 to make a medical condition assessment decision in the field in response to a potential medical event (e.g., on the sidelines or on the field at a soccer game or other sporting event in response to a potential concussion event, in response to a car accident, etc.).
[0132] In one embodiment, the detection module 1106 is configured to provide a medical condition assessment and / or other diagnosis to the user based on an analysis of the one or more user responses received from the response module 1104. In various embodiments, the detection module 1106 described above may comprise, be in communication with, and / or be substantially similar to the acoustic feature computation component 210, the speech recognition component 220, the linguistic feature computation component 230, and / or the medical condition classifier 240.
[0133] In one embodiment, the detection module 1106 may determine a medical condition assessment or other diagnosis for the user (e.g., to determine changes in the user's voice, changes in the user's responses, etc.) based on both the test case response data for the user and previously received baseline response data for the same user (e.g., whether the user has a medical condition, the likelihood that the user has a medical condition, the estimated severity of the condition, etc.). In other embodiments, the detection module 1106 may determine a medical condition assessment or other diagnosis for the user based on the test case response data for the user and also based on response data for a different user (e.g., a different user who has previously been diagnosed with a medical condition, etc.). In yet other embodiments, the detection module 1106 may determine a medical condition assessment or other diagnosis for the user based on the test case response data for the user, baseline response data for the same user, and response data for a different user, etc.
[0134] In particular embodiments, as described above with respect to the acoustic feature computation component 210, the speech recognition component 220, the linguistic feature computation component 230, and / or the medical condition classifier 240, the detection module 1106 may extract one or more audio features (e.g., acoustic features and / or linguistic features) from the audio recordings (e.g., the baseline response data and / or the test case response data) and input the one or more extracted audio features into a model associated with the medical condition (e.g., a machine learning model such as a Gaussian mixture model, an acoustic model, a language model, a neural network, a deep neural network, a classifier, a support vector machine, a multilayer perceptron, etc.), which may output an assessment or other diagnosis of the medical condition based on the one or more extracted audio features.
[0135] In yet other embodiments, in addition to inputting the extracted audio features into the model to diagnose a medical condition, the detection module 1106 may also input other supplemental data related to the user into the model and diagnose a medical condition based on the results. For example, the detection module 1106 may input sensor data from the user's computing device 102 into the model (e.g., along with the extracted audio features or other audio data) to determine a medical condition assessment or other diagnosis for the user.
[0136] In one embodiment, the detection module 1106 may extract one or more image features from image data (e.g., one or more images, videos, etc. of a user, the user's face, other body parts of the user associated with a medical condition, etc.) from an image sensor, such as a camera, of the computing device 102, and may input the one or more image features into the model (e.g., along with extracted audio features, etc.). In yet other embodiments, the detection module 1106 may make an assessment or other diagnosis based, at least in part, on touch input received from the user on a touch screen, touch pad, etc. of the computing device 102.
[0137] For example, the query module 1102 may provide an interactive video game or the like on an electronic display of the computing device 102, where the interactive video game may be configured to test the user for one or more symptoms of a medical condition (e.g., testing reflexes, agility, reaction time, etc.), and the detection module 1106 may extract one or more features from touch input received from the user during the interactive video game (e.g., a score in the video game, the user's reaction time, touch accuracy metrics for the user, etc.) and input the one or more extracted features into a model for diagnosing the medical condition (e.g., along with one or more extracted audio features, etc.). In particular embodiments, the detection module 1106 may extract one or more features from motion information about the user measured by an accelerometer, gyroscope, and / or other motion sensors of the mobile computing device 102, and input the one or more extracted features into a model for diagnosing the medical condition (e.g., along with one or more extracted audio features, etc.).
[0138] As previously described, in certain embodiments, the detection module 1106 may determine an assessment or other diagnosis for a medical condition, including a neurological condition such as a concussion. In other embodiments, the detection module 1106 may determine an assessment or other diagnosis for a medical condition, including one or more of depression, stress, stroke, cognitive stability, mood, conscience, Alzheimer's disease, Parkinson's disease, cancer, etc.
[0139] In particular embodiments, the detection module 1106 may also be configured to determine a medical condition assessment or other diagnosis based on one or more acoustic features of the received verbal response data, regardless of one or more linguistic features of the received verbal response data (e.g., without any linguistic features, using only one or more predetermined linguistic features, without any automatic speech recognition, etc.). Thus, in an embodiment, the detection module's assessment and / or diagnosis may be independent of the language and / or dialect of the received verbal responses, such that the detection module 1106 may provide an assessment and / or diagnosis to a user in different languages using the acoustic features of the received verbal response data. In other embodiments, the detection module 1106 may assess and / or diagnose a medical condition based on both the acoustic and linguistic features of the received verbal response data.
[0140] In certain embodiments, the detection module 1106 may determine the evaluation and / or diagnosis of a medical condition solely on the user's mobile computing device 102. For example, in an emergency situation, a diagnosis may be required as soon as possible, and there may not be time to upload the recorded verbal responses to the backend server 108 for processing, or the mobile computing device 102 may not have a connection to the data network 106, 130 or may not have a fast enough connection. In one embodiment, the detection module 1106 may use one or more models configured to execute using the processing power, volatile memory capacity, and / or non-volatile storage capacity available on the mobile computing device 102. For example, by limiting matrix multiplications in the model (e.g., using no matrix multiplications, only a predetermined number of matrix multiplications, etc.), even though using additional matrix multiplications could improve the accuracy of the evaluation and / or diagnosis, the model used by the detection module 1106 on the mobile computing device may minimize the size of the classifier (e.g., the required volatile and / or non-volatile storage capacity).
[0141] In one embodiment, the detection module 1106 determines the sole and / or exclusive rating and / or diagnosis on the mobile computing device 102. In yet another embodiment, the detection module 1106 may determine a first rating and / or diagnosis (e.g., a first score) on the mobile computing device 102, and another detection module 1106 may determine a second rating and / or diagnosis (e.g., a second score, a more accurate and / or more detailed rating, etc.). In other embodiments, the detection module 1106 may determine the sole and / or exclusive rating and / or diagnosis on the backend server device 108.
[0142] In certain embodiments, multiple audio modules 104 may be configured to conduct one or more clinical trials with users, including trial participants (e.g., to determine the efficacy of a medical treatment based on analysis of the participants' audio data). In such embodiments, the detection module 1106 may determine an evaluation of the efficacy of a medical treatment for a medical condition related to the clinical trial. For example, users, such as trial participants, may be divided into at least a placebo group that does not receive the medical treatment and a different group that receives the medical treatment, or into multiple groups that receive different medical treatments, etc.
[0143] Multiple distributed detection modules 1106 can be configured to provide blinded assessments of a medical condition for both a placebo group and one or more groups receiving a medical treatment, allowing one or more administrators of a clinical trial to determine the efficacy of the medical treatment. For example, the detection modules 1106 can determine the severity of a medical condition, the severity of one or more symptoms of a medical condition, etc. for a placebo group and for a group receiving a medical treatment and compare the two. As used herein, a "blind" assessment refers to an assessment that is not based on whether a participant is in a placebo group or a group receiving a medical treatment. For example, in certain embodiments, the detection modules 1106 may use the same models, the same analyses, etc. for both the placebo group trial participants and the group receiving a medical treatment.
[0144] In certain embodiments, instead of making an assessment based solely on the efficacy of the medical treatment in treating the condition associated with the clinical trial, the detection module 1106 is configured to make an assessment based, at least in part, on one or more biomarkers indicative of the user's quality of life (e.g., verbal response data, sensor data, etc.) in the received response data. For example, in addition to assessing the condition associated with the clinical trial, the detection module 1106 may also assess one or more quality of life biomarkers indicative of physical fatigue, tiredness, mental fatigue, stress, anxiety, depression, and / or other parameters related to the user's quality of life. A biomarker, as used herein, includes a measurable indicator from the user of some biological state and / or condition of the user (e.g., the presence of a disease and / or injury, the presence of one or more symptoms, the user's current quality of life, etc.). In certain embodiments, a biomarker may include a feature objectively identifiable by the detection module 1106 in response to data from the user, such as an acoustic feature, a linguistic feature, a feature identifiable in sensor data, etc.
[0145] In certain embodiments, the detection module 1106 may initially use baseline response data from a user (prospective clinical trial participant) to screen participants for clinical trials (e.g., to determine the user's eligibility for a clinical trial aimed at a certain medical condition, etc.). For example, an anti-cancer drug treatment may be effective, but may be detrimental to the quality of life of the individual using the anti-cancer drug treatment. Instead of subjectively measuring the quality of life and / or behavioral consequences of receiving a treatment or drug using a questionnaire or similar means, in certain embodiments, the detection module 1106 may objectively use biomarkers or other indicators in verbal response data from the user to identify one or more quality of life changes in the user (e.g., clinical trial participant).
[0146] In particular embodiments, the interface module 1108 cooperates with the query module 1102 to display one or more questions and / or prompts to the user (e.g., instead of, in addition to, audibly asking the user questions, etc.). The response module 1104 can display one or more user interface elements (e.g., a play button, a replay button, a next question button, a previous question button, etc.) to allow the user to navigate through one or more questions of the query module 1102. In one embodiment, the determination module 1106 can determine whether an answer from the user to a question from the query module 1102 is correct or incorrect (e.g., based on speech analysis using a machine learning model, etc.), and the interface module 1108 can symbolically indicate whether the answer is correct or incorrect (e.g., dynamically, during management of ratings for the query module 1102, etc.). In yet other embodiments, the determination module 1106 may use automatic speech recognition to record the user's voice response from the response module 1104 and convert it to text, which the interface module 1108 may display to the user (e.g., dynamically, in real time, etc.).
[0147] In particular embodiments, the interface module 1108 (e.g., in cooperation with the query module 1102) can prompt the user to repeat a passage (e.g., a passage including a sentence, a set of words, letters, numbers, monosyllables, etc.). The interface module 1108 may prompt the user to repeat the same passage and / or set of passages each time data is collected (e.g., for a collection of baseline response data, a collection of test case response data, a collection of clinical trial screening data, a collection of clinical trial data, etc.).
[0148] In some embodiments, the interface module 1108 may also prompt the administrator of the assessment (e.g., a coach, parent, medical professional, etc.) with instructions to administer one or more physical examinations of the user being assessed. For example, the interface module 1108 may provide instructions for balance tests, motor coordination tests, range of motion tests, flexibility tests, tactile tests, strength tests, etc., and may provide an interface for the administrator to record the results (e.g., administrator observations) for the response module 1104.
[0149] In one embodiment, the interface module 1108 provides one or more users with access to response data (e.g., audio recordings, baseline response data, test case response data, sensor data, etc.) received from the response module 1104, access to evaluations and / or other diagnostics from the detection module 1106, etc. The interface module 1108 may provide users with access to the response data, evaluations and / or other diagnostics, etc. received from multiple locations (e.g., from a mobile app on a mobile computing device 102, from a web browser on a different computing device 102 accessing a web server on the backend server 108, etc.).
[0150] For example, the interface module 1108 may present to the user baseline ratings and / or scores based on baseline response data, test case ratings and / or scores based on test case response data, follow-up ratings and / or scores based on subsequent responses (e.g., follow-up home assessments during recovery from a previously assessed / diagnosed condition), etc., each presented through the same graphical user interface on one or more computing devices 102 along with the associated response data for each rating and / or score. The interface module 1108 may display the baseline ratings and / or scores next to (e.g., side-by-side with) the current (e.g., test case) ratings and / or scores for comparison, and may display the difference between the baseline ratings and / or scores and the current (e.g., text case) ratings and / or scores, etc. In one embodiment, the interface module 1108 may display a breakdown of the ratings and / or scores using subscores for different categories.
[0151] In some embodiments, the interface module 1108 may aggregate response data, scores, or other evaluations, etc., for users from multiple sports, teams, schools, etc. and display them within a single graphical user interface. In this manner, the interface module 1108 may provide medical professionals, coaches, administrators, etc. with a more complete history and / or status of a user's health, injury history, etc., to make more informed medical decisions.
[0152] In certain embodiments, the interface module 1108 may implement access control permissions by authenticating a user (e.g., via username and password or other authentication credentials) and granting the user access to audio recordings or other response data, assessments or other diagnostics, etc. based on the access control permissions associated with the user (e.g., for personal protection, security, HIPAA compliance, etc.). In certain embodiments, the interface module 1108 implements hierarchical access control permissions for different users, with users at each level in the hierarchy able to access data associated with any level below their level in the hierarchy.
[0153] For example, in embodiments in which the audio module 104 is configured to diagnose concussions and / or other medical conditions for athletes, athletes, parents, and / or guardians may be granted access to the athletes' own individual response data (e.g., audio recordings, evaluations, and / or other diagnoses); coaches may have access to similar data for each team member (e.g., multiple athletes or other users); school or league administrators may have access to similar data for team members of multiple teams (e.g., each team at a school, each team at a league, etc.); district or regional administrators may have access to similar data for team members at multiple schools or leagues, etc. In certain embodiments, the interface module 1108 may provide personalized information to individuals and their coaches, but may average or otherwise anonymize the data for other levels of the hierarchy (e.g., by team, school, location, league, etc.), thereby anonymizing data for particular users (e.g., response data such as audio recordings and / or sensor data, evaluations, and / or other diagnoses, etc.).
[0154] In embodiments in which the audio module 104 is conducting a medical study, the hierarchical access control permissions may allow the interface module 1108 to prohibit individual users (e.g., participants in a medical study) from accessing at least some of their own data (e.g., response data, assessments or other diagnoses, both response data and assessments, etc.), while the interface module 1108 may grant one or more administrators of the clinical trial access to data stored about users (e.g., stored baseline recorded verbal responses, test case recorded verbal responses, assessments or other diagnoses, etc.).
[0155] 12 illustrates one embodiment of a method 1200 for voice-based medical assessment. The method 1200 begins with the query module 1102 asking 1202 a question to the user (e.g., audibly through a speaker of the computing device 102, written on an electronic screen of the computing device 102, etc.).
[0156] The response module 1104 receives 1204 a user response (e.g., a verbal response from a microphone of the computing device 102, a touch response from a touchscreen and / or touchpad of the computing device 102, sensor input from one or more sensors of the computing device 102, a selection or click from a mouse or other input device of the computing device 102, a text response entered by the user on a keyboard and / or touchscreen of the computing device 102, etc.). The detection module 1106 assesses 1206 the user for a medical condition based on an analysis of the responses received 1204 from the user, and the method 1200 ends.
[0157] 13 illustrates one embodiment of a method 1300 for voice-based medical assessment. The query module 1102 queries 1302 a user using a user interface (e.g., a microphone, an electronic display screen, a touch screen, and / or one or more other sensors) of the computing device 102. The response module 1104 records 1304 one or more baseline responses (e.g., verbal responses such as audio recordings, text responses, and / or sensor data as a data file or other data structure) of the user to the one or more queries (1302) posed on the computing device 102, 108.
[0158] The detection module 1106 detects 1306 a potential medical event (e.g., based on a user requesting a medical evaluation, based on data from a sensor, and / or based on receiving other triggers). If the detection module 1106 does not detect 1306 a potential medical event, the method 1300 continues until the detection module 1106 detects 1306 a potential medical event.
[0159] In response to the detection module 1106 detecting (1306) a potential medical event (e.g., an impact or other phenomenon that may have caused a concussion, an indicator of a potential medical condition such as depression, stress, stroke, cognitive stability, mood, conscience, Alzheimer's disease, Parkinson's disease, etc., a request from the user, and / or other trigger), the query module 1102 re-asks (1308) one or more questions to the user using a user interface of the computing device 102.
[0160] The response module 1104 records (1310) one or more test case responses of the user to the re-asked one or more questions (1308) on the computing device 102, 108. The detection module 1106 assesses (1312) the likelihood that the user has a medical condition (e.g., concussion, depression, stress, stroke, cognitive stability, mood, conscientiousness, Alzheimer's disease, Parkinson's disease, etc.) based on an audio analysis of the one or more recorded baseline responses (1304) and the one or more recorded test case responses (1310) on the computing device 102, 108. The method 1300 continues until the detection module 1106 detects (1306) a subsequent potential medical event.
[0161] In various embodiments, the means for querying a user (e.g., audibly and / or otherwise) from the computing device 102 may comprise the voice module 104, the device voice module 104a, the backend voice module 104b, the query module 1102, the mobile computing device 102, the backend server computing device 108, electronic speakers of the computing devices 102, 108, headphones, electronic display screens of the computing devices 102, 108, user interface devices, network interfaces, mobile applications, processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may comprise substantially similar or equivalent means for querying a user.
[0162] In various embodiments, the means for receiving a user response (e.g., a verbal response, a text response, sensor data, etc.) on the computing device 102, 108 may comprise the voice module 104, the device voice module 104a, the back-end voice module 104b, the response module 1104, the mobile computing device 102, the back-end server computing device 108, a microphone, a user input device, a touch screen, a touchpad, a keyboard, a mouse, an accelerometer, a gyroscope, an image sensor, a mobile application, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may comprise substantially similar or equivalent means for receiving a response.
[0163] In various embodiments, the means for assessing a user for a medical condition based on a response received from the user may comprise the voice module 104, the device voice module 104a, the back-end voice module 104b, the detection module 1106, the mobile computing device 102, the back-end server computing device 108, a mobile application, machine learning, artificial intelligence, an acoustic feature computation component 210, a speech recognition component 220, a Gaussian mixture model, an acoustic model, a language model, a neural network, a deep neural network, a medical condition classifier 240, a classifier, a support vector machine, a multi-layer perceptron, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may comprise substantially similar or equivalent means for assessing a user for a medical condition.
[0164] In various embodiments, the means for authenticating different users in a hierarchy of users may comprise voice module 104, device voice module 104a, backend voice module 104b, interface module 1108, mobile computing device 102, backend server computing device 108, mobile application, processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may comprise substantially similar or equivalent means for authenticating different users.
[0165] In various embodiments, the means for granting access to different records and / or different ratings to different users (e.g., based on hierarchical access control permissions for a hierarchy of users, etc.) may comprise the voice module 104, the device voice module 104a, the back-end voice module 104b, the interface module 1108, the mobile computing device 102, the back-end server computing device 108, the mobile application, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may comprise substantially similar or equivalent means for granting access to different records and / or different ratings to different users.
[0166] The methods and systems described herein may also be deployed, in part or in whole, by a machine that executes computer software, program code, and / or instructions on a processor. As used herein, the term "processor" is meant to include at least one processor, and the plural and singular terms should be understood as interchangeable unless the context clearly indicates otherwise. Any aspect of the present disclosure may be realized as a method on a machine, a system or apparatus as part of or relating to a machine, or a computer program product embodied in a computer-readable medium executing on one or more machines. The processor may be part of a server, client, network infrastructure, mobile computing platform, stationary computing platform, or other computing platform. The processor may be any type of computing or processing device capable of executing program instructions, code, binary instructions, etc. The processor may be or include a single processor, digital processor, embedded processor, microprocessor, or any variant, such as a coprocessor (e.g., math coprocessor, graphics coprocessor, communications coprocessor, etc.) that can directly or indirectly facilitate the execution of stored program code or program instructions. Additionally, a processor may enable the execution of multiple programs, threads, and code. Multiple threads may be executed simultaneously to improve processor performance and facilitate simultaneous processing of applications. In one embodiment, the methods, program code, program instructions, etc. described herein may be implemented in one or more threads. Threads may spawn other threads and may be assigned priorities associated therewith, and the processor may execute these threads based on priority or any other order based on instructions provided in the program code.The processor may include memory that stores methods, code, instructions, and programs, as described herein and elsewhere. The processor may access, via an interface, a storage medium that may store methods, code, and instructions, as described herein and elsewhere. Storage media associated with the processor for storing methods, programs, code, program instructions, or other types of instructions that can be executed by a computing or processing device may include, but are not limited to, one or more of a CD-ROM, DVD, memory, hard disk, flash drive, RAM, ROM, cache, etc.
[0167] A processor may include one or more cores, which can increase the speed and performance of a multiprocessor. In embodiments, a process may be a dual-core processor, a quad-core processor, or other chip-level multiprocessor that combines two or more independent cores (called a die).
[0168] The methods and systems described herein may be deployed, in part or in whole, by machines executing computer software on servers, clients, firewalls, gateways, hubs, routers, or other such computer and / or networking hardware. Software programs may be associated with servers, which may include file servers, print servers, domain servers, Internet servers, intranet servers, and other variants such as secondary servers, host servers, distributed servers, etc. Servers may include one or more of the following: memory, processors, computer-readable media, storage media, ports (physical and virtual), communication devices, and interfaces that allow access to other servers, clients, machines, and devices through wired or wireless media. Methods, programs, or code as described herein and elsewhere may be executed by a server. Additionally, other devices required for the execution of methods as described herein may be considered part of the infrastructure associated with the server.
[0169] A server may provide an interface to other devices, including, but not limited to, clients, other servers, printers, database servers, print servers, file servers, communication servers, distributed servers, etc. Additionally, this coupling and / or connection may facilitate remote execution of programs across a network. Networking some or all of these devices may facilitate parallel processing of a program or method in one or more locations without departing from the scope of this disclosure. Additionally, any device attached to a server via an interface may include at least one storage medium capable of storing methods, programs, code, and / or instructions. A central repository may provide program instructions for execution on different devices. In this embodiment, a remote repository may act as a storage medium for program code, instructions, and programs.
[0170] Software programs may also be associated with clients. Clients may include file clients, print clients, domain clients, Internet clients, intranet clients, and other variations such as secondary clients, host clients, distributed clients, etc. Clients may include one or more of memory, processors, computer-readable media, storage media, ports (physical and virtual), communication devices, and interfaces that can access other clients, servers, machines, and devices through wired or wireless media. Methods, programs, or code as described herein and elsewhere may be executed by a client. Additionally, other devices required for execution of methods as described herein may be considered part of the infrastructure associated with the client.
[0171] A client may provide an interface to other devices, including, but not limited to, servers, other clients, printers, database servers, print servers, file servers, communication servers, distributed servers, etc. Additionally, this coupling and / or connection may facilitate remote execution of programs across a network. Networking some or all of these devices may facilitate parallel processing of a program or method in one or more locations without departing from the scope of this disclosure. Additionally, any device attached to a client via an interface may include at least one storage medium capable of storing methods, programs, applications, code, and / or instructions. A central repository may provide program instructions for execution on different devices. In this embodiment, a remote repository may act as a storage medium for program code, instructions, and programs.
[0172] The methods and systems described herein may also be deployed, in part or in whole, over a network infrastructure. The network infrastructure may include elements such as computing devices, servers, routers, hubs, firewalls, clients, personal computers, communication devices, routing devices, and other active and passive devices, modules, and / or components known in the art. The computing and / or non-computing device(s) associated with the network infrastructure may include, among other components, storage media such as flash memory, buffers, stacks, RAM, ROM, etc. The processes, methods, program codes, and instructions described herein and elsewhere may be executed by one or more of the network infrastructure elements.
[0173] The methods, program codes, and instructions described herein and elsewhere may also be implemented on a cellular network having multiple cells. The cellular network may be either a Frequency Division Multiple Access (FDMA) network or a Code Division Multiple Access (CDMA) network. The cellular network may include mobile devices, cell sites, base stations, repeaters, antennas, towers, etc. The cellular network may be a GSM, GPRS, 3G, EVDO, mesh, or other network type.
[0174] The methods, program codes, and instructions described herein and elsewhere may also be implemented on or through a mobile device. Mobile devices may include navigation devices, cell phones, mobile telephones, mobile personal digital assistants, laptops, palmtops, netbooks, pagers, e-readers, music players, etc. These devices may include, among other components, storage media such as flash memory, buffers, RAM, ROM, and one or more computing devices. A computing device associated with a mobile device may be capable of executing program code, methods, and instructions stored thereon. Alternatively, a mobile device may be configured to execute instructions in cooperation with other devices. A mobile device may be configured to communicate with a base station interfaced with a server and execute program code. A mobile device may also communicate over a peer-to-peer network, a mesh network, or other communications network. Program code may be stored on a storage medium associated with a server and executed by a computing device embedded within the server. A base station may include a computing device and a storage medium. The storage device may store program codes and instructions executed by computing devices associated with the base station.
[0175] The computer software, program code, and / or instructions may be stored on and / or accessed from a machine-readable medium. Machine-readable media can include computer components, devices, and recording media that hold digital data used for calculations at certain intervals of time; semiconductor storage known as random access memory (RAM); mass storage, typically for more permanent storage, such as optical disks, hard disks, tapes, drums, cards, and other types of magnetic storage; processor registers, cache memory, volatile memory, non-volatile memory; optical storage such as CDs and DVDs; removable media such as flash memory (e.g., USB sticks or keys), floppy disks, magnetic tape, paper tape, punch cards, standalone RAM disks, Zip drives, removable mass storage, offline, etc.; and other computer memory such as dynamic memory, static memory, read / write storage, mutable storage, read-only, random access, sequential access, position-addressable, file-addressable, content-addressable, network-attached storage, storage area networks, bar code, magnetic ink, etc.
[0176] The methods and systems described herein can transform physical and / or intangible items from one state to another, and the methods and systems described herein can transform data representing physical and / or intangible items from one state to another.
[0177] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced within their scope.
Claims
1. A method for operating a device, the method comprising: determining a likelihood that a person has any of one or more medical conditions, said determining said likelihood comprising: in response to receiving audio data comprising at least one sample of the person's voice, processing the audio data using at least one trained model, said at least one trained model having been trained using a plurality of voice data items of a plurality of people, each voice data item comprising one or more features and corresponding to a known medical condition, said one or more features comprising acoustic and linguistic features of the voice; and processing the audio data using said at least one trained model comprising analyzing the one or more acoustic and linguistic features of the person's voice using said at least one trained model; outputting the likelihood that the person has any of the one or more medical conditions; A method comprising:
2. The method includes, before determining the likelihood that the person has any of one or more medical conditions: selecting one or more prompts for the person; receiving the audio data including a voice of the person, the audio data including one or more verbal responses from the person to the one or more prompts; The method of claim 1 , further comprising:
3. The method of claim 2, wherein each prompt corresponds to a prompt selection score, and the one or more prompts are selected based on (i) information about the target condition to be diagnosed or features useful for diagnosing the target condition, and (ii) a prompt selection score for each prompt, the prompt selection score indicating a correlation between known conditions and medical diagnostic scores for the audio data items of each prompt.
4. The method of claim 3, further comprising determining the prompt selection score for each prompt based on a known medical condition corresponding to the audio data item and a diagnostic score generated by the at least one trained model for the audio data item.
5. The method of claim 1, wherein processing the audio data includes extracting one or more acoustic features of the person's voice, the one or more acoustic features including any one or more of mel-frequency cepstral coefficients, perceptual linear prediction, tone, energy, voicing probability, jitter, fluctuation, or harmonic to noise ratio.
6. The method of claim 1, wherein processing the audio data includes extracting one or more linguistic features of the person's speech, the one or more linguistic features including any one or more of speaking rate, number of filler words per second, number of filler words per word, word difficulty, or part of the speech pattern following a filler word.
7. The method of claim 1, wherein the method is performed on a mobile computing device and the audio data is recorded by a microphone of the mobile computing device.
8. The method of claim 7, wherein processing the audio data is performed on a remote computing device that communicates with the mobile computing device via a network; The method of claim 1 , further comprising receiving, at the remote computing device, the audio data captured by a microphone of the mobile computing device.
9. The method of claim 8, wherein processing the audio data is performed on a computing device remote from a location where the person uttered the voice of the audio data; The method of claim 1 , further comprising receiving, at the computing device, the audio data captured by a microphone at the location.
10. The method of claim 1, wherein determining the likelihood that the person has any of one or more medical conditions includes, for a medical condition, generating a prediction of the severity of the medical condition for the person.
11. The method of claim 1, wherein determining the likelihood that the person has any of one or more medical conditions includes, for a medical condition, generating a prediction of the type of the medical condition for the person.
12. The method comprising: receiving one or more sections of the audio data over time, at least one of the one or more sections of the audio data comprising speech audio data; identifying the at least one sample of the person's voice from among the one or more segments of the audio data; The method of claim 1 further comprising:
13. Identifying the at least one sample of the person's voice from the audio data comprises: when the segment of audio data includes speech, determining whether the speech is the speech of the person; responsive to determining that the section of audio data includes the person's voice, determining whether audio data of at least a portion of the person's voice included in the section of audio data satisfies one or more criteria for use as a voice sample; in response to determining that the at least some audio data of the person's voice satisfies the one or more criteria for use as a voice sample, including the at least some audio data as a sample of the person's voice in the at least one sample of the person's voice; 13. The method of claim 12, comprising:
14. The method of claim 12, wherein determining the likelihood that the person has any of one or more medical conditions comprises repeating the determination of the likelihood that the person has any of one or more medical conditions repeatedly over time, and in each of the repeated repeating, determining the likelihood is performed using a different portion of the one or more segments of the audio data received over time.
15. The iteratively repeating includes at least a first iteration and a second iteration; In the first iteration, determining the likelihood is performed using a first set of segments of the audio data, the first set of segments of the audio data being less than all of the segments of the audio data received over time; in the second iteration, determining the likelihood is performed using a second set of segments of the audio data, the second set of segments of the audio data being less than all of the segments of the audio data received over time; 15. The method of claim 14, wherein the first and second sets of partitions of the audio data partially overlap.
16. The method described in claim 15, wherein the first and second repetitions are performed simultaneously with the person uttering the first and second sets of voices for the multiple segments of the audio data.
17. The method described in claim 10, wherein the determining is performed simultaneously with the person uttering the voice.
18. The method described in claim 1, wherein processing the audio data further includes processing medical history information using the at least one trained model, and the at least one trained model has also been trained using multiple medical history information for the multiple people.
19. The method described in claim 1, wherein processing the audio data using the at least one trained model further includes analyzing one or more non-speech features using the at least one trained model.
20. The method of claim 1, wherein processing the audio data using the at least one trained model further includes analyzing one or more images of the person using the at least one trained model.
21. The method of claim 1, wherein processing the audio data using the at least one trained model further comprises analyzing one or more touch inputs of the person using the at least one trained model.
22. The method of claim 1, wherein processing the audio data using the at least one trained model further includes analyzing the person's medical history using the at least one trained model.
23. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method; The method comprises: determining a likelihood that a person has any of one or more medical conditions, said determining said likelihood comprising: in response to receiving audio data comprising at least one sample of the person's voice, processing the audio data using at least one trained model, said at least one trained model having been trained using a plurality of voice data items of a plurality of people, each voice data item comprising one or more features and corresponding to a known medical condition, said one or more features comprising acoustic and linguistic features of the voice; and processing the audio data using said at least one trained model comprising analyzing the one or more acoustic and linguistic features of the person's voice using said at least one trained model; outputting the likelihood that the person has any of the one or more medical conditions; 1. A non-transitory computer-readable storage medium, comprising:
24. An apparatus, comprising: at least one processor; at least one computer-readable storage medium; Equipped with the at least one computer-readable storage medium having executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform a method; The method comprises: determining a likelihood that a person has any of one or more medical conditions, said determining said likelihood comprising: in response to receiving audio data comprising at least one sample of the person's voice, processing the audio data using at least one trained model, said at least one trained model having been trained using a plurality of voice data items of a plurality of people, each voice data item comprising one or more features and corresponding to a known medical condition, said one or more features comprising acoustic and linguistic features of the voice; and processing the audio data using said at least one trained model comprising analyzing the one or more acoustic and linguistic features of the person's voice using said at least one trained model; outputting the likelihood that the person has any of the one or more medical conditions; 1. An apparatus comprising: