Artificial intelligence-based method and system for cognitive assessment through speech analysis
Patent Information
- Application Number
- PCT/US2026/016469
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-07-21
- Filing Date
- 2026-02-24
- Publication Date
- 2026-08-27
Smart Images

Figure US2026016469_27082026_PF_FP_ABST
Abstract
Description
1611.002W01ARTIFICIAL INTELLIGENCE-BASED METHOD AND SYSTEM FOR COGNITIVE ASSESSMENT THROUGH SPEECH ANALYSISRELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 834,129, filed February 24, 2025, U.S. Provisional Application No. 63 / 808,344, filed May 19, 2025, U.S. Provisional Application No. 63 / 841,575, filed July 10, 2025, and U.S.Provisional Application No. 63 / 847,805, filed July 21, 2025, the contents of which are incorporated by reference herein in their entirety.BACKGROUND
[0002] Cognitive function assessment is an essential step in diagnosing, monitoring, and managing many neurological and psychiatric conditions. Speech-based features such as prosodic changes, lexical patterns, semantic coherence, and acoustic variability can reflect underlying neurological changes, which enable accurate measurement of patient impairments. Ongoing artificial intelligence (“Al”) advancements, including conversational Al and automated speech-analysis algorithms, create opportunities for early cognitive change detection through patient speech analysis.SUMMARY
[0003] The subject matter described herein includes a method and non-transitory computer-readable medium for automated cognitive assessment and generation of cognitive assessment scores from patient cognitive assessment data. In Example 1, a method for automated cognitive assessment and generation of cognitive scores from patient audio data includes a server that hosts an artificial intelligence (Al) agent. The method uses the server and Al agent to deliver one or more oral cognitive assessment to a patient. The method then collects a collection of audio data retrieved during the one or more oral cognitive assessment. The method uses a backend analysis server to extract a plurality of features derived from the collection of audio data using a feature-embedding model. The method next uses the backend analysis server to generate at least one cognitive assessment score from the plurality of features derived from the collection of audio data using a machine-learning classifier trained on at least one clinically validated dataset, wherein generating the at least one cognitive assessment score includes comparing the plurality of features derived from the collection of audio data to diagnostic indicators. Finally, the method uses the backend1611.002W01analysis server to transmit the at least one cognitive assessment score to a clinician-accessible device.
[0004] Example 2 includes the method of Example 1, wherein the one or more oral cognitive assessment includes a server hosting an Al agent initiating an audio communication session with the patient.
[0005] Example 3 includes the method of Example 1 and / or 2, wherein the one or more oral cognitive assessment includes an interaction primer which comprises an oral narrative selected from a library of clinically validated narratives.
[0006] Example 4 includes the method of Example 1 and / or 2, wherein the one or more oral cognitive assessment includes categorical or letter verbal fluency, short-term memory word recall, text descriptions, and picture descriptions.
[0007] Example 5 includes the method of any one of Examples 1, 2, 3, and / or 4, wherein the one or more oral cognitive assessment includes delivering one or more immediate recall audio prompts for the patient to orally recall information from an oral narrative.
[0008] Example 6 includes the method of Example 1 and / or 5, wherein the one or more oral cognitive assessment further includes engaging the patient in an oral topic discussion phase incorporating targeted follow-up questions dynamically generated based on the patient’s responses.
[0009] Example 7 includes the method of any one of Examples 1, 2, and / or 3, wherein the one or more oral cognitive assessment further includes delivering one or more delayed recall audio prompts for the patient to orally recall information from an oral narrative.
[0010] Example 8 includes the method of Example 1 and / or 7, wherein the backend analysis server generates, using an automatic speech recognition (ASR) model, a time-aligned transcript from the collection of audio data retrieved during the one or more oral cognitive assessment, wherein words in the time-aligned transcript are associated with corresponding audio data timestamps of the collection of audio data.
[0011] Example 9 includes the method of Example 1 and / or 8, wherein the plurality of features include a plurality of speech-derived features derived from the collection of audio data retrieved during the one or more oral cognitive assessment and time-aligned transcript, and wherein the plurality of speech-derived features comprise prosodic features, acoustic characteristics, lexical features, and semantic-coherence metrics.
[0012] Example 10 includes the method of Example 1 and / or 9, wherein the at least one cognitive assessment score is based on a plurality of speech-derived features derived from the collection of audio data.1611.002W01
[0013] Example 11 includes the method of any one of Examples 1, 2, and / or 3, wherein selection of the oral narrative is based on patient metadata including demographics or existing impairment indicators provided by an ordering physician.
[0014] Example 12 includes the method of any one of Examples 1, 5, and / or 7, wherein immediate and delayed recall audio prompts relate to semantic units comprising settings, characters, events, and thematic elements extracted from an oral narrative.
[0015] Example 13 includes the method of Example 1 and / or 6, wherein oral topic discussions conducted by the Al agent include oral discussion topics generated based on patient responses and metadata, and in which follow-up queries are generated based on previous patient responses, and wherein oral discussion topics are generated from the group consisting essentially of the patient’s hobbies, family matters, occupation, daily activities, wellbeing, and personal opinions.
[0016] Example 14 includes the method of Example 1 and / or 2, wherein a backend analysis server obtains patient session feedback prior to termination of the audio communication session.
[0017] Example 15 includes the method of any one of Examples 1, 6, and / or 10, wherein a session record is generated from collected audio communication session data, the audio communication session data comprising a collection of audio data retrieved during the one or more oral cognitive assessment, cognitive assessment results, topic-phase data, device information, environmental data, session metadata, and patient identifiers.
[0018] Example 16 includes the method of any one of Examples 1, 6, and / or 10 wherein a session record is generated from collected audio communication session data, the session record including device diagnostic information, patient identifying information, session metadata, and environmental information.
[0019] Example 17 includes the method of Example 1 and / or 15, wherein a featureembedding model is configured to generate feature vectors comprising articulation rates, speech segmentation, voice onset time, Mel-frequency Cepstral Coefficients, fundamental frequency, jitter, shimmer, contextual relevance metrics, harmonics-to-noise ratio, and formant.
[0020] Example 18 includes the method of Example 1, wherein a machine- learning classifier is trained to generate the at least one cognitive assessment score for the measurement and prediction of disorders and conditions including but not limited to Mild Cognitive Impairment, Alzheimer’s Disease, preclinical Alzheimer’s Disease, age-related cognitive decline, Vascular Dementia, Dementia with Lewy Bodies, Frontotemporal Dementia,1611.002W01Parkinson’s Disease with or without Dementia, Depression, Stroke, Huntington’s Disease, Fatal Familial Insomnia, Traumatic Encephalopathy, and Normal Pressure Hydrocephalus.
[0021] Example 19 includes the method of Example 1, wherein the generated at least one cognitive assessment score includes at least one of an impairment probability score, risk of impairment probability score, or an impairment etiology (e.g., Alzheimer’s disease, vascular contributions, depression) probability score.
[0022] Example 20 includes the method of Example 1 and / or 17, wherein the at least one generated cognitive assessment score includes at least one of a lexical complexity score, fluency control score, lexical diversity score, grammatical sophistication score, executive control score, descriptiveness score, content accuracy score, coherence score, and scores for cognitive domains including memory, executive functioning, language use, visuospatial processing, and attention.
[0023] Example 21 includes the method of Example 1, wherein the at least one generated cognitive assessment score includes transparent explanatory data and human-in-thc-loop control elements compliant with the EU Al Act, FDA PCCP, YY / T 1833, or equivalent Artificial Intelligence system design and security regulations and standards.
[0024] Example 22 includes the method of Example 1 , wherein the transmitted at least one cognitive assessment score includes standardized clinical data fields compliant for integration with an electronic health records system.
[0025] Example 23 includes the method of Example 1, wherein data erasure complies with SP 800-88 or equivalent media sanitization standards.
[0026] Example 24 includes the method of Example 1, wherein integrations with electronic health records systems comply with HL7 Fast Healthcare Interoperability Resources standards, HIPAA, YY / T 1843, GDPR, or equivalent medical data privacy and security regulations and standards.
[0027] Example 25 includes the method of Example 1 , wherein collected and transmitted data is encrypted using cryptographic modules compliant with FIPS 140-3 Level 2 or equivalent data security standards.
[0028] Example 26 includes the method of Example 1, wherein a feedback-based training procedure is executed using clinician-generated feedback and diagnosis data to update parameters of at least one model used by the method.
[0029] In Example 27, a non-transitory computer-readable medium for automated cognitive assessment and generation of cognitive scores from patient audio data includes instructions executable by one or more processing device. The instructions’ execution by the one or more1611.002W01processing device implements an artificial intelligence (Al) based cognitive assessment system. The system uses a server and an Al agent to deliver one or more oral cognitive assessment to the patient. The system then collects a collection of audio data retrieved during the one or more oral cognitive assessment. The system uses a backend analysis server to extract a plurality of features derived from the collection of audio data. The system next uses the backend analysis server to generate at least one cognitive assessment score from the plurality of features derived from the collection of audio data using a machine-learning classifier trained on at least one clinically validated dataset, wherein generating the at least one cognitive assessment score includes comparing the plurality of features derived from the collection of audio data to diagnostic indicators. Finally, the system uses the backend analysis server to transmit the at least one cognitive assessment score to a clinician-accessible device.
[0030] Example 28 includes the non-transitory computer-readable medium of Example 27, wherein the one or more oral cognitive assessment includes a server hosting an Al agent initiating an audio communication session.
[0031] Example 29 includes the non-transitory computer-readable medium of Example 27 and / or 28, wherein the one or more oral cognitive assessment includes an interaction primer which comprises an oral narrative selected from a library of clinically validated narratives.
[0032] Example 30 includes the non-transitory computer-readable medium of Example 27 and / or 28, wherein the one or more oral cognitive assessment includes categorical or letter verbal fluency, short-term memory word recall, text descriptions, and picture descriptions.
[0033] Example 31 includes the non-transitory computer-readable medium of any one of Examples 27, 28, 29, and / or 30, wherein the one or more oral cognitive assessment includes delivering one or more immediate recall audio prompts for the patient to orally recall information from the oral narrative.
[0034] Example 32 includes the non-transitory computer-readable medium of Example 27 and / or 31, wherein the one or more oral cognitive assessment further includes engaging the patient in an oral topic discussion phase incorporating targeted follow-up questions dynamically generated based on the patient’s responses.
[0035] Example 33 includes the non-transitory computer-readable medium of any one of Examples 27, 28, and / or 29, wherein the one or more oral cognitive assessment further includes delivering one or more delayed recall audio prompts for the patient to orally recall information from the oral narrative.1611.002W01
[0036] Example 34 includes the non-transitory computer-readable medium of Example 27 and / or 33, wherein a backend analysis server generates, using an automatic speech recognition (ASR) model, a time-aligned transcript from the collection of audio data retrieved during the one or more oral cognitive assessment , wherein words in the transcript are associated with corresponding audio data timestamps of the collection of audio data.
[0037] Example 35 includes the non-transitory computer-readable medium of Example 27 and / or 34, wherein the plurality of features include a plurality of speech-derived features derived from the collection of audio data retrieved during the one or more oral cognitive assessment and time-aligned transcript, and wherein the plurality of speech-derived features comprise prosodic features, acoustic characteristics, lexical features, and semantic -coherence metrics.
[0038] Example 36 includes the non-transitory computer-readable medium of Example 27 and / or 35, wherein the generated cognitive assessment score(s) are based on a plurality of speech-derived features derived from the collection of audio data.
[0039] Example 37 includes the non-transitory computer-readable medium of any one of Examples 27, 28, and / or 29, wherein selection of the oral narrative is based on patient metadata including demographics or existing impairment indicators provided by an ordering physician.
[0040] Example 38 includes the non-transitory computer-readable medium of any one of Examples 27, 31, and / or 33, wherein the immediate and delayed recall audio prompts relate to semantic units comprising settings, characters, events, and thematic elements extracted from the oral narrative.
[0041] Example 39 includes the non-transitory computer-readable medium of Example 27 and / or 32, wherein oral topic discussions conducted by the Al agent include oral discussion topics generated based on patient responses and metadata, and in which follow-up queries are generated based on previous patient responses, and wherein oral discussion topics are generated from the group consisting essentially of the patient's hobbies, family matters, occupation, daily activities, wellbeing, and personal opinions.
[0042] Example 40 includes the non-transitory computer-readable medium of Example 27 and / or 28, wherein a backend analysis server obtains patient session feedback prior to termination of the audio communication session.
[0043] Example 41 includes the non-transitory computer-readable medium of any one of Examples 27, 32, and / or 36, wherein a session record is generated from collected audio communication session data, the audio communication session data comprising audio data,1611.002W01cognitive assessment results, topic-phase data, device information, environmental data, session metadata, and patient identifiers.
[0044] Example 42 includes the non-transitory computer-readable medium of any one of Examples 27, 28, and / or 29 wherein a session record is generated from collected audio communication session data, the session record comprising device diagnostic information, patient identifying information, session metadata, and environmental information.
[0045] Example 43 includes the non-transitory computer-readable medium of Example 27 and / or 36, wherein a feature-embedding model is configured to generate feature vectors comprising articulation rates, speech segmentation, voice onset time, Mel-frequency Cepstral Coefficients, fundamental frequency, jitter, shimmer, contextual relevance metrics, harmonics-to-noise ratio, and formant.
[0046] Example 44 includes the non-transitory computer-readable medium of Example 27, wherein a machine -learning classifier is trained to generate at least one cognitive assessment score for the measurement and prediction of disorders and conditions including but not limited to Mild Cognitive Impairment, Alzheimer’s Disease, preclinical Alzheimer’s Disease, age-related cognitive decline, Vascular Dementia, Dementia with Lewy Bodies, Frontotemporal Dementia, Parkinson’s Disease with or without Dementia, Depression, Stroke, Huntington’s Disease, Fatal Familial Insomnia, Traumatic Encephalopathy, and Normal Pressure Hydrocephalus.
[0047] Example 45 includes the non-transitory computer-readable medium of Example 27, wherein the generated at least one cognitive assessment score includes at least one of an impairment probability score, risk of impairment probability score, or an impairment etiology (e.g., Alzheimer’s disease, vascular contributions, depression) probability score.
[0048] Example 46 includes the non-transitory computer-readable medium of Example 27 and / or 43, wherein the generated at least one cognitive assessment score includes at least one of a lexical complexity score, fluency control score, lexical diversity score, grammatical sophistication score, executive control score, descriptiveness score, content accuracy score, coherence score, and scores for cognitive domains including memory, executive functioning, language use, visuospatial processing, and attention.
[0049] Example 47 includes the non-transitory computer-readable medium of Example 27, wherein the generated at least one cognitive assessment score includes transparent explanatory data and human-in-the-loop control elements compliant with the EU Al Act, FDA PCCP, YY / T 1833, or equivalent Artificial Intelligence system design and security regulations and standards.1611.002W01
[0050] Example 48 includes the non-transitory computer-readable medium of Example 27, wherein transmitted data includes standardized clinical data fields compliant for integration with an electronic health records system.
[0051] Example 49 includes the non-transitory computer-readable medium of Example 27, wherein data erasure complies with SP 800-88 or equivalent media sanitization standards.
[0052] Example 50 includes the non-transitory computer-readable medium of Example 27, wherein integrations with electronic health records systems comply with HL7 Fast Healthcare Interoperability Resources standards, HIPAA, YY / T 1843, GDPR, or equivalent medical data privacy and security regulations and standards.
[0053] Example 51 includes the non-transitory computer-readable medium of Example 27, wherein collected and transmitted data is encrypted using cryptographic modules compliant with FIPS 140-3 Level 2 or equivalent data security standards.
[0054] Example 52 includes the non-transitory computer-readable medium of Example 27, wherein a feedback-based training procedure is executed using clinician-generated feedback and diagnosis data to update parameters of at least one model used by the method.ENUMERATED EMBODIMENTS
[0055] 1. A method for automated cognitive assessment and generation of cognitive assessment scores, the method comprising:delivering to a patient, by a server hosting an artificial intelligence (“Al”) agent, one or more oral cognitive assessment;collecting a collection of audio data retrieved during the one or more oral cognitive assessment;extracting, using a feature-embedding model, a plurality of features derived from the collection of audio data;generating at least one cognitive assessment score from the plurality of features derived from the collection of audio data;wherein generating the at least one cognitive assessment score includes comparing the plurality of features derived from the collection of audio data to diagnostic indicators;1611.002W01wherein diagnostic indicators are generated by a machine-learning classifier trained on at least one clinically validated dataset; and transmitting the at least one cognitive assessment score to a clinician- accessible device.
[0056] 2. The method of Embodiment 1, wherein the one or more oral cognitive assessment includes the server hosting an Al agent initiating an audio communication session with the patient.
[0057] 3. The method of Embodiment 1 or 2, wherein the one or more oral cognitive assessment includes an interaction primer which comprises an oral narrative selected from a library of clinically validated narratives.
[0058] 4. The method of Embodiment 1 or 2, wherein the one or more oral cognitive assessment includes categorical or letter verbal fluency, short-term memory word recall, text descriptions, and picture descriptions.
[0059] 5. The method of any one of Embodiments 1, 2, 3, or 4, wherein the one or more oral cognitive assessment includes delivering one or more immediate recall audio prompts for the patient to orally recall information from the oral narrative.
[0060] 6. The method of Embodiment 1 and / or 5, wherein the one or more oral cognitive assessment further includes engaging the patient in an oral topic discussion phase incorporating targeted follow-up questions dynamically generated based on the patient’s responses.
[0061] 7. The method of any one of Embodiments 1, 2, and / or 3, wherein the one or more oral cognitive assessment further includes delivering one or more delayed recall audio prompts for the patient to orally recall information from the oral narrative.
[0062] 8. The method of any one of Embodiment 1 and / or 7, wherein a backend analysis server generates, using an automatic speech recognition (ASR) model, a time-aligned transcript from the collection of audio data, wherein words in the transcript are associated with corresponding audio data timestamps from the collection of audio data.
[0063] 9. The method of Embodiment 1 and / or 8, wherein the plurality of features derived from the collection of audio data further includes features derived from the time-aligned transcript, and wherein the plurality of features comprises prosodic features, acoustic characteristics, lexical features, and semantic-coherence metrics.1611.002W01
[0064] 10. The method of Embodiment 1 and / or 9, wherein the generated at least one cognitive assessment score is based on a plurality of speech-derived features derived from the collection of audio data.
[0065] 11. The method of any one of Embodiments 1, 2, and / or 3, wherein selection of an oral narrative is based on patient metadata including demographics or existing impairment indicators provided by an ordering physician.
[0066] 12. The method of any one of Embodiments 1, 5, and / or 7, wherein the immediate and delayed recall audio prompts relate to semantic units comprising settings, characters, events, and thematic elements extracted from the narrative.
[0067] 13. The method of Embodiment 1 and / or 6, wherein oral topic discussions conducted by the Al agent include oral discussion topics generated based on patient responses and metadata, and in which follow-up queries are generated based on previous patient responses, and wherein oral discussion topics are generated from the group consisting essentially of the patient's hobbies, family matters, occupation, daily activities, wellbeing, and personal opinions.
[0068] 14. The method of Embodiment 1 and / or 2, wherein a backend analysis server obtains patient session feedback prior to termination of the audio communication session.
[0069] 15. The method of any one of Embodiments 1, 6, and / or 10 wherein a session record is generated from collected audio communication session data, the audio communication session data comprising a collection of audio data retrieved during the one or more oral cognitive assessment, cognitive assessment results, topic-phase data, device information, environmental data, session metadata, and patient identifiers.
[0070] 16. The method of any one of Embodiments 1 6, and / or 10, wherein a session record is generated from collected audio communication session data, the session record comprising device diagnostic information, patient identifying information, session metadata, and environmental information.
[0071] 17. The method of Embodiment 1 and / or 15, wherein a feature-embedding model is configured to generate feature vectors comprising articulation rates, speech segmentation, voice onset time, Mel-frequency Cepstral Coefficients, fundamental frequency, jitter, shimmer, contextual relevance metrics, harmonics-to-noise ratio, and formant.
[0072] 18. The method of Embodiment 1, wherein a machine-learning classifier is trained to generate the at least one cognitive assessment score for the measurement and prediction of disorders and conditions including but not limited to Mild Cognitive Impairment, Alzheimer’s Disease, preclinical Alzheimer’s Disease, age-related cognitive decline,1611.002W01Vascular Dementia, Dementia with Lewy Bodies, Frontotemporal Dementia, Parkinson’s Disease with or without Dementia, Depression, Stroke, Huntington’s Disease, Fatal Familial Insomnia, Traumatic Encephalopathy, and Normal Pressure Hydrocephalus.
[0073] 19. The method of Embodiment 1, wherein the generated at least one cognitive assessment score includes at least one of an impairment probability score, risk of impairment probability score, or an impairment etiology (e.g., Alzheimer’s disease, vascular contributions, depression) probability score.
[0074] 20. The method of Embodiments 1 and 16, wherein the generated at least one cognitive assessment score includes at least one of a lexical complexity score, fluency control score, lexical diversity score, grammatical sophistication score, executive control score, descriptiveness score, content accuracy score, coherence score, and scores for cognitive domains including memory, executive functioning, language use, visuospatial processing, and attention.
[0075] 21. The method of Embodiment 1, wherein the generated at least one cognitive assessment score includes transparent explanatory data and human-in-the-loop control elements compliant with the EU Al Act, FDA PCCP, YY / T 1833, or equivalent Artificial Intelligence system design and security regulations and standards.
[0076] 22. The method of Embodiment 1, wherein the transmitted at least one cognitive assessment score includes standardized clinical data fields compliant for integration with an electronic health records system.
[0077] 23. The method of Embodiment 1, wherein data erasure complies with SP 800-88 or equivalent media sanitization standards.
[0078] 24. The method of Embodiment 1 , wherein integrations with electronic health records systems comply with HL7 Fast Healthcare Interoperability Resources standards, HIPAA, YY / T 1843, GDPR, or equivalent medical data privacy and security regulations and standards.
[0079] 25. The method of Embodiment 1, wherein data is encrypted using cryptographic modules compliant with FIPS 140-3 Level 2 or equivalent data security standards.
[0080] 26. The method of Embodiment 1, wherein a feedback-based training procedure is executed using clinician-generated feedback and diagnosis data to update parameters of at least one model used by the method.
[0081] 27. A non-transitory computer- readable medium comprising:1611.002W01a non-transitory computer readable medium including instructions which, when executed by one or more processing devices, instruct the one or more processing devices to:implement an artificial intelligence (“Al’’) based cognitive assessment system that is configured to:deliver to the patient, using an Al agent, one or more oral cognitive assessment;collect a collection of audio data retrieved during the one or more oral cognitive assessment:extract, using a feature-embedding model, a plurality of features derived from the collection of audio data;generate at least one cognitive assessment score from the plurality of features derived from the collection of audio data;wherein generating the at least one cognitive assessment score includes comparing the plurality of features derived from the collection of audio data to diagnostic indicators;wherein diagnostic indicators are generated by a machine-learning classifier trained on at least one clinically validated dataset; and transmit, using a communication interface, the at least one cognitive assessment score to a clinician-accessible device.
[0082] 28. The non-transitory computer-readable medium of Embodiment 27, wherein the one or more oral cognitive assessment includes a device hosting an Al agent initiating an audio communication session with the patient.
[0083] 29. The non-transitory computer-readable medium of Embodiment 27 and / or 28, wherein the one or more oral cognitive assessment includes an interaction primer which comprises an oral narrative selected from a library of clinically validated narratives.
[0084] 30. The non-transitory computer-readable medium of Embodiment 27 and / or 28, wherein the one or more oral cognitive assessment includes categorical or letter verbal fluency, short-term memory word recall, text descriptions, and picture descriptions.1611.002W01
[0085] 31. The non-transitory computer-readable medium of any one of Embodiments 27, 28, 29, and / or 30, wherein the one or more oral cognitive assessments include delivering one or more immediate recall audio prompts for the patient to orally recall information from the oral narrative.
[0086] 32. The non-transitory computer-readable medium of Embodiment 27 and / or 30, wherein the one or more oral cognitive assessment further includes engaging the patient in an oral topic discussion phase incorporating targeted follow-up questions dynamically generated based on the patient’s responses.
[0087] 33. The non-transitory computer-readable medium of any one of Embodiments 27, 28, and / or 29, wherein the one or more oral cognitive assessment further includes delivering one or more delayed recall audio prompts for the patient to orally recall information from the oral narrative.
[0088] 34. The non-transitory computer-readable medium of any one of Embodiment 27 and / or 33, wherein the backend analysis server generates, using an automatic speech recognition (ASR) model, a time-aligned transcript from the collection of audio data, wherein words in the transcript are associated with corresponding audio data timestamps from the collection of audio data.
[0089] 35. The non-transitory computer-readable medium of Embodiment 27 and / or 34, wherein the plurality of features derived from the collection of audio data further includes features derived from the time-aligned transcript, and wherein the plurality of features comprises prosodic features, acoustic characteristics, lexical features, and semantic-coherence metrics.
[0090] 36. The non-transitory computer-readable medium of Embodiment 27 and / or 35, wherein the generated at least one cognitive assessment score is based on a plurality of speech-derived features derived from the collection of audio data.
[0091] 37. The non-transitory computer-readable medium of any one of Embodiments 27, 28, and / or 29, wherein selection of an oral narrative is based on patient metadata including demographics or existing impairment indicators provided by an ordering physician.
[0092] 38. The non-transitory computer-readable medium of any one of Embodiments 27, 31, and / or 33, wherein the immediate and delayed recall audio prompts relate to semantic units comprising settings, characters, events, and thematic elements extracted from the narrative.
[0093] 39. The non-transitory computer-readable medium of Embodiment 27 and / or 32, wherein the oral topic discussions conducted by the Al agent include oral discussion topics1611.002W01generated based on patient responses and metadata, and in which follow-up queries are generated based on previous patient responses, and wherein oral discussion topics are generated from the group consisting essentially of the patient’s hobbies, family matters, occupation, daily activities, wellbeing, and personal opinions.
[0094] 40. The non-transitory computer-readable medium of Embodiment 27 and / or 28, wherein a backend analysis server obtains patient session feedback prior to termination of the audio communication session.
[0095] 41. The non-transitory computer-readable medium of any one of Embodiments 27, 32, and / or 36 wherein a session record is generated from collected audio communication session data, the audio communication session data comprising a collection of audio data retrieved during the one or more oral cognitive assessment, cognitive assessment results, topic-phase data, device information, environmental data, session metadata, and patient identifiers.
[0096] 42. The non-transitory computer-readable of any one of Embodiments 27, 32, and / or 36, wherein a session record is generated from collected audio communication session data, the session record comprising device diagnostic information, patient identifying information, session metadata, and environmental information.
[0097] 43. The non-transitory computer-readable medium of Embodiment 27 and / or 41, wherein a feature-embedding model is configured to generate feature vectors comprising articulation rates, speech segmentation, voice onset time, Mel-frequency Cepstral Coefficients, fundamental frequency, jitter, shimmer, contextual relevance metrics, harmonics-to-noise ratio, and formant.
[0098] 44. The non-transitory computer-readable medium of Embodiment 27, wherein a machine-learning classifier is trained to generate the at least one cognitive assessment score for the measurement and prediction of disorders and conditions including but not limited to Mild Cognitive Impairment, Alzheimer’s Disease, preclinical Alzheimer’s Disease, age-related cognitive decline, Vascular Dementia, Dementia with Lewy Bodies, Frontotemporal Dementia, Parkinson’s Disease with or without Dementia, Depression, Stroke, Huntington’s Disease, Fatal Familial Insomnia, Traumatic Encephalopathy, and Normal Pressure Hydrocephalus.
[0099] 45. The non-transitory computer-readable medium of Embodiment 27, wherein the generated at least one cognitive assessment score includes at least one of an impairment probability score, risk of impairment probability score, or an impairment etiology (e.g„ Alzheimer’s disease, vascular contributions, depression) probability score.1611.002W01
[0100] 46. The non-transitory computer-readable medium of Embodiment 27 and / or 43, wherein the generated at least one cognitive assessment score includes at least one of a lexical complexity score, fluency control score, lexical diversity score, grammatical sophistication score, executive control score, descriptiveness score, content accuracy score, coherence score, and scores for cognitive domains including memory, executive functioning, language use, visuospatial processing, and attention.
[0101] 47. The non-transitory computer-readable medium of Embodiment 27, wherein the generated at least one cognitive assessment score includes transparent explanatory data and human-in-the-loop control elements compliant with the EU Al Act, FDA PCCP, YY / T 1833, or equivalent Artificial Intelligence system design and security regulations and standards.
[0102] 48. The non-transitory computer-readable medium of Embodiment 27, wherein the transmitted at least one cognitive assessment score includes standardized clinical data fields compliant for integration with an electronic health records system.
[0103] 49. The non-transitory computer-readable medium of Embodiment 27, wherein data erasure complies with SP 800-88 or equivalent media sanitization standards.
[0104] 50. The non-transitory computer-readable medium of Embodiment 27, wherein integrations with electronic health records systems comply with HL7 Fast Healthcare Interoperability Resources standards, HIPAA, YY / T 1843, GDPR, or equivalent medical data privacy and security regulations and standards.
[0105] 51. The non-transitory computer-readable medium of Embodiment 27, wherein data is encrypted using cryptographic modules compliant with FIPS 140-3 Level 2 or equivalent data security standards.
[0106] 52. The non-transitory computer-readable medium of Embodiment 27, wherein a feedback-based training procedure is executed using clinician-generated feedback and diagnosis data to update parameters of at least one model used by the method.BRIEF DESCRIPTION OF THE DRAWINGS
[0107] It is to be understood that the drawings depict only exemplary embodiments and are not therefore to be considered limiting in scope. The exemplary embodiments will be described with additional specificity and detail in the accompanying drawings, in which:
[0108] Figure 1 is a block diagram of an example system including an Al conversational agent, a patient using a patient-operated communications device with associated audiocapture components, and related backend system servers.1611.002W01
[0109] Figure 2 is a flow diagram of an example audio-based conversational assessment session illustrating narrative selection and delivery, assessment prompting, and context-dependent topic discussions.
[0110] Figure 3 is a flow diagram of an example process for generating conversational discussion topics, including narrative selection, dynamic topic generation, and dynamic follow-up topic generation.
[0111] Figure 4 is a flow diagram of an example method for generating cognitive assessment scores using acquired audio data, ASR-generated transcripts, speech-derived feature extraction, and machine-learning-based classifiers.
[0112] Figure 5 is a block diagram of an example patient-accessible communications device configured to execute or interface with the Al conversational agent, such as a telephone, mobile phone, tablet, computer, wearable device, or virtual home assistant.DESCRIPTION
[0113] Traditional cognitive assessments are frequently conducted through in-person evaluations and structured interviews, which typically require trained clinicians, rely on subjective interpretation, and may be difficult to perform regularly or at-scale. Further, many individuals lack convenient access to clinical facilities, which may delay recognition of cognitive decline. Existing clinical workflows rarely incorporate automated extraction of speech-derived clinical diagnostic features, and available tools are typically not scalable, require manual audio transcription, and / or cannot be deployed in convenient and accessible environments such as the patient’s home. Existing approaches often fail to capture subtle, early-stage symptoms, as mild cognitive impairment and early signs of neurodegenerative disorders frequently appear gradually in speech patterns and behaviors, and such early indicators may be missed during infrequent clinical encounters. Current systems generally do not offer continuous, real-world cognitive function monitoring, and many patients thereby lose early intervention opportunities. Further, existing systems struggle to extract clinically meaningful speech features in situations involving varied environmental noise conditions, varied conversation dynamics, and spontaneous utterance patterns.
[0114] There is therefore a need for improved, patient-accessible systems that collect and analyze patient speech using objective, automated, and reproducible methods. Such systems should enable non-invasive cognitive assessment via mobile devices, tablets, computers, virtual home assistants, and other dedicated collection hardware incorporating voicerecognition, noise isolation, and / or source-separation capabilities. The improved systems1611.002W01should further support population data comparisons, integrate with clinical workflows, and facilitate early detection of cognitive impairment and measurement of current cognitive functioning through continuous or periodic monitoring without the need for specialized personnel intervention. Improved systems should additionally provide affirmative population screening capabilities, longitudinal monitoring of cognitive assessment data, and means to implement real-time clinician alerts for early intervention. The improved systems should enable clinician access to historic and current assessment data via standardized healthcare data systems and via remote or resource-constrained systems.DRAWINGS
[0115] Figure 1 is a block diagram of an example system 100 for conducting an Al-driven cognitive assessment communication session. The system 100 comprises a patient 102 operating a patient-controlled device 104, which includes audio-capture components 106, and a plurality of backend analysis servers 108 that arc communicatively coupled with the patient device 104. The system further comprises an Al conversational agent 110 executed by software and communicatively coupled to the patient device 104 and the backend analysis servers 108. The Al conversational agent 110, the patient device 104, and the backend analysis servers 108 can be communicatively coupled together over one or more networks 112, including, but not limited to, a landline -based telephone network, a cellular communications network, a local area network, intranet, and / or the internet. Communication between the system’s 100 components can use any appropriate protocol including Signaling System No. 7 (SS7), Q.931, Internet Protocol (IP), H.323, Voice Over IP (VOIP), Hypertext Transfer Protocol (HTTP), Hypertext Transfer Protocol Secure (HTTPS), and other telephonic- and internet-based protocols. Session data 114 may be locally or remotely stored depending on implementation. All patient data will be encrypted and maintained in an access-controlled, audit-enabled secure directory to ensure compliance with HIPAA, FDA ERES, SP 800-111, SP 800-122, YY / T 1843, GDPR, or equivalent healthcare data privacy, security, and records maintenance regulations and standards. Typically, any cryptographic modules or data transmission modules will implement validated algorithms, ensure authentication, and maintain data integrity to ensure compliance with FIPS 140-3 Level 2, ISO / IEC 19790 Level 2, or equivalent data security standards. Although specific components are depicted, embodiments of the system 100 may include additional components or combine functionality among components depending on implementation requirements.1611.002W01
[0116] An Al conversational agent 110 is an artificial intelligence system configured to conduct clinical assessment interviews with a patient 102. A clinical assessment interview may be structured or semi-structured. In various embodiments, the Al conversational agent 110 can be executed by software on the backend analysis servers 108, by an application on the patient’s device 104, through an intermediary service such as Voiceflow or VAPI, or by a combination of software on the backend analysis servers 108 and an application on the patient’s device 104 with or without integration with intermediary services. In some embodiments, the Al conversational agent 110 generates a speaking voice to engage the patient 102 in conversation using a service such as Elevenlabs, Deepgram, or an equivalent Al voice generation service. In other embodiments, the Al conversational agent 110 generates text on or through the patient’s device 104 to engage the patient 102 in conversation, in the event that the patient 102 is hard-of-hearing or otherwise hearing-impaired.
[0117] The patient device 104 is a communications device operated or accessible by the patient 102 and configured to host or interface with the Al conversational agent 110. The patient device 104 may include, for example, a mobile phone, telephone, tablet, laptop computer, desktop computer, wearable device (e.g., a pendant, a hearing aid, headphones, or the like), virtual home assistant (e.g., a smart speaker, a smart television, an automated companion, or the like), or other communications hardware capable of supporting audio communications or text-based communications. Typically, any embodiment comprising a wearable device will utilize validated known low-risk design materials which minimize toxicity and produce negligible irritation to ensure compliance with ISO 10993, GB / T 16886, or equivalent device safety standards. The patient device 104 may execute a client application, software module, or other patient interface module that facilitates a cognitive assessment communication session, manages local interactions with the Al conversational agent 110, and manages user-facing interface elements such as privacy notifications or feedback prompts. Typically, any patient interface module will be configured for the intended user profile of applicable patients, will include design features that mitigate reasonably foreseeable misuse, and will remain safe despite exposure to common hazards to ensure compliance with IEC 62366, IEC 60601, or equivalent device usability standards. The patient device 104 may establish secure communication channels with cloud-based Al components and backend analysis servers 108, enabling encrypted transmission of session records 114.1611.002W01
[0118] A suitable audio-capture component 106 can comprise one or more sensors, microphones, and / or audio-processing components associated with the patient device 104 configured to obtain speech data from the patient 102 during a cognitive assessment communication session. In some embodiments, an audio-capture component 106 may be integrated directly into the patient device 104, such as a telephone microphone. In other embodiments, an audio-capture component 106 may include an external and / or dedicated audio devices such as a wearable microphone, wired or wireless headsets, or other standalone audio-capture device that interface with the patient device 104 via wired connections, Bluetooth, and / or wireless communication protocols. An audio-capture component 106 may incorporate adaptive noise-isolation, echo-cancellation, source-separation technology, or the like to improve audio capture quality and / or ensure that speech data remains suitable for downstream analysis. The audio-capture components 106 may further provide real-time audio quality indicators, detect environmental noise conditions, and apply noisecompensation filters.
[0119] A backend analysis server 108 may comprise one or more cloud-based or hardware-hosted processing devices. In some embodiments, a backend analysis server 108 can include a non- transitory computer-readable medium comprising instructions that instruct one or more processing devices of the backend analysis server to implement one or more step of initiating an Al conversational agent, analyzing audio data, generating time-aligned transcripts, extracting speech-derived features, and / or generating cognitive assessment scores. The backend analysis server 108 can include an attribution module which generates transparent explanatory data detailing its cognitive assessment score generation (using, e.g., SHAP or LIME) to ensure compliance with the EU Al Act, FDA PCCP, FDA Artificial Intelligence-Enabled Device Software Functions guidance, YY / T 1833, GB / T 45654, or equivalent Artificial Intelligence system design and security regulations and standards. The backend analysis server 108 can comprise human-in-the-loop logging and algorithmic modification protocols to ensure compliance with the EU Al Act, FDA PCCP, FDA Artificial Intelligence-Enabled Device Software Functions guidance, YY / T 1833, GB / T 45654, or equivalent Artificial Intelligence system design and security regulations and standards. In some embodiments, the backend analysis servers 108 may enforce security controls and / or provide access via APIs or web-based portals with graphical user interfaces for transmitting assessment results to authorized recipients or clinician-accessible locations to ensure compliance with HL7 Fast Healthcare Interoperability Resources, HIPAA, YY / T 1843, GDPR, or equivalent medical data privacy and security regulations and standards.1611.002W01
[0120] Figure 2 is a flow diagram illustrating an example method 200 for conducting an AI-driven cognitive assessment communication session. Although specific steps are depicted, embodiments of the method 200 may include additional steps or combine functionality among steps depending on implementation requirements.
[0121] At step 202, an Al conversational agent initiates an audio communication session with the patient via the patient’s device 202. In some embodiments, the communications session is configured to use a communications link coupled with a client application on the patient’s device 202. In other embodiments, the communications session is configured to only use the native communications features present on the patient’s device 202. The audio communication session may be scheduled, automatically triggered by predefined conditions, or patient-prompted.
[0122] At step 204, the Al conversational agent delivers a narrative selected from a library of relevant cognitive-assessment narratives. The selection may be based on patient metadata including demographics or determination of existing impairment from impairment indicators provided by the ordering physician. Delivery of the narrative utilizes audio components on the patient’s device 202, including microphones and speakers.
[0123] At step 206, the agent queries the patient with an immediate recall free-response prompt to assess short-term memory retrieval. Delivery of the prompt utilizes audio components on the patient's device 202, including microphones. In an alternative embodiment, the query is delivered to the patient’s device 202 as a text prompt.
[0124] At step 208, the Al conversational agent engages the patient in a topic discussion phase. Broad discussion topics are presented to the patient (e.g., “Tell me about your day”), followed by between one and ten follow-up questions based on the patient’ s response to the initial broad discussion topic. Multiple discussion topics are presented to the patient to elicit speech detailed enough for analysis. If the patient does not generate enough volume of speech, the agent will prompt the patient with further discussion topics until responses are sufficient for analysis, calculated based on the number of tokens generated by the patient per discussion topic (minimum 150 tokens, equivalent to approximately 100 words) as well as the duration of the total session (minimum 10 minutes). The discussion phase is designed to elicit spontaneous speech and conversational behavior for analysis, and follow-up queries may be generated in real time to maintain a natural conversational flow. This multi-stage relevance-driven topic generation architecture improves structural stability and increases the reliability of downstream scoring and assessment by improving speech collection under1611.002W01variable linguistic and situational conditions. Delivery of the topic discussion utilizes audio components on the patient’s device 202, including microphones and speakers.
[0125] At step 210, the Al conversational agent delivers additional assessment tasks, including but not limited to delayed recall prompts to assess long-term memory recall or recognition or fluency tasks including categorical or letter verbal fluency. The patient’s responses provide data indicative of cognitive functioning following a brief conversational distraction. Delivery of additional assessment tasks utilizes audio components on the patient’s device 202, including microphones and speakers.
[0126] At step 212, session data is collected and stored in a session record on the patient’s device 202. In some embodiments, a client application on the patient’ s device 202 controls the process of saving the session record. The collected session data may include audio from all conversational phases, patient responses, recall results, and metadata related to environmental conditions or device performance.
[0127] At step 214, a backend analysis server uses an Automatic Speech Recognition (ASR) model to process the audio data to produce a time-aligned transcript. In some embodiments, the ASR model is implemented by the backend analysis server using its processor, memory, and storage. In other embodiments, an external service provider implements the ASR model. Transcript words are associated with corresponding timestamps to facilitate alignment with acoustic features and conversational segments.
[0128] At step 216, a backend analysis server uses a feature-embedding model to extract a plurality of speech-derived features from the audio data and corresponding time-aligned transcript, including prosodic, acoustic, lexical, and semantic-coherence metrics. In some embodiments, the feature-embedding model is implemented by the backend analysis server using its processor, memory, and storage. In other embodiments, an external service provider implements the feature-embedding model. In some embodiments, the model comprises multiple feature-processing stages to separately extract audio characteristics and transcript-derived linguistic features. Extracted features may include articulation rates, speech segmentation, lexical frequencies and diversities, parts-of-speech usage and ratios, disfluency and pause rates and durations, and other measures relevant to cognitive analysis. In certain embodiments, feature extraction models may include methods such as tagging parts-of-speech, linguistic analysis, audiometric analysis, tagging function and content words, identifying speech error / repair markers, applying LSTM, BERT, or other models to generate embeddings from recorded and transcribed speech, or equivalent feature extraction methods. Semantic-coherence metrics may include cosine similarities between consecutive utterance1611.002W01embeddings, a cumulative topic-drift measure of the difference between early-session and late-session embedding features, graph-based coherence measures derived from an utterance similarity graph, or equivalent semantic-coherence metrics. Analysis of semantic-coherence metrics will also generate content and memory features, including but not limited to forgetting rates for content units (names, locations, and other details), accuracy and breadth of descriptions, and integration of acoustic and lexical features (e.g., duration of unfilled pauses in descriptions of location content units).
[0129] At step 218, a backend analysis server uses a machine-learning classifier trained on clinically validated datasets to generate at least one cognitive assessment scores by comparing extracted features to diagnostic indicators. In some embodiments, the machinelearning classifier comprises a multi-stream, multi-stage architecture. Input streams may include data separated by modality (e.g., lexical data derived from transcripts vs audiometric data derived from waveforms, etc.), by task (e.g., narrative recall data indicating memory performance, interview data indicating presence or absence of cognitive leisure activities, etc.), or a combination of the two. The first stage of analysis integrates each input stream into specialized subsets of models based on domains and subdomains (e.g., memory, attention, language use, social functioning, physical activity, cognitive activity, etc.), with features assigned to each model subset based on a combination of domain expertise, feature selection, feature engineering, and / or dimensionality reduction. These first-stage models will use multimodel multi-modal pipelines that integrate architectures including but not limited to random forests with and without gradient boosting, support vector machines, logistic and linear regressions, and neural networks. Each input stream may also be processed by multiple stages of models to produce domain and subdomain subscores. A second stage will integrate the outputs of these models (final layer for multilayer perceptrons, confidence scores for random forest or logistic regression classifiers, regression scores for regression models, etc.) in combination with other engineered features to further refine the collected data into explainable subscores to support impairment classification and risk assessment. In some embodiments, the generated assessment scores following this stage may include impairment probabilities, risk scores, cognitive domain scores, deviation indices, or other clinically relevant metrics, each based on outputs derived from first stage subscores, second stage subscores, or a combination of the two. In some embodiments, the backend analysis server may fuse the multiple input streams using a feature-fusion model that applies dynamic weights to each feature type based on its learned relevance to impairment assessment prediction. The backend analysis server may apply several regression or classification1611.002W01models, including but not limited to logistic or linear regressions, support vector machines, random forests, gradient-boosted machines, and neural networks including multilayer perceptrons or transformer models. The multi-stream and fusion-based architecture improves robustness and enhances consistency of impairment assessments across conversational situations relative to single-stream or non-fused classifications.
[0130] At step 220, the cognitive assessment scores, extracted features, and session records are transmitted to a clinician-accessible location, which may include a web portal with a graphical user interface, an electronic health record (EHR) system, a clinical analytics platform, a caregiver device, or a backend storage system. Data transmissions are facilitated using the backend analysis server’s network components, which may include, but are not limited to, Network Interface Cards, optical transceivers, wireless communications transceivers, VSAT terminals, and other equivalent network components. In some embodiments, a clinician may access transmitted scores, features, and records using a computing device such as a tablet, mobile phone, computer, or other clinician-accessible device. In some embodiments, transmission of cognitive assessment scores may include an alert or other indicator provided to the clinician-accessible location to notify the clinician that the patient’s scores warrant follow-up clinician assessment.
[0131] At step 222, some embodiments of the system may collect and incorporate clinician-provided cognitive assessment scores and diagnoses of cognitive or neurodegenerative disorders into downstream model-update procedures, perform population-level assessment comparisons, or initiate follow-up assessments.
[0132] Figure 3 is a flow diagram illustrating an example process 300 for generating and delivering discussion topics based on patient-specific contextual data. Although specific steps are depicted, embodiments of the process 300 may include additional steps or combine functionality among steps depending on implementation requirements.
[0133] At step 302, an Al conversational agent delivers a narrative selected from a library of clinically relevant cognitive-assessment narratives to the patient. Narrative selection may be based on patient metadata including demographics or determination of existing impairment from indicators provided by the ordering physician.
[0134] At step 304, an Al conversational agent queries the patient’s retention of immediate recall prompt information is queried to evaluate short-term memory function using an open-ended free recall question (e.g., “Describe everything you remember about the story”), which is statically linked to the delivered narrative to maintain consistency across patientinteractions. The Al conversational agent queries the patient's information retention using the audio-capture components on the patient’s device.
[0135] At step 306, an Al conversational agent identifies potential discussion topics that align with assessment objectives based on the type of evaluation ordered by the clinician.
[0136] At step 308, an Al conversational agent generates one or more discussion topics for the topic-discussion phase. These topics may concern the patient’s hobbies, family matters, occupation, daily activities, well-being, diet, physical activity, social activity, personal opinions, or other previously identified conversational themes. The Al conversational agent delivers discussion topics using the audio-capture components on the patient’s device.
[0137] At step 310, an Al conversational agent dynamically generates follow-up conversational prompts, enabling the Al conversational agent to maintain a natural conversation while eliciting speech samples relevant to cognitive assessment. The Al conversational agent delivers follow-up prompts using the audio-capture components on the patient’s device.
[0138] At step 312, following the topic-discussion phase, the Al conversational agent queries the patient’s retention of delayed recall prompt information to evaluate long-term memory function using open-ended free recall (“Describe everything you remember about the story”), as well as targeted prompts about specific narrative units (“Do you remember the name of the dog?”). The recall prompts are statically linked to the delivered narrative to maintain consistency across patient interactions. The Al conversational agent queries the patient’s information retention using the audio-capture components on the patient’s device.
[0139] Figure 4 is a flow diagram of an example method 400 for generating cognitive assessment scores from patient speech. Although specific method steps are depicted, embodiments of the method 400 may include additional steps or combine functionality among steps depending on implementation requirements.
[0140] At step 402, an audio communication session between the Al conversational agent and the patient is initiated. In some embodiments, the communications session is configured to use a communications link coupled with a client application on the patient’s device. In other embodiments, the communications session is configured to only use the native communications features on a patient-accessible device. During the session, the Al conversational agent delivers a selected narrative, presents recall prompts, and engages in a topic-discussion phase to collect speech data.
[0141] At step 404, audio-capture components associated with the patient device captures patient audio. The audio-capture components may apply one or more adaptive noise-1611.002W01compensation or signal-enhancement techniques as needed to prepare the audio for analysis, including a Wiener filter, spectral subtraction filter, or a deep learning-based noise suppression model.
[0142] At step 406, session data is collected and stored in a session record and transmitted to backend servers for analysis using communications components on the patient’s device. In some embodiments, a client application on the patient’s device manages saving and transmitting the session record. The session record may include patient responses, task results, and metadata related to environmental conditions or device performance.
[0143] At step 408, the backend analysis servers use an Automatic Speech Recognition (ASR) model to process the session record and generate a time-aligned transcript containing words or tokens associated with corresponding audio timestamps. In some embodiments, the ASR model is implemented by the backend analysis server using its processor, memory, and storage. In other embodiments, an external service provider implements the ASR model.
[0144] At step 410, the backend analysis servers employ a feature-embedding model to extract a plurality of speech-derived features from the audio data and transcript. These speech-derived features may include prosodic, audiometric, lexical, and semantic-coherence features. In some embodiments, the feature-embedding model is implemented by the backend analysis server using its processor, memory, and storage. In other embodiments, an external service provider implements the feature-embedding model.
[0145] At step 412, the backend analysis servers use a machine-learning classifier trained using clinically validated and labeled patient data from one or more datasets to analyze extracted features and generate one or more oral cognitive assessment scores indicative of neurological, psychiatric, or cognitive conditions, including risk probabilities or impairment metrics.
[0146] At step 414, the backend analysis servers transmit the generated scores, extracted features, and session record to one or more clinician-accessible location. These clinician-accessible location may include clinical electronic health record (EHR) platforms, secure storage repositories, a web portal with a graphical user interface, or clinician-facing review tools. Data transmissions are facilitated using the backend analysis server’s network components, which may include, but are not limited to, Network Interface Cards, optical transceivers, wireless communications transceivers, VSAT terminals, and other equivalent network components. In some embodiments, a clinician may access transmitted scores, features, and records using a computing device such as a tablet, mobile phone, computer, or other clinician-accessible device. In some embodiments, transmission of cognitive1611.002W01assessment scores may include an alert or other indicator provided to the clinician-accessible location to notify the clinician that the patient’s scores warrant follow-up clinician assessment. In some embodiments, clinician feedback may be incorporated into iterative model-update workflows.
[0147] Figure 5 is a block diagram of an example patient-accessible communications device 500 configured to execute or interface with the Al conversational agent. The device 500 may include a processor 502, memory 504, audio-capture components 506, and a network interface 508 configured to communicate with an Al-based assessment system via one or more networks. In some embodiments, the device 500 executes a client application 510 to facilitate the cognitive assessment session, manage communications with the Al conversational agent, manage device data operations, and provide user-facing controls and notifications. Although specific device features are depicted, embodiments of the device 500 may include additional features or combine functionality among features depending on implementation requirements.
[0148] The audio capture components 506 of the device 500 may include one or more microphones, audio sensor arrays, or external audio-capture peripherals that interface with the device via wired connectors, Bluetooth, or other wireless protocols. The audio-capture components 506 may incorporate noise-isolation, echo-cancellation, or other audio quality improvement capabilities to ensure that patient speech is captured with sufficient fidelity for downstream analysis.
[0149] In some embodiments, the client application 510 on the device 500 may store session data 512, including buffered audio samples and local metadata used to maintain session continuity. The device 500 includes a secure data enclave 514 used for encryption, authentication, or privacy-secure processing of patient data. All patient data will be encrypted and maintained in an access-controlled, audit-enabled secure directory to ensure compliance with HIPAA, FDA ERES, SP 800-111, SP 800-122, YY / T 1843, GDPR, or equivalent healthcare data privacy, security, and records maintenance regulations and standards. Any cryptographic modules or data transmission modules will implement validated algorithms, implement authentication controls, and maintain data integrity to ensure compliance with FIPS 140-3 Level 2, ISO / IEC 19790 Level 2, or equivalent data security standards. All session data 512 and data within the secure data enclave 514 are removed from the device 500 at the conclusion of each session using cryptographic erasure methods compliant with SP 800-88 or equivalent media sanitization standards. In some embodiments, a patient interface module 516 within the client application 510 may present instructions,1611.002W01progress indicators, privacy statements, or optional feedback prompts to the patient. The device 500 may further support integration with wearable devices or smart-home devices configured to augment or enhance audio capture during the assessment. The device 500 may further include a localized edge-processing module 518 to perform region-specific patient data processing to comply with resident processing requirements.
Claims
1611.002W01CLAIMSWhat is claimed is:
1. A method for automated cognitive assessment and generation of cognitive assessment scores, the method comprising:delivering to a patient, by a server hosting an artificial intelligence (“Al”) agent, one or more oral cognitive assessment;collecting a collection of audio data retrieved during the one or more oral cognitive assessment;extracting, using a feature-embedding model, a plurality of features derived from the collection of audio data;generating at least one cognitive assessment score from the plurality of features derived from the collection of audio data;wherein generating the at least one cognitive assessment score includes comparing the plurality of features derived from the collection of audio data to diagnostic indicators;wherein diagnostic indicators are generated by a machine-learning classifier trained on at least one clinically validated dataset; and transmitting the at least one cognitive assessment score to a clinician- accessible device.
2. The method of Claim 1, wherein the one or more oral cognitive assessment includes the server hosting an Al agent initiating an audio communication session with the patient.
3. The method of Claim 1 or 2, wherein the one or more oral cognitive assessment includes an interaction primer which comprises an oral narrative selected from a library of clinically validated narratives.
4. The method of Claim 1 or 2, wherein the one or more oral cognitive assessment includes categorical or letter verbal fluency, short-term memory word recall, text descriptions, and picture descriptions.1611.002W015. The method of any one of Claims 1, 2, 3, or 4, wherein the one or more oral cognitive assessment includes delivering one or more immediate recall audio prompts for the patient to orally recall information from the oral narrative.
6. The method of Claim 1 and / or 5, wherein the one or more oral cognitive assessment further includes engaging the patient in an oral topic discussion phase incorporating targeted follow-up questions dynamically generated based on the patient’s responses.
7. The method of any one of Claims 1, 2, and / or 3, wherein the one or more oral cognitive assessment further includes delivering one or more delayed recall audio prompts for the patient to orally recall information from the oral narrali vc.
8. The method of any one of Claim 1 and / or 7, wherein a backend analysis server generates, using an automatic speech recognition (ASR) model, a time-aligned transcript from the collection of audio data, wherein words in the transcript are associated with corresponding audio data timestamps from the collection of audio data.
9. The method of Claim 1 and / or 8, wherein the plurality of features derived from the collection of audio data further includes features derived from the time-aligned transcript, and wherein the plurality of features comprises prosodic features, acoustic characteristics, lexical features, and semantic-coherence metrics.
10. The method of Claim 1 and / or 9, wherein the generated at least one cognitive assessment score is based on a plurality of speech-derived features derived from the collection of audio data.
11. A non- transitory computer-readable medium comprising:a non-transitory computer readable medium including instructions which, when executed by one or more processing devices, instruct the one or more processing devices to:implement an artificial intelligence (“Al”) based cognitive assessment system that is configured to:deliver to the patient, using an Al agent, one or more oral cognitive assessment;1611.002W01collect a collection of audio data retrieved during the one or more oral cognitive assessment;extract, using a feature-embedding model, a plurality of features derived from the collection of audio data;generate at least one cognitive assessment score from the plurality of features derived from the collection of audio data;wherein generating the at least one cognitive assessment score includes comparing the plurality of features derived from the collection of audio data to diagnostic indicators;wherein diagnostic indicators are generated by a machine-learning classifier trained on at least one clinically validated dataset; and transmit, using a communication interface, the at least one cognitive assessment score to a clinician-accessible device.
12. The non-transitory computer-readable medium of Claim 11 , wherein the one or more oral cognitive assessment includes a device hosting an Al agent initiating an audio communication session with the patient.
13. The non-transitory computer-readable medium of Claim 11 and / or 12, wherein the one or more oral cognitive assessment includes an interaction primer which comprises an oral narrative selected from a library of clinically validated narratives.
14. The non-transitory computer-readable medium of Claim 11 and / or 12, wherein the one or more oral cognitive assessment includes categorical or letter verbal fluency, short-term memory word recall, text descriptions, and picture descriptions.
15. The non-transitory computer- readable medium of any one of Claims 11, 12, 13, and / or 30, wherein the one or more oral cognitive assessments include delivering one or more immediate recall audio prompts for the patient to orally recall information from the oral narrative.
16. The non-transitory computer- readable medium of Claim 11 and / or 14, wherein the one or more oral cognitive assessment further includes engaging the patient in an oral1611.002W01topic discussion phase incorporating targeted follow-up questions dynamically generated based on the patient’s responses.
17. The non-transitory computer-readable medium of any one of Claims 11, 12, and / or 13, wherein the one or more oral cognitive assessment further includes delivering one or more delayed recall audio prompts for the patient to orally recall information from the oral narrative.
18. The non-transitory computer-readable medium of any one of Claim 11 and / or 17, wherein the backend analysis server generates, using an automatic speech recognition (ASR) model, a time-aligned transcript from the collection of audio data, wherein words in the transcript are associated with corresponding audio data timestamps from the collection of audio data.
19. The non-transitory computer- readable medium of Claim 11 and / or 18, wherein the plurality of features derived from the collection of audio data further includes features derived from the time-aligned transcript, and wherein the plurality of features comprises prosodic features, acoustic characteristics, lexical features, and semantic- coherence metrics.
20. The non-transitory computer- readable medium of Claim 11 and / or 19, wherein the generated at least one cognitive assessment score is based on a plurality of speech- derived features derived from the collection of audio data.