Systems and methods for evaluating medical conditions and / or diseases from video and multimedia resources

The system leverages social media and clinical records with AI/ML to provide instantaneous and accurate medical condition assessment, addressing resource limitations and clinician bias in traditional clinical settings.

WO2026094052A1PCT designated stage Publication Date: 2026-05-07BLUMROZEN GAD YEHOSHUA
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BLUMROZEN GAD YEHOSHUA
Filing Date
2025-11-02
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Clinical evaluation resources, such as clinical facilities and expert clinicians, are limited, leading to inefficiencies in medical diagnostics. Clinical records are cumbersome to collect, biased, and often collected in non-natural environments, affecting diagnosis accuracy and reliability. Social media data, which is rich and reflects real-life scenarios, has been underutilized for medical assessment.

Method used

A system and method using computerized models to analyze social media, clinical records, and direct multimedia streams to determine medical, cognitive, and mental conditions, leveraging artificial intelligence and machine learning to provide near-instantaneous analysis, reducing clinician bias and travel costs.

Benefits of technology

Enables accurate, reliable, and instantaneous medical condition assessment using social media data, overcoming clinician bias and travel limitations, and enhancing diagnostic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050965_07052026_PF_FP_ABST
    Figure IL2025050965_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods provide for the detection of medical, cognitive, and / or mental conditions and / or diseases (hereinafter "conditions and / or diseases") in humans and animals. The systems and methods utilize computerized models to determine (e.g., recognize) whether the human or animal subject has the condition and / or disease, the level or severity of the determined condition and / or disease, and / or the category of the medical condition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Docket: 602575

[0002] SYSTEMS AND METHODS FOR EVALUATING MEDICAL CONDITIONS AND / OR DISEASES FROM VIDEO AND MULTIMEDIA RESOURCES

[0003] CROSS-REFERENCES TO RELATED APPLICATIONS

[0004] This application is related to and claims priority from commonly owned US Provisional Patent Application Serial Number 63 / 714,190, entitled: Systems And Methods For Evaluating A Medical, Cognitive And / Or Mental Condition Of A Subject, filed on 31 October 2024, the disclosure of which is incorporated by reference in its entirety herein.

[0005] TECHNICAL FIELD

[0006] The present disclosure is directed to detecting psychological and mental conditions and / or diseases in humans and animals.

[0007] BACKGROUND

[0008] Human diseases and / or (medical, cognitive, and mental) conditions are typically determined by conventional testing methods and evaluations of clinical records from medical professionals and / or clinicians, such as doctors, neurologist, psychiatrists, therapists, social workers, psychologists, and other medical professionals. The clinical records, typically including interviews, physical examinations, self-scores, and test results, can be examined further by medical professionals, to determine the existence of a suspected condition or disease, typically in a clinical setting. Moreover, results of the determination of a condition or disease are not available instantly.

[0009] SUMMARY

[0010] The present disclosure is directed to systems and methods for detection of medical, cognitive, and / or mental conditions and / or diseases (hereinafter “conditions and / or diseases”) in humans and animals. The system utilizes computerized models to determine (e.g., recognize) whether the human or animal subject has the condition and / or disease, the level or severity of the determined condition and / or disease, and / or the category of the medical condition. The system can substitute or add to the clinical testing and to the clinical records analysis and assist or determine the condition and / or disease, by using direct subject conventional multi-media (video, audio and / or text) streams of the subject from video cameras such as the subject or the clinician, or indirect video, audio and text obtained from the subject or online, such as via social / or text from social media, such as Tick Tock™, Facebook®, Docket: 602575

[0011] Twitter® (now X), Instagram™, Snapchat™, Pinterest™, Yelp™, other reviewing sites, and the like.

[0012] Information on the subject, which resides in the world wide web (WWW), including in social media, additionally filmed video, audio recordings, text and clinical records, upon availability, is fed (input) into computer computation resources, and with the use of analytical, statistical, and artificial intelligence (Al) models, to provide a near instantaneous analysis and determination of conditions and / or diseases in the subject. The analysis exploits the intermedical condition correlations, with context estimation, and thus excludes the need for controlled environments, that do not reflect the real-life scenarios, and use the rich information on the subject that resides in the multi-media on the Internet and social media. The resultant determination or likelihood of the subject having a condition and / or disease, is delivered onsite, directly to the subject or clinician. The result, for example, excludes clinician and subject biases, saving traveling time and costs, clinical site utilities, and efforts.

[0013] The present disclosure is directed to a method, for example, a computer implemented method, for determining the presence of at least one condition and / or disease in a subject. The subject, for example, is a human, and animal, or a robot. The method comprises: processing at least one video associated with the subject into multimodal input for one or more feature extractors of a prediction model; generating at least one label associated with the subject to train the prediction model; training a prediction model with the at least one label, the prediction model for outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease; and the prediction model analyzing the output of each of the one or more extracted features, and outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease.

[0014] The present disclosure is directed to a system, for example, a computer system, for determining the presence of at least one condition and / or disease in a subject. The subject, for example, is a human, and animal, or a robot. The system comprises: a trainable prediction model including one or more feature extractors and configured for analyzing the output of each of the one or more feature extractors, and outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease; a first processor for processing at least one video associated with the subject into multimodal input for one or more feature extractors of a prediction model; a second processor associated with the Docket: 602575 generation of at least one label associated with the subject to train the prediction model; and a third processor for training the prediction model with the at least one label.

[0015] BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Some embodiments of the present disclosure are herein described, by way of example only, with reference to the accompanying drawings. With specific reference to the drawings, it is stressed that the particulars shown are schematic and are by way of example and for purposes of illustrative discussion of embodiments of the disclosure. In this regard, the description taken with the drawings makes it apparent to those skilled in the art how embodiments of the disclosure may be practiced.

[0017] Attention is now directed to the drawings, where like reference numerals or characters indicate corresponding or like components. In the drawings:

[0018] FIG. l is a block diagram of a process for determining the existence of a disease and / or condition and the level or severity of the detected disease and / or condition;

[0019] FIG. 2 is a block diagram of a prediction model for the process of FIG. 1; and

[0020] FIG. 3 is a block diagram of an example system on which the disclosed processes, including those of FIG. 1, are performed.

[0021] DETAILED DESCRIPTION OF EMBODIMENTS

[0022] One or more specific and / or alternative embodiments of the present disclosure will be described below, some with reference to the drawings, which are to be considered in all aspects as illustrative only and not restrictive in any manner. It shall be apparent to one skilled in the art that many alternative embodiments to those depicted in the drawings or those described below are possible, within the scope of the teaching of this disclosure. It should also be noted that to provide a concise description of these embodiments, not all features or details of an actual implementation are described at length in the specification. It is further noted that elements depicted in the drawings are not necessarily to scale, or in correct proportional relationships.

[0023] Before explaining at least one embodiment of the disclosed subject matter in detail, it is to be understood that the disclosure is not necessarily limited in its application to the details of construction and the arrangement of the components and / or methods set forth in the following Docket: 602575 description and / or illustrated in the drawings. The disclosure is capable of other embodiments or of being practiced or carried out in many ways.

[0024] As will be appreciated by one skilled in the art, aspects of the present disclosure may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “circuit,” “module” or “system.” Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more non-transitory computer readable (storage) medium(s) having computer readable program code embodied thereon.

[0025] OVERVIEW

[0026] The shortage of available high-quality clinical evaluation resources, such as clinical facilities, expert clinicians, patient travel time, and clinic space and time limitations, limit clinical procedure efficiency. To analyze a subject, and to compare the subject’s past records to a population, requires an accurate, high-quality, diverse clinical database. Limited clinical databases restrict medical diagnostics downstream. Clinical evaluation, and subject medical records collection into medical databases are often limited to controlled non-natural environments. Additionally, they are restricted due to privacy limitations and require complex scoring procedures that ultimately result in clinician bias.

[0027] Also, clinical records, from which these data sets are made, are difficult to obtain due to privacy regulations, their limited time of availability, they involve in tedious procedure to obtain them, and may be located off subject site, leading to further delays in obtaining these records. Additionally, there are typically costs associated with obtaining and copying the records, along with paperwork, which is completed on site or remotely, which causes delays in obtaining the records. Moreover, the records may be released only in person, and only to the subject himself. Also, simply collecting sufficient clinical records and a sufficient amount thereof, is a cumbersome, time intensive, and sometimes costly process, and furthermore, the existing medical records are biased due to the subjective evaluation process, and do not reflect real-life behavior and affect the diagnosis process.

[0028] Guaranteeing sufficient clinical datasets pertaining to various clinical conditions is essential, as a lack of sufficient databases can adversely affect the medical care. The quality of medical records can affect the way databases represent real-world conditions and limit and Docket: 602575 affect the accuracy of the diagnosis. For example, some existing clinical medical records are collected in non-natural environments and thus do not always reflect the subject medical condition. This also makes the accuracy and reliability of these clinical databases uncertain. Additionally, collecting clinical medical records is a cumbersome process, typically requiring input from medical professionals, and can suffer from rater bias, and subjectivity. Privacy limitations limit the wide use of datasets for many research domains, as is the case for instance with the use of video recordings exposing conditions and diseases, such as individual Parkinson Disease (PD) in subjects.

[0029] Outside of these clinical records, there is a lack of databases which are usable to replace these clinical records to perform the analysis and or determination of diseases or conditions.

[0030] Multimedia data streams (video, audio, and text) provide informative data that can be used as behavior capturing tools and produce behavioral markers for medical condition assessment. Automatic analysis of video recordings in the clinic or other clinical setting, can provide information about severity of disease such as Parkinson’s Disease (PD) and dementia. Vocal recording can be used for early detection of neurological disease such as Parkinson’s disease, for example, by extracting the subjects’ mood, psychological profiling, and medical condition analysis. Text mining can also be used to extract medical conditions such as cognitive decline. However, this data, obtained from clinical settings, is limited, and biased. Moreover, it requires the patient to expend the time, effort and costs of travel time, as well as medical (clinical) staff time.

[0031] To provide an accurate and reliable source of data, which has been underutilized, if even used at all, to date, to analyze and determine conditions and / or diseases, the present disclosure is directed to using social media, with or without clinical records, to analyze and determine conditions and / or diseases. For example, the conditions and / or diseases include Attention- Deficit / Hyperactivity Disorder (ADHD) and Autism Spectrum Disorder (ASD), Parkinson’s Disease, dementia, schizophrenia, bipolar disorder, and other neurologic, mental and cognitive disorders, and behavioral deficits.

[0032] In addition to social media, filmed video, audio and previously drafted text, including answers to questionnaires (self-reports), and other material written by the subject are useful in determining a condition and / or disease according to the present disclosure.

[0033] Social media includes massive amounts of information in video, audio files and text, which currently has not been fully exploited for medical assessment. Presently, in the medical Docket: 602575 field, automatic assessment of medical conditions is still extremely limited, as the information suffers from incorrect labeling, requiring double checking of the data, verifying that the data is real and not artificial (boot), and also, may have been subject to various formatting and recording conditions. In particular, there is not any research to date, which assists in analyzing and / or detecting common neurological diseases like ADHD and ASD where continuous monitoring in a home environment is important for optimizing treatment plans.

[0034] As disclosed herein social media (SM) information include information that can be utilized for continuous assessment of subject behavior and interaction type that can be used for medical condition and / or disease assessment. One advantage of using social media data is the ubiquity of data found in social networks. Social media products, like multimedia, are typically filmed / produced in real time and none controlled environment, and thus reflect daily life experience, and therefore have potentially less clinical bias, as it is usually generated in natural home settings. Additionally, the social media data can be used to extract communication of the subject of interest with other subjects and objects in the scene. The data can be used to extract significant labels and prior knowledge that can be then aggregated with the official clinical records, self-reports, tests, and clinical evaluation in the prediction model. The labels can be scale, or descriptive sentences, acquired from social media data harvesting using multimedia agent (like LLM based) inquiries. Thus, social media as disclosed herein, is used in continuous evaluation of subject condition and treatment efficiency, and continuous optimizing drug dose with minimal clinical assistance and efforts. As disclosed herein, it has been found how to extract social media data, to ensure its reliability, privacy, and accessibility, in order to benefit from using social media.

[0035] The social media, additionally filmed video, audio, and / or additional text, and clinical records, are input into computer models, such as machine learning (ML) models that use artificial intelligence, to provide a near instantaneous (such as contemporaneous) analysis and determination of a disease and / or condition. This result can be delivered on-site, directly to the subject or clinician, saving time and travel costs and efforts.

[0036] DETAILED DESCRIPTION

[0037] FIGs. 1 and 2 show block diagrams as flow diagrams of an example system 100 and example process performed by the system 100. The system 100 utilizes direct user-driven input 110 and system driven input (indirect input) 120. Some direct input and sometimes social media (from the system driven input 120), for example, a video stream or video, is Docket: 602575 demultiplexed 130, with the demultiplexed video, audio, and / or text, known as multimodal input, being input to a prediction model 160. Some user driven input 110 and system driven input 120, which is processed, including by being labeled, by direct clinical utility (human or computerized utility) (arrow 132) and labeled for machine learning (ML) using large language models (LLM) 140 and the data is sent to the prediction model 160. The prediction model 160 outputs a recognition of a condition and / or disease and the severity of the recognized condition and / or disease. The prediction model 160 is a trained and / or trainable model. For example, the prediction model 160 is also trained to recognize multiple conditions and / or diseases for the subject being analyzed.

[0038] Initially, there is a subject 102, to be evaluated for the conditions and / or diseases, including medical, cognitive, and mental conditions and / or diseases, which are typically human conditions and / or diseases, but may also be animal conditions and / or diseases as well. The subject 102 may also be a robot.

[0039] The subject 102 has provided an investigator (INV) 103, which is typically a machine utility but can be human, which enables the access to the multimedia files and permissions. Typically, the materials provided by the subject 102 are in electronic (digital) formats. The investigator 103 is linked to the Internet 122 (over a communications network), so as to be able to obtain social media and other digital materials from the Internet, for example, by computerized search agents 314 (FIG. 3). The subject 102 also links to the Internet 122, to provide the investigator 103 any passwords, codes, or other privileges, needed for full access to the subject’s social media, accounts, or other electronic materials, for example, both public and private to the subject 102.

[0040] The process can also be initiated by the subject clinician or computerized agent 104, for example, a medical doctor, psychiatrist, psychologist, or other clinician, for example, in the field of neurologic and or behavioral conditions and / or diseases. The clinician 104 has worked with the subject 102 and serves to label various aspects of the subject 102, with clinical direct labeling (e.g., arrow 132).

[0041] At block 110 user driven input, also known as “active input”, associated with the subject 102 is obtained. This user driven input includes user-based video streams 112, and text, for example, self-reports 114 and clinical records 116. Text 114, such as self-reports, including questionnaires and other written materials authored by the subject, and the like, in electronic formats. There are also clinical records 116, in electronic form, typically medical or clinician Docket: 602575 based conventional records, of the subject 102, which include, for example, medications prescribed to the subject and doses for each medication, and history of the medication, physical and psychiatric records, written by doctors, psychiatrists, psychologists, clinicians and the like, hospitalization and clinical visit records, transplant records, diagnostics, prescriptions (past and present), blood tests, urine tests, electrocardiograms, X-rays, Computerized Tomography scans, magnetic resonance imaging (MRI) scans, and the like. The clinical records are typically confidential (e.g., protected under Health Insurance Portability and Accountability Act of 1996 (HIPP A) laws in the United States), which require subject (user) permission for access thereto.

[0042] There is also system driven input, known as “passive” or “indirect input”, for example multimedia content 120, including social media, as well as other media generated by or about the subject 102, such as photographs, videos, articles, papers, essays, audio files, and the like, generated by or about the subject 102 in non-clinical settings, for example, real world settings in real time. Social media may include video streams, audio streams / audio files (with the video streams or separate therefrom) and text, with the video and / or the audio stream, or alone. The social media may, for example, be from social media web sites such as, but not limited to, Tick Tock™, Facebook®, X™ (formerly Twitter™), Linkedln®, Instagram™, Pinterest™, Yelp™ and other reviewing sites. The social media is typically discovered by searching techniques via search engines on networks such as the world wide web (WWW) including the Internet 122 and other public and / or Wide Area Networks.

[0043] Both user-driven input 110 and the system driven input 120 streams are processed, at blocks 130 (demultiplexing by a demultiplexer 316 (FIG. 3)) and 140 (labeling by a computerized labeling agent 312 (FIG. 3)). This processing, detailed further below, allows the user driven input 110 and system driven input 120 to demultiplex the data to multimedia streams (video frames, audio, and text), known as multimodal data, extract descriptive measures like labels of the multimedia streams, and extract the priors for the prediction models, all of which are needed to train and test the prediction model 160.

[0044] The user-based video stream 112, or video stream from multimedia from social media in block 120, is demultiplexed at block 130, for example, by a demultiplexer 316 (FIG. 3). The demultiplexing includes separating the video stream into video, audio and text components, to create multimodal data used as input (multimodal input) for the prediction model 160. For example, from the video component 112, the subject 102 may be further extracted therefrom, removing various background objects including other subject. From the voice component, Docket: 602575 voice recognition may be applied to identify the subject’s 102 voice and separate it from background noise or other subjects. Text can be then extracted using voice to text algorithms (transcription) as part of the 130 block.

[0045] Additionally, the video from the multimedia 120, including that from the social media, is, in some cases, demultiplexed into video, audio and text components (or multimodal data) by the demultiplexer 316 in accordance with the aforementioned demultiplexing 130.

[0046] The video data is then processed (by the demultiplexer 316) to multiple streams in block 130. The data is demultiplexed to video frames, text, voice, and the voice can be transcribed to conversational text using transcription models, for example, Amazon® Transcribe, from Amazon Web Services (AWS). Furthermore, the voice (for example, in an audio file), and the frames can further be processed (by a processor 316a associated with the demultiplexer 316) to separate user voice, for example, using Whisper™, an artificial intelligence (Al) powered speech-to-text model, available from OpenAI, and crop the image of interest, based on subject i’th priors if available from the scene using a body or a facial detection algorithm that identifies one or more face regions within an input image.

[0047] The processor 316a further executes a cropping module that automatically extracts the user (e.g., the subject) of interest in block 102, and a region of interest corresponding to the detected face (or body), and then generates a new image containing only the facial (or body) portion. In one embodiment, the facial detection algorithm employs a convolutional neural network (CNN) or a Haar Cascade Classifier (e.g., an algorithm that detects objects in images, irrespective of their scale in the image and location) to determine the bounding box coordinates of the face (or body). The cropping module utilizes these coordinates to define the cropping boundaries and outputs a high-resolution cropped facial image.

[0048] The system 100 may optionally include one or more pre-processing modules (which may be part of or associated with the demultiplexer 316 and processor 316a, which is not shown, and preprocessor 210), which includes a scaling module for example, contrast and lighting adjustments, and a storage module (storage media 306 in FIG. 3) for example, saving the processed facial images or features or prediction model 160 weights, and configuration to a database or transmitting them to an external system. In another embodiment, the system 100 is implemented as software executable on a general-purpose computer, mobile device, or cloudbased platform. The software may be integrated with existing image management systems or social media applications to enable automatic face cropping and tagging operations. The Docket: 602575 described method improves the accuracy, consistency, and speed of face (facial) extraction compared to manual image editing techniques.

[0049] Labeling

[0050] Labeling is important for supervised machine learning models. It includes severity of the medical condition, like severity of ADHD in scale of 0-9, category or sub-class of the ADHD like hyperactivity or attention-based ADHD, or in case of dementia, dementia type like Alzheimer or Lewy body dementia (LBD), and / or recognition of the medical condition, usually 0, or 1. All the labels are input into the prediction model 160. The labels include, for example, clinically generated labels 132 and as per block 142 (direct labeling from self-reports and clinical records) and block 144 (indirect labeling from multimedia including social media), are also used for training the prediction model 160.

[0051] The clinically generated labels of arrow 132 and the direct labels of block 142 are “direct labels”, while the labels of block 144 for the multimedia 120 are “indirect labels”. While a minimum of one label, such as the clinically direct label 132 can be sufficient to train the prediction model 160, the more labels, also including the direct 142 and / or indirect 144 labels in training the prediction model 160 increases the accuracy of the prediction model 160, based on two or three types of labels being input into the prediction model 160 at or close to the same time.

[0052] Clinical Direct Labeling

[0053] The user video stream 112 is, for example, labeled by the clinician 104, for example, at the arrow 132. The clinician 104 is typically human, but can be also a software utility, including a robot, when possible, provides a label as to a condition and / or disease or a manually assigned score for the disease or condition. The clinical direct label 132 is typically the result of a clinician or software utility or robot examination, including, for example, a physical examination, of the subject 102 or an examination of documents, including records and medical records of the subject.

[0054] The score is typically based on an accepted standardized or normalized scale, for example Unified Parkinson’s Disease Ranking Scale (UPDRS) for Parkinson’s disease or Diagnostic and Statistical Manual of Mental Disorders (DSM) for psychiatric disorders and diseases, Test of Variables of Attention (TOVA) for ADHD, ADHD-Adult Self Report Scale (ASRS), for subjects suspected of having ADHD, and Montreal Cognitive Assessment Docket: 602575

[0055] (MOCA) for dementia, so as to be within acceptable ranges for other scores obtained from computerized agents or devices. However, the score assigned may be subject to clinician bias, where the clinician is not familiar with the subject’s culture or environment, or if by an agent (software utility), might be biased due to unbalanced or small trained data. Additionally, the clinician may be subjected to bias, such as ethical, social, and the like. Also, for example, the subject under examination may exhibit subject bias where the subject overacts to suggest evidence of the disease and / or condition, or underact, to avoid behaving like someone with the suspected condition and / or disease.

[0056] Other clinical direct labeling 132 occurs when a clinician 104 examines, observes, or interviews the subject 102 and indicates the appearance of the existence and / or severity of the condition and / or disease, which is the label 132, and inputs this label 132 into the prediction model 160.

[0057] Computerized Labeling

[0058] In addition, or instead of the clinical evaluation, which is not always available, the system 100 provides a computer generated direct labeling, for example, by a computerized labeling agent 312 (FIG. 3) using the self-reports (including computerized tests) 114 and clinical records 116, at block 142. Subject 102 medical condition labeling by direct labeling is performed at block 142. The self-reports 114 and clinical records 116 are labeled by computerized processes, like tabular neural network, optical character recognition (OCR), to extract labels of description of the subject condition and is performed by computerized labeling agents 312. The labels include, for example, a score, descriptive text, of the condition and / or disease, at block 142.

[0059] The subject self-report scores, like ASRS self-report for ADHD, or computerized tests like Tova for ADHD, which, for example, are used to assess the labels directly according to their user manual instruction. The self-reports 114 include, for example, written or computerized tests, typically standardized tests, such as ADHD-ASRS, for subjects suspected of having ADHD, MOCA for dementia, UPDRS for Parkinson’s disease, and TOVA for ADHD. The self-reports are for example typically in digital (electronic) form or converted to digital form.

[0060] The clinical records are assigned a label in the form of a score, sometimes on the normalized scale, as well as category, recognition - Yes / No for the condition and / or disease, or sometimes descriptive text of the subject. Clinical records include, for example, medications prescribed to the subject and doses for each medication, and history of the medication, physical Docket: 602575 and psychiatric records, written by doctors, psychiatrists, psychologists, clinicians and the like, hospitalization and clinical visit records, transplant records, diagnostics, prescriptions (past and present), blood tests, urine tests, electrocardiograms, X-rays, Computerized Tomography scans, magnetic resonance imaging (MRI) scans, and the like.

[0061] For both the self-reports 114 and the clinical records 116, the labeling may include labels for the recognition of the condition and / or disease, the severity (related to the subject medical scores) of the condition and / or disease, the class or type of the condition and / or disease (e.g., ADHD may be hyperactivity or attention deficit), and descriptive text (e.g., words, phrases, word sequences) of the subject condition and / or disease.

[0062] Computer agent indirect labeling for the multimedia is performed at block 144, with indirect labels being, for example, for the recognition of the condition and / or disease, the severity (including scoring) of the condition and / or disease, the class or type of the condition and / or disease (e.g., ADHD may be hyperactivity or attention deficit), and descriptive text (e.g., words, phrases, word sequences) of the subject condition and / or disease. The social media is labeled with an indirect label, at block 144 by computerized processes, for example, by one or more computerized labeling agents 312 (FIG. 3). Labeling of the multimedia 120 video, audio and text, including the video audio and text from the social media, is performed automatically by the computerized labeling agent 312. The labeling can be based on large language model (LLM) processing of the multimedia data, such as video clips self-labeling of subject’s experiences in the video keywords or description. Such self-labeling is mostly of subjects that are clinically diagnosed with a known disease or condition, such as ADHD. For example, with the case of ADHD, the severity is believed to be correlated with the self-labeling under assumption that the subjects want to share with their circle of followers their true experience.

[0063] To validate the data labels, the clinical, direct, and indirect score from blocks 132, 142, and 144, can be cross validated with real subjects and / or with Artificial Intelligence (Al). The block 144 may include also identify the reliability of the multimedia data, and / or verifying whether the data is authentic, to ensure label noise will not be produced when the system 100 compares auto labeling and cross validation with the data content.

[0064] The different labels of the subject and / or of the specific video can be aggregated to one label to train the model by statistical algorithm that incorporates the reliability of the label from the source, the consistency of the labels, and the bias that the labeler might have. The bias can be or example clinician bias due to cultural gap or inaccurate sensing quality, or in case of Docket: 602575 indirect labeling multimedia bias due to fake videos. The aggregation itself can be choosing one label, weighted sum of the labels, majority of voting, or any type of regression that can be trained on similar data.

[0065] The aggregation can be also performed after training the prediction of each model separately, and then aggregated the predictions, using methods like above, with regression that their weights can be trained on multiple subjects in a population.

[0066] Priors to the Prediction Model

[0067] A computerized model in block 144 calculates, from the multimedia inputs using rule based, classical prediction model, or pretrained neural network models, indirect priors known also as “priors”, PD- indirect medical priors of subject i from block 120. The priors are input to the context analysis block 164 in the prediction model 160 to produce the content.

[0068] Obtaining content from the at least one video obtained from the multimedia or additional different video from the multimedia, and producing certain of the obtained content as priors, inputting the priors into a context analyzer to produce context of video and the interaction of the subject with the environment for training the prediction model. These priors can be aggregated with the data from the extracted features, i.e., blocks 162a-162d, and to the prediction model itself to provide information which may explain a condition and / or disease, or lack thereof, and assist in training context dependent prediction model to enhance the prediction model accuracy.

[0069] The priors can include interactions with subjects and objects in the scenes of the video (subject provided or from social media), or may also include metadata like location of the video, or the time when they were taken, or textual description of the scene or the person extracted from utilities like image to text or by using LLM analysis of the multimedia text. The priors are extracted by data harvesting model inside block 144. For example, the data from the multimedia 120, includes posts of the subject 102 that may be analyzed by data harvesting from a type of LLM in block 144, and produce insights about the subject’s 102 state of ADHD following the subject’s 102 patterns in the text. Priors include subject matter in video scenes, metadata of the multimedia 120, including the social media, text from the multimedia 120 taken at different times, comments made by the subject 102 on social media, reviews made by the subject on social media (e.g., Yelp™) tagging of social media, such as tagging certain social media for ADHD. The priors also include, for example, the subject smiling, the number of Docket: 602575 people the subject is associating with, the activity the subject is doing, physical features of the subject, such as reddish skin, indicative of a potential heart or circulation issue.

[0070] For example, priors from social media may also include how many times a user )(subject 102) posts, the frequency of the posts, and the time period over which the posts were made. Also, it is looked at whether the user (subject 102) stopped posting for a time period, and the subject matter of the posts. The text of the posts may be looked at over time for grammar and spelling mistakes. Also, for example, with Facebook®, the adding or deletion of friends and the number of followers and any changes to these numbers. Additionally, the subject’s social media comments to posts are looked at for their substance, tone, spelling and grammar.

[0071] For example, if looking for autism in a child, who is the subject 102, a video of a soccer game may be used. Should the child not be near the ball and not even looking at the ball and the cluster of players around the ball, this may be further evidence of, or a prior suggesting, autism. However, if the child is always close to the ball and reacting to the travel of the ball, this would be a factor of slight or a prior of negligible or lack of autism.

[0072] Also, for example, there may be a video of two people running a marathon. If the subject 102 is talking with another runner, he may be a socially active person. If the subject 102 is just running and not interacting with any other runner, the prior may be inconclusive.

[0073] Prediction Model

[0074] The now z’th subject video streams (multi-modal data streams of image frames, audio streams, and text, also represented as block 166 in FIG. 2), are input into the prediction model 160, which is, for example, a multimodal prediction model. The prediction model 160 can include dedicated feature extraction at block 162, called fundamental features, which are then input to prediction model block 210.

[0075] The prediction model 160 has multiple prediction outputs of the different medical condition and / or disease, which can be further used as features for the prediction of specific medical conditions and / or diseases, and enable exploitation of the correlation between the different medical conditions and / or diseases.

[0076] The i’th subject labels are also input into the prediction model 160. These subject labels include the - clinician medical diagnostics scores of subject i, Llc(arrow 132), the direct medical labels of subject i, LlD, derived from the self score and medical records at block 142, Docket: 602575 and the indirect medical scores of subject i, LlID, derived from the multimedia data at block 144, and the PjD- indirect medical priors of subject i.

[0077] Each of the multi-modal data inputs from each of the blocks 130 (demultiplexed video 112 and demultiplexed video from the multimedia 120, including social media), 132, 142, 144 (labels) is subject to the process at block 162, and includes the context analysis (based on the priors from the labeling block 140, or from the video demultiplexing 130) at block 164 and are represented in FIG. 2 as block 166. These inputs are analyzed in the prediction model 160 for a one or multiple conditions or disease, such as, for example, Attention- Deficit / Hyperactivity Disorder (ADHD) and autism spectrum disorder (ASD), Parkinson’s Disease, dementia, schizophrenia, bipolar disorder, and other neurologic disorders and behavioral disorders.

[0078] The features extracted include, but are not limited to, for example, vocal features 162a, body features 162b, eye features / facial features 162c and general data driven features 162d.

[0079] The vocal features 162, include, for example, the features which are extractable by / from pre-trained models, or direct from the audio file (from demultiplexed video, as provided by the subject 110 or from the social media 120) and is, for example, from the following groups: a) Spectral Features; b) Mel-Frequency Cepstral Coefficients (MFCCs); c) Pitch and Prosody; d) Timbre and Tonal Features; and e) features of interactions with another subject(s) and / or object(s). Vocal features may also include, for example, voice inflections, accents, pronunciation of words and phrases, grammar, diction, speed of speech, and the like.

[0080] The body features 162b are, for example, skeletal based features (from a skeletal extraction pre-trained model, such as MediaPipe™ open source framework from Google of Mountain View, California), or other features embeddings from other pre-trained network of computer vision model, trained on activity detection, or direct features like body part velocity, obtained from computer vision models. Body features may also include, for example, gestures type, controlled and uncontrolled movements, tremor features like intensity and frequency, stability, and the like.

[0081] The facial features / eye features 162c are extractable from either rule-based models, a Mult expert domain model, and / or pretrained models, that extract, for example, facial landmarks, Action Units (AU), or emotional expressions. Facial landmarks are extracted from the subject’ 103 facial region of interest (ROI) in each frame ( v” ). A robust landmark Docket: 602575 detection algorithm, such as OpenF ace’s 68-point 3D model, is used to identify key facial points, including the eyes, nose, mouth, and jawline. OpenF ace, an open-source facial behavior analysis toolkit, employs a pre-trained Constrained Local Model (CLM) combined with a 3D Morphable Model (3DMM) to detect landmarks and estimate 3D face shape, enabling accurate pose and depth analysis. Eye features may also include, for example, gazing, following, cross- eyes, lazy eyes, chaotic or non-chaotic movements, and the like.

[0082] Action Units (AUs), based on the Facial Action Coding System (FACS), are derived from facial landmarks to quantify specific facial muscle movements (e.g., AU1 : inner brow raiser, AU12: lip comer puller). Using the pre-trained Constrained Local Model (CLM) and 3D Morphable Model (3DMM) in OpenF ace, the 3D facial landmarks (L” = from frame (v-1) are processed to estimate the intensity of each AU.

[0083] Facial expressions, representing emotional states such as happiness, sadness, anger, surprise, fear, and disgust, are derived from the face, facial landmarks, and Action Units (AUs) to quantify the subject’s emotional dynamics. Utilizing the pre-trained Constrained Local Model (CLM) and 3D Morphable Model (3DMM) framework of OpenF ace, facial expressions are computed for each frame ( v-1) of the analyzed video based on the 3D facial landmarks ( L”) and the AU ( AU-1) extracted from that frame. The emotion expression for each specific emotion is defined separately using FACS-based rules, with the 3D landmarks ( L” ) (where ( M = 68) for OpenF ace) providing geometric context to enhance accuracy by accounting for head pose and facial structure. The expressions are calculated as follows for each frame

[0084] General data driven or contextual features 162d, are, for example, based on histogram knowledge, or based on pre-trained Large Language Models (LLM) models that can extract personality-based features, mode features, memory-based features, and the specific medical condition features. These features can be derived from the LLM model that can be pre-trained (zero shout) like ChatGPT, or by user defined LLM like transformers. The LLM input can be the integration of specific medical condition from the text, or can be mode general features, or can refer to clinical guideline based features, like ASRS based features for ADHD. General data driven or contextual features include, for example, movement evaluations as to how smooth the movement is made, the type of activity in which the movement is made, and the like. Docket: 602575

[0085] Turning to the other portion of the fundamental features is the context 164, for example, a pretrained model embeddings or using pre-trained mode with transfer learning. The pretrained model is, for example, a machine learning model, which was trained previously for a different purpose than its use now, and its features are significant to describe subject conditions and / or diseases that are mainly data driven and are not always able to be interpreted. In case of neural network, the features of such pretrained models are sometimes referred to as Neural Network embeddings or latent space. Possible pretrained models may be, for example, indicative of human activity, as detected from the body images.

[0086] FIG. 2 provides additional details for prediction model 160, with the model core 200 shown in this figure. The prediction model 160 includes components comprising a preprocessing unit 210, an advance feature extractor 220, a classical machine learning (ML) model 230 and a deep learning (DL) neural network 240.

[0087] The prediction model 160 may be trained with data labels including one or more of, the clinical direct labels from arrow 132, the direct labels and / or scores from block 142 and the indirect labeled and / or scores, from block 144. Aggregation of multiple labels can be performed on label and prediction domains.

[0088] For the label domain, the different labels of the subject and / or of the specific video are aggregated to one label to train the model by statistical algorithms that incorporate the reliability of the label from the source, the consistency of the labels, and the bias that the labeler may possess. The bias, for example, may be clinician bias due to cultural gaps or inaccurate sensing quality, or in case of indirect labeling multimedia bias due to fake videos. The aggregation itself can be choosing one label, a weighted sum of the labels, a majority of voting, or any type of regression that can be trained on similar data.

[0089] For the prediction domain, the aggregation, for example, is performed after training the prediction of each model separately, and then aggregated the data in relation to the predictions, using Al and LLM methods as disclosed above, with regression that their weights can be trained on multiple subjects in a population.

[0090] For example, the aggregation includes taking a weighted average of each of the labels input into the prediction model 160 to determine a conditions of disease or the probable condition and / or disease, ranked list of the specific or probable conditions and / or diseases. Docket: 602575

[0091] Also, for example, the priors are aggregated with data labels, including the clinical, direct labels from arrow 132, the direct labels and / or scores from block 142 and the indirect label and / or scores, from block 144, to train the prediction model 160.

[0092] Initially, the fundamental features 166 are input to the model core 200. The preprocessor 210 functions to remove outliers from the video and audio streams. The preprocessor 210, for example, also functions to filter sample noise, movement effects, and the like, in the video, for example, by applying a Kailman filter, and / or supply compensatory missing data, to overcome discontinuities in the video, and optionally, may provide an additional normalization or standardization for the video.

[0093] The output of the preprocessor 210 is input into the feature extractor 220. In the feature extractor (advanced feature extractor) 220, the input is subject to deriving advanced features exploiting possible expert domain knowledge or output of a pretrained expert neural network. Also, the input data is standardized, for example, for the classical ML model 230 and the DL neural network 240.

[0094] The feature extraction is, for example, performed in the form of a spatial -temporal features extraction. For example, feature extraction using skeletal landmarks can be used to extract spatial -temporal features, to evaluate the subject’s activity profile and kinematics within an observation period. The spatial-temporal features, for example, for the body features, are divided to three groups: 1) intra-joints dynamics; 2) inter-joint synchronization features; and 3) spatial-temporal activity level features measured by Joint Area and Volume.

[0095] The advanced feature extraction (by the advanced feature extractor 220) can include expert domain knowledge group of features like statistical features, dynamic features, and temporal features.

[0096] Advanced extraction of features may be performed on the vocal features, the face / eye features, similar to that for the aforementioned body features, to determine, for example a condition and / or disease. For example, low body activity (body features) coupled with sad facial expressions and eye movements, is indicative of depression in the subject.

[0097] For example, for condition and / or disease detection related to body movements (body gestures) deficits, intra-joints dynamics are measured by the joins kinematics using linear and angular velocity, acceleration, and entropy, for example, as described in Blumrosen, et al., “A Real-Time Kinect Signature-Based Patient Home Monitoring System,” in Sensors, Vol. 16, Docket: 602575

[0098] No. 11, p. 1965, Nov. 2016, doi: 10.3390 / sl6111965, which is incorporated by reference herein. The inter-joint synchronization features are related to movement coordination, and can be estimated with a Pearson correlation coefficient, which estimates the correlation between different joints kinematics pattern over time The spatial -temporal activity level features are measured by joint area and volume. For this a convex hull is derived and employed to measure the physical area and volume of the joints' movements in 3D space. The convex hull is determined of a set of points, which, for example, is the smallest convex polygon that contains all the points. For example, the convex hull of a joint in a skeleton frame represents the area occupied by that joint in 3D space. The convex hull for each joint in each skeleton frame is a measure of the level and range of the joint's movements.

[0099] For all other cases (e.g., vocal features, face / eye features), the advanced features are optionally further filtered to remove outliers and then scaled for standardization.

[0100] Features in Observation Period

[0101] Behavioral based assessment of subject of movie duration includes different diverse scenes that can be exploited to enhance the overall accuracy. The analysis is tailored to capture the diversity of the scene to enable by using each sub-clip (representing a scene) with a different observation period, which is equivalent to clinical experiment repetitions. The observation period needs to be balanced between maximizing the samples diversity between the observations and yet the observation is long enough to capture the scene.

[0102] Features Selection and Feature Extraction

[0103] To avoid overfitting of the prediction model 160 due the relative potentially high number of features, dimensionality reduction, for example, as based SHapley Additive exPlanations (SHAP), for example, as described in S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2017. Accessed: Feb. 17, 2024. [Online], Available: https: / / proceedings.neurips.cc / paper / 2017 / hash / 8a20a8621978632d76c43dfd28b67767- Abstract.html, which is incorporated by reference herein, and Density -Based Spatial Clustering of Applications with Noise DBSCAN, for example, as disclosed in M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise,” which is incorporated by reference herein. The DBSCAN clustering algorithm based on pairwise Pearson coefficient correlation may be used first, followed by, for each feature in the cluster, obtaining their equivalent SHAP values, which represents on its Docket: 602575 contribution to the expected prediction, to select representative features for each cluster with the highest SHAP values. This approach allows for the reduction of the dimensionality of the feature set while retaining the most prediction and diverse features.

[0104] The output of the advanced feature extractor 220 is input for the classical machine learning (ML) model 230 and Deep Learning (DL) Neural Network 240.

[0105] Classical Machine Learning (ML) Model and Deep Learning (DL) Neural Network

[0106] The classical machine learning model 230 includes an algorithm, which learns patterns from data using statistical or mathematical methods, without relying on deep neural networks. These models form the foundations of machine learning and are widely used when the dataset is relatively small, structured, and interpretable. The classical machine learning model works by: 1) extracting features from the input data (e.g., mean, variance, or engineered descriptors);

[0107] 2) training a model that learns relationships between those features and the target output; and

[0108] 3) making predictions or classifications on new, unseen data. The classical machine learning model includes, for example, a Gaussian Naive Bayes, Support Vector Machines (SVM), XGBoost, linear or non-linear regression, clustering algorithms, and the like.

[0109] For the deep learning Neural Network 240 TabNet from Google is suitable for use.

[0110] To train the classical classifier 230 and the deep learning neural network 240, the data set, for example, is split into 1 :4 ratio between the test and training set, with a relatively equal gender balance, if possible. For the classical ML model 230 and the DL neural network 240, hyper parameter tuning is performed, to optimize performance of these components 230, 240. Also, the observation period can be optimized to reflect the condition and / or disease dynamics for enhancing the prediction model 160 performance.

[0111] The output of the prediction model 160 includes the recognition of the subject having or suspected of having the condition and / or disease and the severity thereof. The outcome typically also includes the type or class of the condition and / or disease, such as for ADHD, the type or class may be hyperactivity or attention deficit.

[0112] The output of the prediction model 160 may also include one or more treatment protocols for each detected condition and / or disease. Based on this protocol, one or more effective amounts of a composition, or pharmaceutical composition, exercises, or other treatments may be administered to the subject 102. Docket: 602575

[0113] The output of the prediction model 160 may also be a plurality of conditions and / or diseases (e.g., their recognition and severity), which are correlated, and / or associated with each other, based on recognition of at least one condition and / or disease and / or disease and the severity of the at least one recognized condition and / or disease. Thus, the total set conditions and / or diseases, and their dependencies, can enhance the prediction accuracy of each of condition and / or disease separately. For example, a subject evaluated with autism can be more likely to have ADHD based on the autism recognition.

[0114] Optionally, the output of the prediction model 160 may be sent back to the clinician 104 as part of a clinical decision support tool 170. This tool 170 aids the clinician 104 in future evaluations of the subject 102 as well as other subjects when evaluation a condition and / or disease. This output can also be transmitted directly or indirectly to the subject 102, tailoring a subject specific treatment plan, or behavioral inputs.

[0115] FIG. 3 is a block diagram of an example system 300 in accordance with the disclosure. The system 300 includes, for example, a Central Processing Unit (CPU / Graphics Processing Unit (GPU) 302, in communication with storage / memory 304, and storage media 306, such as databases and other storage. The CPU / GPU 302 is formed, for example, of processing circuitry, including one or more processors (computerized processors), such as microprocessors, to perform the processes of the disclosure.

[0116] The Central Processing Unit (CPU) 302 is formed of one or more processors (computerized processors), including microprocessors, and hardware processors, for obtaining data from the storage media and processing input to the system, as well as transmitting output from the system 300. The processors for the CPU 302 include, for example, conventional processors, such as those used in servers, computers, and other computerized devices. The processors, for example, may comprise general purpose computers, which are programmed in software, to carry out the functions described herein. These processes include, for example, analytical and comparison processes. The software may be downloaded to the computer in electronic form, over a network, for example, or it may, alternatively or additionally, be provided and / or stored on non-transitory tangible media, such as magnetic optical, or electronic memory, which may be part of the storage / memory 304. Docket: 602575

[0117] The GPU 302 is, for example, an electronic circuit which manipulates the storage / memory 304 to generate images for display on a screen. The GPU 304 also functions to deploy the Al and Machine learning models, as disclosed herein.

[0118] Within the CPU / GPU 302 are specific computerized functionalities of labeling agents 312, search agents 314, for example, those which search social media, demultiplexer 316 and associated processor 316a, preprocessors 210, feature extractor 318, advanced feature extractors 220, and the prediction model 160, including the classical machine learning model 230 and the deep learning neural network 240. The functionalities communicate with each other as well as the storage / memory 304 and the storage media 306 directly or indirectly.

[0119] The storage / memory 204 is, for example, any conventional storage media. The storage / memory 304 stores machine executable instructions for execution by the CPU / GPU 302, to perform the processes of the disclosure. The storage / memory 304 also includes machine executable instructions associated with the operation of the storage media 306.

[0120] The storage media 306 stores databased of video, audio files and digital documents, including text, self-reports, clinical records, and the like.

[0121] ASPECTS OF THE DISCLOSURE

[0122] The present disclosure is directed to a method, for example, a computer implemented method, for determining the presence of at least one condition and / or disease in a subject. The method comprises: processing at least one video associated with the subject into multimodal input for one or more feature extractors of a prediction model; generating at least one label associated with the subject to train the prediction model; training a prediction model with the at least one label, the prediction model for outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease; and the prediction model analyzing the output of each of the one or more extracted features, and outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease.

[0123] Optionally, the method is such that the at least one label includes a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject.

[0124] Optionally, the method is such that the generating the at least one label includes generating a plurality of labels including direct and / or indirect labels. Docket: 602575

[0125] Optionally, the method is such that the direct labels include at least one of: 1) a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject, or 2) a direct label from selfreports or clinical records; and an indirect label from multimedia.

[0126] Optionally, the method is such that the direct and / or indirect labels are input into the prediction model and aggregated to train the prediction model.

[0127] Optionally, the method is such that the direct labels from self-reports or clinical records, and the indirect labels from multimedia include one or more of recognition of one or more conditions and / or diseases, at least one category for each of the recognized conditions and / or diseases, or a score for each of the recognized conditions and / or diseases.

[0128] Optionally, the method is such that the at least one condition and / or disease output by the prediction model includes a plurality of conditions and / or diseases.

[0129] Optionally, the method is such that the prediction model recognizes a second condition and / or disease based on the recognition of the first condition or disease.

[0130] Optionally, the method is such that the multimodal input includes video, voice, and or text, from a demultiplexed video stream.

[0131] Optionally, the method is such that the one or more extracted features include vocal features, body features, facial / eye features and general data driven features.

[0132] Optionally, the method is such that the multimedia includes social media.

[0133] Optionally, the method is such that the at least one video associated with the subject is received from the subject or is obtained from multimedia.

[0134] Optionally, the method is such that the processing the at least one video into the multimodal input includes demultiplexing the video.

[0135] Optionally, the method is such that it additionally comprises: obtaining content from the tat least one video obtained from the multimedia or additional different video from the multimedia, and producing certain of the obtained content as priors, inputting the priors into a context analyzer to produce context of the video and interaction of the subject with the environment for training the prediction model.

[0136] Optionally, the method is such that the subject is human or animal, or robot. Docket: 602575

[0137] The disclosure is directed to a system, for example, a computer system, for determining the presence of at least one condition and / or disease in a subject. The system comprises: a trainable or otherwise trained prediction model including one or more feature extractors and configured for analyzing the output of each of the one or more feature extractors, and outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease; a first processor for processing at least one video associated with the subject into multimodal input for one or more feature extractors of a prediction model; a second processor associated with the generation of at least one label associated with the subject to train the prediction model; and a third processor for training the prediction model with the at least one label.

[0138] Optionally, the system is such that the first processor includes a demultiplexer for creating multimodal input for the at least one feature extractor.

[0139] Optionally, the system is such that the at least one label includes a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject.

[0140] Optionally, the system is such that the second processor includes a computerized labeling agent for generating the at least one label as a plurality of direct and indirect labels.

[0141] Optionally, the system is such that the direct labels include at least one of: 1) a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject, or 2) a direct label from selfreports or clinical records; and an indirect label from multimedia.

[0142] Optionally, the system is such that the third processor inputs the direct and indirect labels into the prediction model, and aggregates the direct and / or indirect labels to train the prediction model.

[0143] Optionally, the system is such that the direct labels from self-reports or clinical records, and the indirect labels from multimedia include one or more of recognition of one or more conditions and / or diseases, at least one category for each of the recognized conditions and / or diseases, or a score for each of the recognized conditions and / or diseases.

[0144] Optionally, the system is such that the at least one condition and / or disease output by the prediction model includes a plurality of conditions and / or diseases. Docket: 602575

[0145] Optionally, the system is such that the prediction model is configured to recognize a second condition and / or disease based on the recognition of the first condition or disease.

[0146] Optionally, the system is such that it additionally comprises: a demultiplexer for demultiplexing a video stream to generate the multimodal input including video, voice, and or text.

[0147] Optionally, the system is such that each of the one or more feature extractors extract features including vocal features, body features, facial / eye features and general data driven features.

[0148] Optionally, the system is such that the multimedia includes social media.

[0149] Optionally, the system is such that the at least one video associated with the subject is received from the subject or is obtained from multimedia.

[0150] Optionally, the system is such that it additionally comprises a fourth processor configured for: obtaining content from the at least one video obtained from the multimedia or additional different video from the multimedia, and producing certain of the obtained content as priors, inputting the priors into a context analyzer to produce context of the video and interaction of the subject with the environment for training the prediction model.

[0151] Optionally, the system is such that the subject is human, animal, or a robot.

[0152] The implementation of the method and / or system of embodiments of the disclosure can involve performing or completing selected tasks manually, automatically, or a combination thereof. Moreover, according to actual instrumentation and equipment of embodiments of the method and / or system of the disclosure, several selected tasks could be implemented by hardware, by software or by firmware or by a combination thereof using an operating system or a cloud-based platform (such as those provided by Amazon Web Services™ or Microsoft® Azure™).

[0153] For example, hardware for performing selected tasks according to embodiments of the disclosure could be implemented as a chip or a circuit. As software, selected tasks according to embodiments of the disclosure could be implemented as a plurality of software instructions being executed by a computer using any suitable operating system. In an exemplary embodiment of the disclosure, one or more tasks according to exemplary embodiments of method and / or system as described herein are performed by a data processor, such as a Docket: 602575 computing platform for executing a plurality of instructions. Optionally, the data processor includes a volatile memory for storing instructions and / or data and / or a non-volatile storage, for example, non-transitory storage media such as a magnetic hard-disk and / or removable media, for storing instructions and / or data. Optionally, a network connection is provided as well. A display and / or a user input device such as a keyboard or mouse are optionally provided as well.

[0154] For example, any combination of one or more non-transitory computer readable (storage) medium(s) may be utilized in accordance with the above-listed embodiments of the present disclosure. The non-transitory computer readable (storage) medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a readonly memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0155] A computer readable signal medium may include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0156] As will be understood with reference to the paragraphs and the referenced drawings, provided above, various embodiments of computer-implemented methods are provided herein, some of which can be performed by various embodiments of apparatuses and systems described herein and some of which can be performed according to instructions stored in non-transitory Docket: 602575 computer-readable storage media described herein. Still, some embodiments of computer- implemented methods provided herein can be performed by other apparatuses or systems and can be performed according to instructions stored in computer-readable storage media other than that described herein, as will become apparent to those having skill in the art with reference to the embodiments described herein. Any reference to systems and computer- readable storage media with respect to the following computer-implemented methods is provided for explanatory purposes and is not intended to limit any of such systems and any of such non-transitory computer-readable storage media with regard to embodiments of computer-implemented methods described above. Likewise, any reference to the following computer-implemented methods with respect to systems and computer-readable storage media is provided for explanatory purposes and is not intended to limit any of such computer- implemented methods disclosed herein.

[0157] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0158] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein. Docket: 602575

[0159] As used herein, the singular form “a,” “an” and “the” include plural references unless the context clearly dictates otherwise.

[0160] The word “exemplary” is used herein to mean “serving as an example, instance or illustration.” Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments and / or to exclude the incorporation of features from other embodiments.

[0161] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the disclosure. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments unless the embodiment is inoperative without those elements.

[0162] The above-described processes including portions thereof can be performed by software, hardware, and combinations thereof. These processes and portions thereof can be performed by computers, computer-type devices, workstations, cloud-based platforms, processors, micro-processors, other electronic searching tools and memory and other non- transitory storage-type devices associated therewith. The processes and portions thereof can also be embodied in programmable non-transitory storage media, for example, compact discs (CDs) or other discs including magnetic, optical, etc., readable by a machine or the like, or other computer usable storage media, including magnetic, optical, semiconductor storage, or other source of electronic signals.

[0163] The processes (methods) and systems, including components thereof, herein have been described with exemplary reference to specific hardware and software. The processes (methods) have been described as exemplary, whereby specific steps and their order can be omitted and / or changed by persons of ordinary skill in the art to reduce these embodiments to practice without undue experimentation. The processes (methods) and systems have been described in a manner sufficient to enable persons of ordinary skill in the art to readily adapt other hardware and software as may be needed to reduce any of the embodiments to practice without undue experimentation and using conventional techniques.

[0164] Although the disclosure has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to Docket: 602575 those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

Docket: 602575CLAIMS1. A method for determining the presence of at least one condition and / or disease in a subject comprising: processing at least one video associated with the subject into multimodal input for one or more feature extractors of a prediction model; generating at least one label associated with the subject to train the prediction model; training a prediction model with the at least one label, the prediction model for outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease; and the prediction model analyzing the output of each of the one or more extracted features, and outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease.

2. The method of claim 1, wherein the at least one label includes a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject.

3. The method of any one of claims 1 or 2, wherein the generating the at least one label includes generating a plurality of labels including direct and / or indirect labels.

4. The method of any one of claims 1 to 3, wherein the direct labels include at least one of: 1) a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject, or 2) a direct label from self-reports or clinical records; and an indirect label from multimedia.

5. The method of any one of claims 1 to 4, wherein the direct and / or indirect labels are input into the prediction model and aggregated to train the prediction model.

6. The method of any one of claims 3 to 5, wherein the direct labels from self-reports or clinical records, and the indirect labels from multimedia include one or more of recognition of one or more conditions and / or diseases, at least one category for each of the recognized conditions and / or diseases, or a score for each of the recognized conditions and / or diseases.Docket: 6025757. The method of claim 1, wherein the at least one condition and / or disease output by the prediction model includes a plurality of conditions and / or diseases.

8. The method of claim 7, wherein the prediction model recognizes a second condition and / or disease based on the recognition of the first condition or disease.

9. The method of claim 1, wherein the multimodal input includes video, voice, and or text, from a demultiplexed video stream.

10. The method of claim 1, wherein the one or more extracted features include vocal features, body features, facial / eye features and general data driven features.

11. The method of any one of claims 4 or 6, wherein the multimedia includes social media.

12. The method of claim 1, wherein the at least one video associated with the subject is received from the subject or is obtained from multimedia.

13. The method of claim 12, wherein the multimedia includes social media.

14. The method of claim 1, wherein the processing the at least one video into the multimodal input includes demultiplexing the video.

15. The method of claim 1, additionally comprising: obtaining content from the tat least one video obtained from the multimedia or additional different video from the multimedia, and producing certain of the obtained content as priors, inputting the priors into a context analyzer to produce context of the video and interaction of the subject with the environment for training the prediction model.

16. The method of any one of claims 1 to 15, wherein the subject is human or animal, or robot.

17. A system for determining the presence of at least one condition and / or disease in a subject comprising: a trainable prediction model including one or more feature extractors and configured for analyzing the output of each of the one or more feature extractors, and outputting recognition of at least one condition and / or disease and the severity of the recognized at least one condition and / or disease; a first processor for processing at least one video associated with the subject into multimodal input for one or more feature extractors of a prediction model; a second processor associated with the generation of at least one label associated with the subject to train the prediction model; and a third processor for training the prediction model with the at least one label.Docket: 60257518. The system of claim 17, wherein the first processor includes a demultiplexer for creating multimodal input for the at least one feature extractor.

19. The system of any one of claims 17 or 18, wherein the at least one label includes a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject.

20. The system of any one of claims 17 to 19, wherein the second processor includes a computerized labeling agent for generating the at least one label as a plurality of direct and indirect labels.

21. The system of claim 20, wherein the direct labels include at least one of: 1) a clinical direct label indicating at least one condition and / or disease for the subject as determined by a clinician or machine, based on an examination of the subject, or 2) a direct label from self-reports or clinical records; and an indirect label from multimedia.

22. The system of any one of claims 17, 20 or 21 wherein, the third processor inputs the direct and / or indirect labels into the prediction model, and aggregates the direct and indirect labels to train the prediction model.

23. The system of any one of claims 17, or 20 to 22, wherein the direct labels from selfreports or clinical records, and the indirect labels from multimedia include one or more of recognition of one or more conditions and / or diseases, at least one category for each of the recognized conditions and / or diseases, or a score for each of the recognized conditions and / or diseases.

24. The system of claim 17, wherein the at least one condition and / or disease output by the prediction model includes a plurality of conditions and / or diseases.

25. The system of any one of claims 17 or 24, wherein the prediction model is configured to recognize a second condition and / or disease based on the recognition of the first condition or disease.

26. The system of claim 17, additionally comprising: a demultiplexer for demultiplexing a video stream to generate the multimodal input including video, voice, and or text.

27. The system of claim 17, wherein each of the one or more feature extractors extract features including vocal features, body features, facial / eye features and general data driven features.

28. The system of claim 23, wherein the multimedia includes social media.

29. The system of claim 18, wherein the at least one video associated with the subject is received from the subject or is obtained from multimedia.

30. The system of claim 29, wherein the multimedia includes social media.Docket: 60257531. The system of claim 17, additionally comprising a fourth processor configured for: obtaining content from the tat least one video obtained from the multimedia or additional different video from the multimedia, and producing certain of the obtained content as priors, inputting the priors into a context analyzer to produce context of the video and interaction of the subject with the environment for training the prediction model.

32. The system of any one of claims 17 to 31, wherein the subject is human, animal, or a robot.

Citation Information

Patent Citations

  • Systems and methods for mental health assessment

    US20210110894A1

  • Machine learning classification of video for determination of movement disorder symptoms

    US20240087743A1

  • Machine learning method for predicting a health outcome of a patient using video and audio analytics

    US20240120050A1

  • Multimodal (audio / text / video) screening and monitoring of mental health conditions

    WO2023235527A1