Systems and methods for monitoring early-stage symptoms of parkinson's disease

WO2026169868A1PCT designated stage Publication Date: 2026-08-13NEUROFORE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-08-13

Smart Images

  • Figure US2026014085_13082026_PF_FP_ABST
    Figure US2026014085_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods are provided for generating a prediction indicative of a neurodegenerative disease risk using non-motor symptom data collected over time. The system may obtain longitudinal symptom inputs for a user and determine feature-importance weights with an ensemble learning model. The weights are applied to form weighted feature vectors that emphasize clinically informative symptoms. A temporal sequence is encoded of the weighted vectors to capture dependencies and progression patterns. The encoded representations are adjusted with a delay-aware attention mechanism that conditions the encoded sequences on an inferred disease-onset reference. An inference module with multiple prediction heads produces an onset estimate, a diagnostic-delay estimate, and a risk value. A risk output is determined based at least in part on the encoded sequence and the inference outputs and may be communicated to a user device.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR MONITORING EARLY-STAGE SYMPTOMS OF PARKINSON’S DISEASECROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 755,432 filed February 7, 2025, entitled “ SYSTEMS AND METHODS FOR MONITORING EARLY-STAGE SYMPTOMS OF PARKINSON’S DISEASE,” which is incorporated herein by reference in its entirety.BACKGROUND

[0002] Parkinson’s Disease (“PD”) is a progressive neurodegenerative disorder, characterized by the progressive loss of dopaminergic neurons in the substantia nigra, a region of the brain essential for movement control. As these dopamine-producing cells diminish, movement regulation and coordination become increasingly difficult, and patients experience motor impairments alongside other pathological changes, such as neuronal loss, depigmentation, and gliosis in key brain regions.

[0003] PD staging commonly relies on the Hoehn and Yahr scale and the Unified Parkinson’s Disease Rating Scale. Early stages of the disease frequently appear as non-motor symptoms, including age-related factors, depression, anxiety, cognitive dysfunction, and autonomic disturbances, while later stages exhibit more recognizable motor symptoms like tremor, rigidity, bradykinesia, muscle stiffness, and postural instability. Clinical diagnostics, however, have largely focused on motor symptoms, which often means that non-motor signs indicating PD’s onset are overlooked. By the time the classic motor symptoms emerge, a significant loss (e.g., of up to about 70%) of dopaminergic neurons in the substantia nigra pars compacta may have already occurred, drastically affecting patient quality of life. Reliance on motor symptom presentation tends to result in delayed diagnosis and intervention, allowing neurological damage to progress. Recognizing PD in its initial stages through non-motor manifestations offers the potential for targeted intervention, which could slow disease progression, enhance outcomes, improve quality of life, and alleviate economic pressures on families and the healthcare system.

[0004] From a technological perspective, prevailing computational methods encounter further challenges. Verified PD diagnoses often trail the actual disease onset by years, preventing conventional supervised learning approaches from learning to recognize early non-motor symptom patterns as disease-relative, undermining accuracy in early detection. Approaches that can explicitly address temporal uncertainty and learn from pre-di agnostic nonmotor symptom data may offer more robust, timely insights into disease progression and risk.SUMMARY

[0005] The following embodiments and aspects thereof are described and illustrated in conjunction with systems, devices, and processes that are meant to be exemplary and illustrative, not limiting in scope. Aspects of the present disclosure generally relate to a system operable to detect and / or monitor neurodegenerative diseases such as Parkinson’ s Disease (PD) by prioritizing and tracking non-motor signs and symptoms that may relate to various stages of the neurodegenerative disease, thus enabling earlier detection and management.

[0006] While some studies have examined individual non-motor symptoms, none have integrated the full spectrum of non-motor signs and symptoms of PD into a comprehensive model for detecting and / or monitoring PD. One advantage of the systems and methods hereof the detection system may include the ability to predict the existence of PD based on a comprehensive analysis of non-motor signs and symptoms. Such comprehensive analysis may reveal manifestations of PD that cannot otherwise be detected by focusing on isolated symptoms and / or motor symptoms. In some instances, the systems and methods may implement strategically trained tools, models, techniques, and algorithms described herein. By leveraging a combination of non-motor signs and symptoms, the detection system of the present disclosure provides a scalable solution that can provide enhanced and earlier detection of PD, thus allowing for earlier intervention. Accordingly, in comparison to existing solutions, the detection system and methods described herein overcome the limitation of current diagnostic processes that, in contrast, rely heavily or primarily on the presence of motor symptoms.

[0007] According to various aspects, the detection system may include one or more user devices and an analysis platform or data processing center. Any of the one or more user devices may be an electronic device capable of executing algorithms. In various instances, the user devices and the analysis platform may be a phone, tablet, laptop, PC, and the like. The analysis platform may be decentralized between the user devices or may be remote from the user devices. In some instances, the analysis platform may include a centralized server and database in communication with the user devices. The user devices may communicate with the analysis platform through a wireless connection such as Wi-Fi, Bluetooth, cellular network, LongRange, Satellite, Infrared, and the like. Alternatively, the user devices may communicate with the analysis platform via a wired connection. The analysis platform may include memory, a communication interface, and a processor configured to execute one or more detection models. In various instances, the detection models may include one or more supervised learning models, unsupervised learning models, and / or reinforcing learning models. Additionally, the detection models may include one or more model-agnostic methods. In some instances, the Detection system may be configured to integrate with a cloud system.

[0008] According to one aspect, a system is provided for generating a prediction indicative of PD risk. The system may include a memory storing instructions and a processor configured to execute the instructions to obtain symptom data for a user across multiple times. The processor may be further configured to identify feature-importance weights for the symptom data and to apply those weights to generate weighted feature vectors that reflect the relative predictive value of various non-motor symptoms. The system may encode a temporal sequence of the weighted feature vectors; adjust the encoding using a delay-aware attention mechanism associated with an inferred disease-onset reference; generate an inference output that includes at least an onset estimate, a diagnostic-delay estimate, and a risk value; and determine a risk output based at least in part on the encoded sequence and the prediction-head outputs.

[0009] In some embodiments, the symptom data obtained and processed by the system may include non-motor symptoms, such as cognitive changes, autonomic irregularities, sleep-related disturbances, or other patient-reported indicators relevant to early Parkinson’s Disease detection. In other embodiments, the symptom data for a user may be aggregated over time, enabling the system to analyze longitudinal patterns, detect changes in severity or frequency, and incorporate temporal relationships into the prediction process. In yet other embodiments, the processor may identify or compute the feature-importance weights using at least one ensemble classifier configured to rank symptom features by their predictive contribution and apply those weights so that symptoms of greater diagnostic significance are given proportionally higher influence during analysis.

[0010] In some embodiments, the delay-aware attention mechanism may utilize disease-relative positional information derived from an inferred onset reference, allowing the system to prioritize one or more observations that are more informative for identifying early-stage symptom patterns. In other embodiments, the processor may generate the onset estimate, diagnostic-delay estimate, and risk value using multiple prediction heads trained jointly, enabling consistent inference of disease timing and risk within a shared predictivearchitecture. In yet other embodiments, the system may further provide an uncertainty value, such as a confidence score or calibrated interval, which reflects uncertainty associated with the predicted diagnostic output.

[0011] In some embodiments, the system may select or recommend one or more clinical actions from a constrained action set, such as recommending additional monitoring, earlier re-evaluation, or a referral to a specialist based on the determined risk output. In other embodiments, the constrained action set may be continuously updated based at least in part on outcomes associated with prior predictions, enabling the system to adapt its action-selection behavior over time. In yet other embodiments, the processor may store, retrieve, and process temporal sequences, encoded representations, calibration artifacts, and model checkpoints using a database that is designed to support longitudinal clinical records and ongoing system refinement.

[0012] According to another aspect, a system is provided for generating a prediction indicative of PD risk. The system may include a memory storing instructions and a processor configured to execute the instructions to receive non-motor symptom data for a subject over multiple times, identify and apply feature-importance weights to the symptom data to form weighted feature vectors, and encode a sequence of the weighted feature vectors using a temporal processing architecture. The processor may further be configured to identify, through multiple prediction heads operating on the encoded sequence, at least an onset estimate, a diagnostic-delay estimate, and a risk value, and to generate a risk output based at least in part on the encoded sequence and the outputs of the prediction heads. The processor may additionally be configured to provide a predicted diagnostic value based on the risk output to a user device.

[0013] In various embodiments, the system may be configured to implement a delay-aware temporal analysis framework that enables the detection models to interpret symptom progression in a disease-relative timeframe rather than a calendar-relative manner. In such embodiments, the system may construct a disease-relative temporal frame by transforming at least a portion of the timestamps associated with symptom observations into disease-relative timestamps based at least in part on an inferred onset estimate. These timestamps may then be used to compute corresponding disease-relative positional encodings, which are injected or incorporated into one or more self-attention layers of a transformer-based encoder. By conditioning the temporal encoding on the estimated onset rather than absolute time, the system may emphasize clinically meaningful intervals, compensate for diagnostic delay, and enhancethe identification of early-stage symptom patterns that precede confirmed diagnosis.

[0014] In some embodiments, the processor may identify feature-importance weights using an ensemble learning model configured to rank symptom features by their relative predictive contribution. The identified weights may then be applied to emphasize symptoms determined to be more indicative of early-stage PD.

[0015] In other embodiments, encoding the sequence of weighted feature vectors may include identifying one or more patterns of symptom progression, such as changes in the severity, frequency, or temporal relationships among non-motor symptoms over time. In yet other embodiments, the processor may detect a distribution shift in newly received symptom data and, upon detecting such a shift, may trigger recalibration, retraining, or adjustment of decision thresholds to maintain predictive reliability in changing clinical or population conditions.

[0016] In some embodiments, the processor may select or recommend a clinical action from a constrained action set, such as suggesting additional monitoring, requesting further evaluation, or prompting referral to a specialist based on the determined risk output. In other embodiments, the processor may continuously refine at least one of the feature-importance weights, temporal-encoder parameters, prediction-head parameters, or calibration artifacts using newly acquired symptom data and verified diagnoses, thereby enabling the system to maintain or improve predictive performance over time.

[0017] According to yet another aspect, a method is provided for generating a prediction indicative of a neurodegenerative disease. The method may include receiving multiple data entries for a user, each entry including symptom information and an associated timestamp and analyzing the symptom information to identify feature-importance weights that emphasize symptoms having greater predictive contribution. The method may further include applying the identified feature-importance weights to generate weighted feature vectors and encoding a sequence of the weighted vectors using a transformer encoder comprising multi-head self-attention. To support temporally informed inference under conditions of delayed clinical diagnosis, the method may include adjusting the encoded sequence with a delay-aware attention mechanism that constructs a disease-relative temporal frame by transforming at least a portion of the timestamps into disease-relative timestamps based at least in part on an inferred onset estimate. In various embodiments, these disease-relative timestamps may substantially guide attention operations within the transformer encoder and facilitate generation of inferenceoutputs including an onset estimate, a diagnostic-delay estimate, and a risk value. A final risk output may be determined based at least in part on the encoded sequence and the inference outputs, and a diagnostic prediction may be provided to a user device.

[0018] In some embodiments, the method may further include ranking symptom features of the non-motor symptom data and emphasizing those features having higher predictive contribution by applying corresponding weights when forming the weighted feature vectors. Such ranking may allow the method to prioritize clinically meaningful symptoms when generating temporally processed representations. In other embodiments, generating the risk output may include refining a risk evaluation by sampling a plurality of candidate onset estimates (e.g., from a posterior distribution), computing respective risk values for the sampled candidate onset estimates, and aggregating the resulting values to generate a temporally informed and delay-aware risk output.

[0019] In yet other embodiments, the method may further include continuously refining at least one of the feature-importance weights, encoder parameters, onset-estimation parameters, diagnostic-delay parameters, or risk-value parameters based on newly received non-motor symptom data and verified diagnoses. Through this continual refinement, the method may maintain or improve its predictive performance as additional user data becomes available over time.

[0020] According to various embodiments, the processor may continuously refine at least one of the feature-importance weights, calibration sets, or model parameters based at least in part on newly acquired data and verified diagnoses, thereby maintaining or improving predictive performance over time. In some embodiments, the system may use symptom data that includes non-motor symptoms reported by the user. In other embodiments, the symptom data is collected and aggregated over time. In such embodiments, the data collected over time may be used to detect changes or emerging trends and patterns in the user’s condition. In yet other embodiments, the processor determines feature-importance weights using at least one ensemble classifier, which ranks symptom features according to their predictive contribution and applies corresponding weights to emphasize those features that are more informative for assessing PD risk.

[0021] According to various aspects the detection system may be configured to execute a process for detecting PD that can enable a diagnosis of PD in an individual. The process may include the steps of receiving input data, executing data preprocessing techniques, executingdata processing techniques, predicting a likelihood of a Parkinson’s Disease diagnosis, providing the predictions indicative of the likelihood of the Parkinson’s Disease diagnosis, reinforcing data preprocessing and processing techniques, and training detection models. In various instances, input data may include, but is not limited to, non-motor signs and symptoms, medical records, medical history, and family medical history. In such instances, non-motor signs and symptoms may include relatively commonly assessed symptoms associated with PD (e.g., cognitive changes, sleep disturbance, and constipation) and relatively uncommonly assessed symptoms associated with PD (e.g., impaired taste and smell, hallucinations, sleep apnea, and dysphagia). In various instances, data preprocessing may include normalizing the input data. Further, data processing may include processing the data using the various detection models. Based at least in part on the processed data, the detection system may predict or determine a likelihood of a PD diagnosis. In various instances, the steps of reinforcing data preprocessing and processing techniques and training detection models may enhance the accuracy of the detection system during the prediction step. Generally, the steps of the diagnostic process of the detection system may be separated into a data entry phase, a processing phase, a prediction phase, a learning phase, and a training phase. In various instances, the different phases of the diagnostic process may be cyclical, meaning the training phase may lead into a future data entry phase to enhance the accuracy during the processing and prediction phases.

[0022] In certain embodiments, the detection system utilizes a temporal uncertainty learning framework designed to analyze longitudinal non-motor symptom data, particularly in cases where confirmed PD diagnoses occur after the actual onset of the disease. For example, the learning framework may enable the system to simultaneously estimate the latent onset time of the disease (T); the diagnostic delay (8); and the probability of PD classification (y). The system is capable of generating risk predictions at multiple pre-defined intervals, such as three, five, or seven years, prior to the confirmed diagnosis. In some instances, the temporal uncertainty learning framework incorporates an amortized inference network that processes a subject’s longitudinal data using a transformer-based encoder and employs distinct prediction heads to perform onset estimation, delay estimation, and PD classification.

[0023] In some embodiments, to substantially ensure the model substantially focuses on meaningful patterns in symptoms linked to Parkinson’s Disease, rather than simply learning the timing of diagnosis, the detection system may incorporate one or more safeguards. For example, the safeguards may include the use of medically informed structured priorities toguide how diagnostic delays are modelled, the introduction of additional tasks that require the model to reconstruct symptom progressions based on estimated disease onset, and the application of simulation-based validation where the true onset and delay timings are already established. This approach may maintain the integrity and reliability of the system’ s predictions by preventing it from defaulting to simplistic or trivial solutions.

[0024] In some embodiments, the system may implement delay-aware attention mechanisms that weight symptom observations according to their informativeness relative to an inferred onset reference, rather than calendar time, allowing the system to interpret symptom patterns in a disease-relative framework. Such mechanisms may include adjusting temporal positions based on the inferred onset, considering multiple possible onset times, and propagating related uncertainty into the predicted diagnostic values.

[0025] In other embodiments, the system may provide an uncertainty quantification. For example, the uncertainty quantification may be provided as a horizon-specific uncertainty quantification to support clinical decision-making under temporal uncertainty. For example, the system may generate calibrated prediction ranges or calibrated risk estimates for different forecast horizons, enabling the system to distinguish uncertainty associated with shorter-term predictions from that of longer-term predictions. In yet other embodiments, the system may output not only a likelihood of PD but also a recommended clinical action selected from a constrained action set. Such recommended actions may include, for example, additional monitoring, earlier or confirmatory evaluation, or scheduling follow-up at a defined interval based at least in part on the predicted risk and corresponding uncertainty.

[0026] In some embodiments, the system may employ a hybrid machine-learning architecture that integrates ensemble methods with deep-learning models. In such embodiments, ensemble methods may be used to determine feature-importance values that quantify the predictive contribution of various non-motor symptoms, and the resulting weighted feature vectors may be processed by a temporal neural network configured to capture complex temporal relationships and account for uncertainty related to diagnostic delay.

[0027] In other embodiments, the system may implement a validation framework designed to evaluate early-detection performance in contexts where ground-truth labels are delayed. Such a framework may include temporally aware data partitioning, evaluation across multiple prediction horizons, computation of lead-time or utility metrics, assessment of data-distribution shifts, stress-testing under missing or irregularly sampled data, and estimation of confidenceintervals to assess robustness.

[0028] In yet another aspect, the system may be configured to execute a training process for improving the accuracy and reliability of its diagnostic predictions. The training process may include receiving historical symptom data and verified diagnoses, training one or more detection models, and scaling those models over time. Such training may enable the system to learn symptom patterns associated with positive diagnoses and to differentiate Parkinson’s Disease from other conditions with overlapping symptoms. The training process may further employ a combined learning objective that includes, for example, a component for disease-classification accuracy, a component encouraging medically plausible diagnostic-delay estimates, a component for reconstructing symptom trajectories relative to inferred onset timing, and a component discouraging decreases in predicted disease risk as symptom severity increases. The training may be optimized using gradient-based techniques that can include regularization strategies, gradient management, adaptive learning-rate scheduling, and early-stopping criteria to maintain stable and efficient learning behavior.

[0029] These and other aspects and advantages of the present disclosure will become apparent to those skilled in the art after considering the following detailed description in connection with the accompanying drawings. It will be understood by those skilled in the art that one or more aspects of this disclosure can meet certain objectives, while one or more other aspects can lead to certain other objectives. Various modifications to the illustrated embodiments will be readily apparent to those skilled in the art, and the generic principles herein can be applied to other embodiments and applications without departing from embodiments of the disclosure. Other objects, features, benefits, and advantages of the present disclosure will be apparent in this summary and descriptions of the disclosed embodiments, and will be readily apparent to those skilled in the art. Such objects, features, benefits, and advantages will be apparent from the above as taken in conjunction with the accompanying figures and all reasonable inferences to be drawn therefrom.BRIEF DESCRIPTION OF THE DRAWINGS

[0030] FIG. 1 is schematic view of an example detection system according to the teachings of the present disclosure;

[0031] FIG. 2 is another schematic view of the example detection system of FIG. 1;

[0032] FIG. 3 is a flow chart of a method for detecting a neurodegenerative disorderaccording to the teachings of the present disclosure; and

[0033] FIG. 4 is a flow chart of a training process according to the teachings of the present disclosure.

[0034] While the disclosure is susceptible to various modifications and alternative forms, a specific embodiment thereof is shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the drawings and detailed description presented herein are not intended to limit the disclosure to the particular embodiment disclosed, but to the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claimsDETAILED DESCRIPTION

[0035] Before any embodiments are described in detail, it is to be understood that the disclosure is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings, which is limited only by the claims that follow the present disclosure. The disclosure is capable of other embodiments, and of being practiced, or of being carried out, in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings.

[0036] The following description is presented to enable a person skilled in the art to make and use embodiments of the disclosure. Various modifications to the illustrated embodiments will be readily apparent to those skilled in the art, and the generic principles herein can be applied to other embodiments and applications without departing from embodiments of the disclosure. Thus, embodiments of the disclosure are not intended to be limited to embodiments shown but are to be accorded the widest scope consistent with the principles and features disclosed herein. The following detailed description is to be read with reference to the figures, in which like elements in different figures have like reference numerals. Skilled artisans will recognize the examples provided herein have many useful alternatives and fall within the scopeof embodiments of the disclosure.

[0037] The present disclosure generally relates to systems and platforms designed to identify, monitor, and predict the onset or progression of neurodegenerative diseases by analyzing patterns in patient-reported information collected over time. The disclosed system may integrate data from multiple devices and sources, enabling the aggregation, processing, and interpretation of both real-time and historical symptom information. By leveraging modem data-acquisition and machine learning techniques, the system can evaluate temporal changes in non-motor symptoms, generate diagnostic insights, and provide risk-related outputs to users or clinicians. The system can also support remote access and interaction across different user devices, allowing patients and healthcare professionals to submit, review, and share health-related information. As such ,the system provides a flexible foundation for diagnostic support, enabling earlier detection, improved monitoring, and more informed clinical decision-making.

[0038] In some embodiments, the detection system described herein may be employed to predict the onset of Parkinson’s Disease (“PD”). In such embodiments, the system utilizes an analysis platform that acquires patient data, retrieves previous records from various storage solutions, and applies processing techniques, including supervised, unsupervised, and reinforcement learning models, to interpret the information and determine a diagnosis. The platform continuously refines its predictive algorithms by incorporating both raw patient inputs and verified diagnostic outcomes, ultimately delivering nuanced PD diagnosis predictions directly to user devices.

[0039] Additionally, while the following discussion may describe features associated with specific devices or embodiments, it is understood that additional devices and / or features can be used with the described systems and methods, and that the discussed devices and features are used to provide examples of possible embodiments, without being limited.

[0040] FIG. 1 is a schematic diagram of a detection system 100 designed to provide a PD diagnosis in accordance with a disclosed example. The detection system 100 may include at least one or more machine learning techniques, algorithms, models, or frameworks and one or more user devices (e.g., a smart phone, a computer, or another electronic device capable of performing tasks and processing and storing data). For example, the detection system 100 may include a first user device 105 that may be, for example, associated with an individual or patient; a second user device 110 that may be, for example, associated with a clinician,physician, or other medical personnel; a communication network 115; a database 117; and a analysis platform or analysis platform 119 having at least one processor configured to provide a PD diagnosis prediction.

[0041] The first and second user devices 105, 110 may each comprise any suitable personal computing equipment, including but not limited to desktop computers, laptop computers, terminals, smart televisions, electronic whiteboards, tablet computers, smart telephones, wearable devices, or similar electronic devices. While the illustrated embodiment depicts two user devices, it is contemplated that any number of user devices may be incorporated within the diagnostic system 100, including configurations with more than two, tens, or even hundreds of such devices.

[0042] The first and second user devices 105, 110 may be of identical type or may differ in type depending on the particular implementation. For instance, one embodiment may utilize a first user device 105, such as a tablet located at a clinical facility, wherein multiple patients possess individual login credentials to access their respective profiles within the detection system 100. Concurrently, the second user device 110 may be a personal computer associated with the same clinic, allowing medical professionals, each with their unique login credentials, to access various patient profiles and data within the detection system 100. In other embodiments, the first user device 105 may be a personal cellular device that is subsequently or retroactively connected to the detection system 100, whereby the patient utilizes a downloadable application to access their respective profile on the detection system 100. In this scenario, the second user device 110 may be a personal computer located at a clinic, enabling one or more medical professionals, via their individual login credentials, to monitor patient uploaded data and review the corresponding PD diagnosis prediction. Among the advantages provided by the first and second user devices 105, 110, whether implemented as smart devices, smartphones, laptops, personal computers, tablets, or other suitable electronic devices, is the inherent portability. This portability enables individuals to access healthcare services, records, and information from virtually any location, thereby reducing or eliminating delays associated with traveling to physical clinics or hospitals.

[0043] The communication network 115 may be a local area network (LAN), a wide area network (WAN), or other types of networks, supporting protocols like transmission control protocol / Internet protocol (TCP / IP) over Ethernet. In some instances, the communication network 115 may include Wi-Fi, Bluetooth, cellular network, Long Range, Satellite, Infrared, and the like. In some instances, the user devices 105, 110 may communicate with the analysisplatform 119 and the database 117 via a direct or indirect wired connection. In some instances, the detection system 100 may also support or include cloud integration for larger data storage, faster data sharing, and enhanced remote diagnosis.

[0044] The database 117 may be configured to store information input patient data (e.g., demographics, medical history, clinical visits), symptom data (e.g., non-motor symptoms, motor symptoms, symptom severity, and symptom onset timeline), longitudinal data (e.g., time-stamp symptom assessment and progression data), diagnostic outcomes (e.g., confirmed diagnoses and diagnostic uncertainty), raw input data (e.g., patient inputted responses, sensor or device data, lab results), model data, training data, external data (e.g., reference data sets and guidelines) and the like. The database 117 may be accessed any of the first and second user devices 105 and the analysis platform 119 via the communication network 115.

[0045] In some instances, the database 117 may be provided in the form of one or more data stores, memory units, processors, elastic cache systems, cloud storage systems, logic tables, relational databases (e.g., Structured Query Language (SQL) databases), NoSQL databases, time-series databases, graph databases, data warehouses, data lakes, embedded databases, vector databases, hybrid databases, distributed databases, or similar data storage mechanism, or a combination thereof.

[0046] Different types of data may be stored in a single database 117 or may be stored in separate databases, or any combination. Each of the one or more databases 117 may be any type of centralized database, distributed database, or the like. For example, a centralized database may reside in one central location which may be accessed by various users via the communication network 115. In another aspect, a distributed database may be spread across multiple servers or nodes in at least one network, with each server managing part of the database 117. Within a distributed database, a user may retrieve data seamlessly as if it were stored within a centralized database.

[0047] The analysis platform 119 or data processing center is designed to process, interpret, and manage the flow of patient and system data within the detection system 100. In various embodiments, the analysis platform 119 may comprise one or more electronic processes configured to acquire, aggregate, and analyze data from the connected user devices 105, 110, as well as from databases 117 and external sources. This analysis platform 119 is configured to execute a range of machine learning models (including supervised, unsupervised, and hybrid approaches) to evaluate patient information, historical records, and other relevantdatasets. In some embodiments, the analysis platform 19 may be configured to execute one or more machine learning models (e.g., ensemble models, transformer neural networks, prediction heads models, SHAP, K-means clustering, reinforcement leaning, and the like). In such embodiments, the analysis platform 119 may be configured as a centralized data processing center. In other embodiments, the analysis platform 119 may be decentralized.

[0048] In some embodiments, the analysis platform 119 may orchestrate the preprocessing of raw, non-normalized data, normalizing and encoding inputs to substantially ensure compatibility with one or more machine learning algorithms. The analysis platform 119 may be configured to execute predictive models for a PD diagnosis by ingesting and receiving user-submitted information, comparing the data to existing medical records, and applying data processing techniques to provide nuanced PD prediction outcomes. Additionally, the analysis platform 119 may be configured to iteratively refine its predictive accuracy by incorporating verified diagnostic results, leveraging reinforcement learning as needed. In some embodiments, the analysis platform 119 may also facilitate adaptive questioning, wherein data acquisition may be modified based on prior responses and integrating new information in real time to enhance future diagnostic predictions and system performance. Generally, the analysis platform 119 is designed to provide a robust, scalable, and secure environment for executing sophisticated data processing and predictive analytics.

[0049] Referring to FIG. 2, another schematic of the detection system 110 is provided. According to various aspects, the analysis platform 119 may be designed to ingest user input data, including responses from patients and clinicians, in order to determine a PD diagnosis prediction. In some embodiments, the analysis platform 119 may include its own memory 120, which is configured to temporarily or permanently store incoming data, intermediate processing results, or learned parameters from ongoing model training and a communication interface 125, for example, a transceiver, that allows the analysis platform 119 to communicate with external devices, for example one or more servers over a communication network as noted above. Further, the analysis platform 119 may employ supervised learning models 130 and unsupervised learning models 135 to facilitate data acquisition and diagnostic prediction. In certain embodiments, the analysis platform 119 may utilize a hybrid approach, integrating both feature selection techniques and deep learning models, to enhance the accuracy and robustness of PD diagnosis predictions by leveraging the strengths of multiple analytical strategies.

[0050] In various instances, the analysis platform 119 may be configured to perform various functions including, but not limited to: acquire data from one or more individuals viathe one or more user devices 105, 110; retrieve previous data entries from memory 120 or, if integrated, a cloud system (not illustrated); perform advanced data preprocessing and processing techniques using various detection models (e.g., machine learning approaches including but not limited to supervised learning models 130 and unsupervised learning models 135); comprehensively analyze the acquired data, saved data, and / or any other medical history or record; provide a PD diagnosis prediction (e.g., a prediction of a likelihood of PD or a likelihood of a PD diagnosis); output said PD diagnosis prediction to the user devices 105, 110 via a communication interface 135; and receive additional information such as verified PD diagnoses for training one or more reinforcing learning models 140 to improve the accuracy of future PD diagnosis predictions.

[0051] The memory 120 of the analysis platform 119 may be configured to store incoming data, including patient and clinician inputs, intermediate processing results, and learned parameters temporarily or permanently from ongoing model training. Accordingly, the memory 120 may permit the analysis platform 119 to efficiently manage and retrieve historical records, compare new patient information against previously stored data, and support iterative refinement of predictive models. Additionally, the memory 120 may facilitate seamless data access for both real-time and batch processing, ensuring that diagnostic predictions are based on comprehensive and up-to-date information. In certain embodiments, the memory 120 may also support integration with cloud-based systems, allowing for scalable storage and enhanced data sharing across devices and users.

[0052] The communication interface 125 of the analysis platform 119 may be designed to facilitate data acquisition, exchange, and system interaction within other components of the detection system 100 (e.g., the first and second user devices 105, 110 and the database 117). In some embodiments, the communication interface 125 may comprise hardware and software elements such as transceivers, network adapters, and protocol stacks that enable connectivity between the analysis platform 119 and external devices, including user devices 105, 110, servers, and cloud-based resources. This communication interface 125 may initiate and manage the display of initial prompts on connected user devices (e.g., the first and second user devices 105, 110), thereby facilitating the acquisition of patient and clinician data. In some embodiments, the communication interface 125 may be configured to receive diverse forms of raw, non-normalized input, including textual, numerical, and categorical data, directly from end-user devices, ensuring that the analysis platform 119 can ingest both spontaneous patient responses and structured medical records.

[0053] In other embodiments, The communication interface 125 further supports the receipt of data pertaining to previously verified PD diagnoses, encompassing both confirmed positive and negative cases. For example, the communication interface 125 may be configured to receive a verified positive PD diagnosis, meaning the individual is diagnosed with PD, and a verified negative PD diagnosis, meaning the individual is not diagnosed with PD at the time. Such verified outcomes are critical for supervised and reinforcement learning processes, as they enable the analysis platform 119 to refine prediction models and enhance diagnostic accuracy over time. Beyond direct patient or clinician inputs, the communication interface 125 may be capable of interacting with external programs, systems, and databases through standardized APIs and secure data exchange protocols. In such instances, the communication interface 125 allows the analysis platform 119 to request and integrate supplementary data, such as family history, prescription records, and longitudinal health information, thereby providing a comprehensive foundation for diagnostic analysis. In some embodiments, the communication interface 125 may be configured to request and receive data from the first and second user devices 105, 110.

[0054] Central to user interaction, the communication interface 125 may include a graphical user interface (“GUI”) on the display screens of user devices 105, 110. In some embodiments, the GUI may be implemented using frameworks such as Tkinter or other crossplatform libraries. The GUI may provide clinicians and patients with a virtual environment for entering symptoms, medical history, and additional health-related data. The GUI may comprise logic-based questionnaires featuring multiple question types such as drop-down menus, multiple choice, fill-in-the-blank formats, and scales for numerical input (e.g., “1 to 10,” “0 to 5”) to capture detailed and nuanced information about both common and uncommon nonmotor symptoms associated with PD. The interface also accommodates the entry of vital sign metrics and allows for optional free-text notes, enabling users to document observations not covered by standardized questions.

[0055] In some embodiments, the communication interface 125 may be designed to process and encode incoming data for compatibility with machine learning pipelines. In such embodiments, the communication interface 125 may perform one or more preprocessing steps to prepare the data for further analysis. For example, the communication interface 125 may normalize the data formats, encode binary responses (e.g., converting “yes” / “no” to 1 / 0), and transform categorical labels into integer values to facilitate subsequent analysis. The communication interface 125 may also support real-time adaptive questioning, dynamicallyadjusting the content and structure of future questionnaires based on previous responses and evolving diagnostic criteria. Additionally, the communication interface 125 enables secure transmission and temporary or permanent storage of acquired data in the platform’s memory 120 or database 117, supporting both immediate analysis and long-term model refinement.

[0056] In various embodiments, the analysis platform 119 may comprise a hybrid machine learning module that includes an integrated system combining multiple modeling approaches. For example, the hybrid machine learning module may combine models having feature selection techniques and deep learning algorithms to enhance diagnostic accuracy and robustness. This hybrid module may enable the analysis platform 119 to leverage both supervised and unsupervised learning models 130, 135, facilitating comprehensive data analysis and nuanced prediction capabilities. By integrating ensemble methods and advanced neural networks, the detection system 100 may substantially effectively identify relevant diagnostic indicators, capture complex relationships within patient data, and support the timely and precise assessment of PD risk and progression. Such a hybrid approach allows the detection system 100 to provide clinicians and patients with improved predictive outcomes.

[0057] After and / or in conjunction with data preprocessing, the detection system 100 may be configured to perform data processing. Integrating robust preprocessing into the analysis platform 119 offers several technical advantages for PD diagnosis prediction. Preprocessing may substantially ensure that raw, non-normalized data is systematically transformed and encoded, which improves compatibility with downstream machine learning models and enhances overall prediction reliability. By normalizing inputs, handling missing values, and converting categorical or textual data into numerical formats, the system reduces the risk of data inconsistencies and model errors. This workflow enables the analysis platform to produce more accurate feature representations, which are critical for both ensemble methods and neural network architectures. In some embodiments, the detection system 100 may require incoming data to be fully preprocessed before initiating predictive modeling. In such embodiments, subsequent stages such as feature importance weighting, temporal sequence encoding, and classification must wait until the preprocessing pipeline completes. In other embodiments, waiting for preprocessing ensures that all inputs are standardized, encoded, and validated may maintain the integrity of the analysis and preventing downstream propagation of errors. Moreover, this sequencing may enable real-time adaptive questioning, as the platform can dynamically adjust prompts or request additional clarifying data if preprocessing identifies anomalies or gaps, further optimizing diagnostic outcomes and system efficiency.

[0058] In other embodiments, data processing may occur for data sets not requiring data preprocessing. In such embodiments, the hybrid modeling approach combines feature selection and deep learning. In various instances, the deep learning aspect of the analysis platform 119 may include at least two machine learning models. The hybrid modeling approach of the detection system 100 may include a first model relating to handling large data sets for producing robust predictions based on highlighted and prioritized values, and a second model related to capturing complex, non-linear relationships within data, allowing for nuanced predictions based on the interplay of multiple data sets.

[0059] Generally, the analysis platform 119 may incorporate one or more analytical stages designed to support diagnostic functions. For example, the analysis platform 119 may utilize machine learning tools to identify relevant patterns in patient data and enhance the accuracy of its diagnostic predictions. Each stage may be designed to process, analyze, and interpret the input and / or stored health information in order to better support clinicians and patients in making informed decisions about PD risk and diagnosis.

[0060] In various embodiments, the detection system 100 (e.g., the analysis platform 119) may implement a multi-stage machine-learning, interference pipeline configured to identify, prioritize, and analyze non-motor symptom patterns associated with early-stages of PD. The pipeline may comprise at least two primary stages: a first stage related to an ensemble-based feature-importance and weighting stage, and a second stage related to a temporal neural -network stage designed to model symptom progression and address diagnostic delay uncertainties. In some embodiments, the first stage may be substantially directed to ensemblebased feature weighting. In such embodiments, the first stage (i.e., Stage 1) may receive raw patient symptom data; process the data to compute feature importance weights; and provide weighted feature vectors that identify, underscore, and emphasize relatively meaningful diagnostic symptoms. As such, the first sage may be substantially directed to reducing noise and prioritizing clinically relevant symptoms before further neural network processing. Further, a second stage (i.e., Stage 2) may be directed to forming a temporal neural network. During Stage 2, the detection system 100 may receive the weighted feature vectors from Stage 1. The weighted feature vectors may be used to processed to determine at least a prediction onset time, a diagnostic delay, and / or a PD probability. Accordingly, Stage 2 may be configured to perform temporal modeling of symptom progression, while accounting for diagnostic delay, and providing a final PD risk prediction with calibrated uncertainty intervals.

[0061] Together, these stages operate cooperatively to generate relatively accurate,temporally contextualized PD diagnosis predictions. These two stages may substantially constitute the system's inference architecture, which processes input data to generate predictions. The system additionally implements training, calibration, and reinforcement learning operations that optimize and refine the inference pipeline over time, as described in subsequent sections. In some embodiments, the output of both stages may include a PD diagnosis prediction, a calibrated risk estimate with uncertainty intervals, estimated disease onset time and diagnostic delay, feature importance scores, and a recommended clinical action.

[0062] As part of the first stage, the detection system 100 (e.g., the analysis platform 119) may be configured to identify, quantify, and prioritize non-motor symptom features that exhibit diagnostic relevance to early-stage Parkinson’s Disease. The first stage may employ one or more supervised ensemble learning models, such as a Random Forest classifier, an XGBoost (e.g., extreme Gradient Boosting) classifier, or related gradient-boosted tree architectures, to compute and identify a set of feature-importance values that reflect each symptom’s predictive contribution. In some embodiments, these ensemble models may be trained on historical, annotated patient datasets that include longitudinal non-motor symptom measurements and verified diagnostic outcomes.

[0063] During training, the ensemble models may construct a plurality of decision trees (e.g., between approximately 100 and 500 independent trees in the case of a Random Forest, or sequentially boosted trees in the case of XGBoost). Each tree may be trained on a bootstrapped or subsampled subset of the training data, thereby capturing heterogeneous symptom relationships across diverse patient profiles. For each decision tree, branch-splitting operations may identify symptom dimensions that at least partially reduce impurity or maximize predictive gain when distinguishing PD versus non-PD cases. In some embodiments, the ensemble aggregates these contributions across the full set of trees to produce numerical feature-importance scores for each symptom, where higher scores correspond to features demonstrating greater discriminatory power in early PD detection.

[0064] In some embodiments, Random Forest implementations may include computing importance scores using mean-decrease-in-impurity (i.e., Gini importance), permutation importance, or related impurity-reduction metrics. In alternative embodiments, XGBoost may compute feature importance using gain-based measures, cover-based metrics, frequency -based metrics, or combinations thereof. XGBoost may further employ regularization techniques, including LI (Lasso) and L2 (Ridge) penalties to reduce model overfitting, maximum tree-depth constraints to limit structural complexity, and column- and row-subsamplingprocedures, to enhance generalization performance, particularly in heterogeneous clinical datasets.

[0065] Upon completion of the importance-computation step, the ensemble model may generate a feature-importance vector:W = [Wi,W2, ...,Wn]

[0066] where each element (wi) corresponds to the learned predictive value of symptom “i.” These weights are then applied to patient-specific symptom observations:X = [x1,x2, ...,xn]

[0067] to produce a weighted feature representation:x_weighted = [wi • xltw2• x2, wn• xn]

[0068] This weighted representation increases the relative emphasis of clinically meaningful symptoms, such as those empirically correlated with prodromal PD, while diminishing the influence of less informative or noisy features. By amplifying high-value diagnostic signals prior to downstream neural -network processing, this stage provides a technically advantageous, noise-reduced input representation that improves prediction stability and enhances the system’s ability to detect subtle early-stage symptom patterns.

[0069] In some embodiments, the ensemble model may further compute time-dependent or context-dependent importance values, enabling dynamic weighting of symptoms relative to a patient’s inferred disease-onset timeline T. In such embodiments, separate importance scores for symptom magnitudes, rates of change, or temporal patterns may be calculated thereby allowing the system to prioritize symptoms not only by absolute severity but also by their temporal proximity to estimated onset. These weighted symptom representations may be provided as standardized, diagnostically optimized inputs for subsequent processing stages, including the temporal neural -network architecture described herein.

[0070] As part of the second stage, the detection system 100 (e.g., the analysis platform 119) may incorporate a second supervisor learning model 125 (e.g., Neural Network) to analyze weighted symptom data over time and generate PD diagnosis predictions. In some instances, the most relevant non-motor symptoms, such as sleep disturbances, cognitive impairment, autonomic dysfunction indicators, or other prioritized features, identified during the first stage may be provided as input to the second supervisor learning model 125 (e.g.,Neural Network).

[0071] The second supervisor learning model 125 may comprise multiple specialized sub-networks organized into a temporal uncertainty learning architecture. In some embodiments, the architecture of the second supervisor learning model 125 may be designed to model longitudinal symptom progression patterns in the presence of systematically delayed clinical diagnoses. Because verified PD diagnoses typically occur years after symptom onset, the temporal neural network may incorporate mechanisms to infer disease-relative timescales and to estimate latent parameters such as onset time and diagnostic delay.

[0072] In some embodiments, the temporal network may include a transformer-based encoder composed of multiple layers of multi-head self-attention mechanisms. The encoder may process sequences of weighted symptom vectors corresponding to distinct observation times and compute relevance scores between symptoms observed at different times. Through this self-attention mechanism, the model may identify long-range temporal dependencies, non-linear progression patterns, early prodromal signals, and other clinically meaningful symptom trajectories.

[0073] To account for diagnostic delay, the temporal neural network may incorporate an amortized inference network configured to jointly estimate a latent disease onset time T; a diagnostic delay 8; and a PD classification likelihood y. Each of these quantities may be predicted using dedicated multi-layer perceptron (“MLP”) heads, thereby enabling multi-task learning that integrates onset prediction, delay estimation, and disease-likelihood modeling in a unified framework.

[0074] In certain embodiments, the model may further include a delay-aware attention mechanism that transforms absolute observation timestamps into disease-relative timestamps (t' = t - T) and injects these relative timestamps into the attention layers. By conditioning temporal encoding on inferred onset time rather than absolute calendar time, the system aligns symptom trajectories across patients and emphasizes observations most indicative of early PD progression. In some embodiments, the system may propagate onset uncertainty by sampling multiple candidate onset times from a posterior distribution and marginalizing classification predictions across the sampled values. This probabilistic approach yields a more robust PD risk estimate that accounts for uncertainty in symptom timing.

[0075] Further, the temporal neural network may incorporate auxiliary reconstruction tasks that require the model to reconstruct historical symptom vectors conditioned on the inferredonset time. These tasks help ensure that the onset estimate plays a meaningful role in explaining symptom evolution, preventing collapse to trivial solutions, and improving physiological plausibility.

[0076] The integration of the first and second stages into a cohesive machine-learning pipeline provides multiple technical benefits. Whereas the first stage may reduce noise and dimensionality by emphasizing clinically significant features, the second stage may model complex temporal patterns and compensates for delayed diagnostic labels. Together, these stages (i.e., the first and second stages) may enable early detection of PD-related patterns, improved robustness to missing or irregular data, enhanced interpretability through feature importance attribution, and more reliable prediction outcomes in real-world clinical environments.

[0077] In some embodiments, the second stage of the machine-learning pipeline, may employ a transformer encoder neural network configured to process sequences of weighted symptom feature vectors over time. In this architecture, patient data may be represented as a temporal sequence S = {si, S2, ..., st], where each element Si corresponds to a weighted feature vector representing observed non-motor symptoms at a particular time ti. These weighted vectors may be generated during the first stage through the application of ensemble-derived feature-importance values, thereby ensuring that the inputs supplied to the transformer emphasize diagnostically meaningful symptom attributes.

[0078] In some embodiments, the transformer encoder may comprise L layers (typically L = 6) of multi-head self-attention mechanisms, with H attention heads per layer (typically H = 8). Each layer may include a multi-head self-attention sublayer configured to compute attention weights between all pairs of time points in the sequence, enabling the identification of short-range and long-range temporal dependencies in symptom progression; a position-wise feed-forward sublayer incorporating non-linear activation functions to transform intermediate representations; and residual connections and layer-normalization operations to stabilize training, improve gradient flow, and support deeper temporal modeling.

[0079] During operation, the self-attention mechanism may compute pairwise relevance scores among symptom observations across different times, allowing the model to capture clinically meaningful patterns such as gradual intensification of prodromal symptoms, fluctuating symptom clusters, and non-linear symptom trajectories that may precede a clinically verified PD diagnosis by several years. The position-wise feed-forward networksmay refine these contextualized representations, and the residual connections may preserve information continuity while mitigating vanishing-gradient effects.

[0080] The transformer encoder may output a sequence of encoded temporal representations E = {ei, e?, et}, where each encoded vector ei lies in a d-dimensional representation space (e.g., d = 256). Each encoded representation captures not only the symptom state at the corresponding time point ti but also contextual dependencies derived from the full temporal history of the patient. These encoded representations may then be provided to subsequent components of the temporal uncertainty learning architecture, including the onset-estimation network, delay-estimation network, and disease-classification head described elsewhere herein.

[0081] This transformer-based temporal encoding strategy provides several technical advantages. First, multi-head self-attention enables the model to weigh symptom observations differentially across time, allowing it to focus on critical intervals associated with early-stage PD development. Second, the architecture supports flexible modeling of irregularly sampled or incomplete symptom histories because the attention mechanism does not require uniform time intervals. Third, by producing rich contextual embeddings that capture both temporal dynamics and symptom-severity trends, the transformer encoder provides an informative intermediate representation that enhances the accuracy and robustness of downstream PD risk predictions.

[0082] In some embodiments, as part of the second stage, the detection system 100 may implement a delay aware attention mechanism that conditions temporal encoding on an inferred disease onset time rather than absolute calendar time. By aligning symptom trajectories to an estimated onset, the system weights observations by their informativeness relative to disease progression stage, thereby mitigating distortions introduced by systematically delayed clinical diagnoses. In such embodiments, the analysis platform 119 may transform chronological timestamps into disease-relative time before applying attention. For each observation timestamp t a relative time may be computed as:

[0083] Where T denotes the inferred latent onset time. Positional encodings are then generated from t- using sinusoidal functions:, ( t- ( t- PE(tf, 2 / ) = sin I 7777777777 7777 I, PE(t 2 / + 1) = COS 7777777777777 \100002' / d / 7 7\100002' / d /

[0084] and injected into the attention layers of the temporal encoder so that attention weights are computed with respect to disease relative positions rather than absolute time. This conditioning allows the network to emphasize intervals most indicative of prodromal and early progression phase. Given a sequence of weighted symptom feature vectors S = {s1(s2, ... , st] produced in the first stage and a transformer-based temporal encoder in Stage 2, the model applies multi-head self-attention using the disease-relative positional encodings above. The encoder thereby computes relevance scores among all pairs of time points ( k) in the diseaserelative frame, producing an encoded sequence E = {e1(e2, .... et] in which each representation captures both local symptom state and long-range dependencies aligned to the inferred onset.

[0085] Further, because the onset time T is inferred and may be uncertain, some embodiments may marginalize downstream prediction over a posterior distribution p(i I x). For example, the system may sample the K candidate onset times— ^k} fromadistribution centered at the point estimate with variance reflecting uncertainty; compute model predictions yk(e.g., PD risk) for each candidate; and aggregate predictions by a probability-weighted average of:yfinal

[0086] where wkis normalized posterior probabilities associated with each candidate Tfc. This posterior marginalization propagates onset uncertainty into the final PD risk estimate and improves robustness under irregular sampling and delayed labeling. The delay aware attention operates within the temporal neural network architecture in the second stage and may interface with at least one of the transformer encoder, which receives disease-relative positional encodings and produces context-rich temporal embeddings; and the amortized inference network with three prediction heads (onset, delay, classification), which consumes the encoder’s embeddings after any pooling step (e.g., mean or attention-weighted pooling) to output T, <5, and y.

[0087] In some embodiments, the second stage may include employing an amortized inference network configured to interpret the transformer-encoded temporal representations and produce joint predictions related to disease onset, diagnostic delay, and Parkinson’s Disease classification likelihood. In such embodiments, the amortized interference network may operate downstream relative to the transformer encoder and leverages the contextualizedtemporal embeddings to infer clinically meaningful latent variables that cannot be directly observed from patient data due to systematically delayed diagnostic labeling.

[0088] In some embodiments, the encoded temporal sequence E = {ei, e?, .. et] generated by the transformer encoder may be aggregated into a single fixed-dimensional representation e agg G Rd. In such embodiments, the aggregation may be performed using one or more pooling techniques, including but not limited to mean pooling, attention-weighted pooling, or a combination thereof. The aggregated representation may capture both short- and long-range temporal dependencies, producing a holistic summary of the patient’s symptom progression aligned to clinically relevant temporal structure.

[0089] The aggregated vector e agg may then be processed by an amortized inference network comprising the three parallel multi-layer perceptron (“MLP”) prediction heads, each dedicated to a separate inference task. For example, one task may be related to an onset time estimator. This MLP may follow an architecture of [d — d / 2 — d / 4 — 1] and may be configured to estimate a continuous-valued latent disease onset time T. In some embodiments, the onset time estimator may employ a linear activation function at the output layer to allow unbounded real-valued onset estimates. This prediction head enables the system to infer the approximate time at which PD-related physiological changes began, independent of when clinical diagnosis occurred.

[0090] In some embodiments, a second task may be a delay estimator. The delay estimator may be configured to employ an architecture (e.g., [d — d / 2 — d / 4 — 1]) but may utilize a soft plus activation function at its output defined as 6 = log(l + exp(z)), ensuring that the estimated diagnostic delay 5 is positive. The delay estimator may be designed to predict the duration between the inferred disease onset time and the verified clinical diagnosis, thus allowing the model to adjust its temporal interpretation of symptom patterns to account for systematic diagnostic lag.

[0091] In some embodiments, a third task may be a disease classifier. The disease classifier may be configured to employ an architecture (e.g., [d — d / 2 — d / 4 — 1]) and may predict a PD classification probability y G [0, 1] using a sigmoid activation function. The classifier may output the likelihood that the patient exhibits prodromal or early-stage PD at the time of evaluation.

[0092] Accordingly, these three prediction heads (i.e., the onset time estimator, delay estimator, and disease classifier) may be trained jointly using a multi-task learning objective,which may integrate losses corresponding to onset prediction, delay prediction, and disease probability estimation. Multi-task training provides several technical advantages including improved parameter sharing, enhanced generalization, and the ability to enforce consistency between related temporal inferences. By learning interconnected tasks simultaneously, the amortized inference network ensures that predictions of onset time, diagnostic delay, and disease likelihood mutually reinforce one another and reflect physiologically plausible relationships.

[0093] The combination of the transformer-encoded temporal sequence and the amortized inference network’s multi-headed architecture enables the analysis platform 119 to capture subtle early symptom patterns, compensate for delayed diagnostic labels, and make temporally informed predictions that would not be achievable using traditional machine-learning models. The hybrid use of ensemble-derived feature weights from the first stage and deep temporal inference from the second stage allows the system to detect complex symptom interactions, such as the co-occurrence of severe constipation, disrupted REM sleep, and cognitive decline, which are strongly indicative of early-stage PD. This hybrid design enhances diagnostic accuracy, improves computational efficiency, and provides robust predictions even when symptom data is irregular, incomplete, or temporally misaligned.

[0094] To prevent collapse to relatively trivial solutions in which the model relies on diagnosis timing rather than learning physiologically meaningful PD-relevant symptom patterns, the Detection system 100 may incorporate a set of architectural and training-based safeguards. In some embodiments, the system may implement structured priors encoding medically-informed diagnostic-delay distributions, wherein a log-normal prior (pprior(<5) = LogNormal z = log (4), <7 = 0.5)) is imposed via a prior-loss term Lprior= KL(pmodel(<5) || Pprior(<5))to penalize delay predictions that may deviate from physiologically plausible ranges; an auxiliary symptom-reconstruction task, in which a decoder network predicts symptom vectors S(T, t) conditioned on an inferred onset time rand query time t, with reconstruction loss:

[0095] thereby forcing the temporal model to utilize the inferred onset time meaningfully to explain observed symptom-progression patterns rather than ignoring onset structure; and simulation-based validation, wherein synthetic patient data with known ground-truth onsettimes and diagnostic delays is generated using disease-progression models, enabling verification that the system correctly recovers true latent parameters and does not instead learn spurious correlations tied to the timing of clinical diagnoses. Together, these mechanisms constrain the temporal uncertainty learning architecture to capture symptom dynamics.

[0096] Further, the hybrid approach may enable the analysis platform 119 to continuously learn as more data is entered and stored. In various instances, the detection system 100 may learn and train from not only input user and clinician data but also verified PD diagnoses. For example, a list of signs and symptoms for a patient with a verified positive PD diagnosis may be entered into the detection system 100. In this example, the detection system 100 may learn and weigh the symptoms that ultimately lead to the positive PD diagnosis and update the detection models by adjusting the weighing scale to incorporate the newly learned data for diagnosing PD. In other words, because of the various detection models incorporated within the detection system 100, the detection system 100 may train and learn from a combination of one or more patient profiles and various PD diagnoses to increase accuracy and specificity for future PD diagnosis predictions. Moreover, the detection system 100 may utilize one or more weighted input vectors, adjusted according to the learned symptom weights, to refine diagnostic accuracy based at least in part on the relative importance of each symptom. The learned symptom weights may then be used to adjust input vectors during a prediction phase, thereby refining the precision and accuracy of the predictions provided by the detection system 100 based on the relative importance of each symptom.

[0097] In addition to the two-stage inference pipeline, the detection system 100 may include one or more auxiliary components that provide explainability, interpretability, and exploratory analysis capabilities. In various embodiments, the auxiliary components may operate on the outputs of Stage 2 to support clinical decision-making and system validation. The detection system 100 may additionally or alternatively include one or more model-agnostic methods and / or unsupervised learning models 135 to enhance explainability, interpretability, and transparency of the system's diagnostic outputs. In some instances, the detection system 100 may incorporate a method, model, algorithm, or technique directed to explaining the output of other detection models such as the supervised learning models 130 (e.g., Random Forest and Neural Network). For example, the detection system 100 may incorporate SHapley Additive exPlanations (“SHAP”) to provide a detailed and transparent breakdown of each model’s predictions into contributions attributable to individual symptoms (i.e., assign importance scores to individual symptoms). Based at least in part on the output from sHAP, the detectionsystem 100 may recognize, compare, and flag different contributions from various signs and symptoms for predicting a PD diagnosis. For example, the detection system 100 may flag that sleep disturbances approximately 40% of the model’s predictive signal, anxiety contributes approximately 30%, and remaining symptoms constitute the final 30%. Such interpretability insights may assist clinicians, researchers, and end-users in understanding the rationale behind the diagnostic outcome.

[0098] In other instances, the detection system 100 may include one or more unsupervised learning techniques configured to perform exploratory analyses of patient symptom data. For example, the detection system 100 may utilize K-means clustering to group patients based at least in partially on similarities in their non-motor symptoms. These cluster-based groupings may reveal potential subtypes of PD or distinct symptom-progression trajectories, allowing deeper insight into disease heterogeneity and variability among patients. Additionally, such unsupervised techniques (e.g., clustering) may improve downstream diagnostic performance by uncovering structure in the data that is not easily captured by supervised models alone, thereby enabling the Detection system 100 to better understand early-stage patterns, symptom progression, and variability.

[0099] In some embodiments, the detection system 100 may implement a comprehensive training framework configured to optimize the parameters of both stages of the inference pipeline and the auxiliary components. The training framework or training process may occur prior to deployment and may be repeated periodically as new data becomes available. In some embodiments, the training framework may be organized into one or more phases. For example, the training framework may include phases directed to dataset preparation, model optimization, and quantitative validation using standard classification metrics. In various embodiments, the detection system 100 may implement a rigorous training and evaluation framework to optimize the detection models and validate performance prior to clinical deployment. The framework may be organized into phases that include dataset preparation, model optimization, and quantitative validation using standard classification metrics. Representative metrics may include accuracy, precision, recall, and Fl score, as well as confusion matrices to visualize true-positive / false-positive trade-offs and guide targeted refinements. For example, if the system occasionally misclassifies severe sleep disturbances as a non-disease state, additional training data emphasizing sleep-related features may be incorporated to correct the error and improve sensitivity and specificity for the relevant non-motor domain.

[0100] In some embodiments, during training, the system may optimize a combinedobjective that integrates classification accuracy, onset and delay inference, medically informed priors, and auxiliary consistency constraints:" " "

[0101] where ^classification may be a binary cross-entropy loss for disease prediction; LOnsetar|d Ldelaymay be mean-squared errors for onset-time and diagnostic-delay estimation; Lprior may be a KL divergence term constraining predicted delays to medically informed distributions; ^reconstruction may require symptom -trajectory reconstruction conditioned on inferred onset; and ^monotonicity may penalize configurations in which increased symptom severity would otherwise decrease predicted risk. In some embodiments, training may employ an Adam optimizer with a learning rate in a range from about le — 5 to about le — 2 (e.g., the Adam optimizer may be imparted with a learning rate of about le — 4). Additionally, the Adam optimizer may employ a batch size in a range from about 8 to about 256 (e.g., the batch size may be about 32). A gradient clipping may be imparted with a maximum norm in a range from about 0.5 to about 5.0 (e.g., the gradient clipping may be imparted with a maximum norm of about 1.0). A rate of a dropout regularization may be in a range from about 0.0 to about 0.5 (e.g., the rate may be about 0.1 to about 0.3). An L2 weight decay may be imparted with a coefficient in a range from about le — 6 to about le — 3, (e.g., the L2 weight decay coefficient may be about le — 5). An early stopping with patience may be in a range from about 5 to about 50 epochs (e.g., about 20 epochs). The loss weights may be set with a w n a range from about 0.5 to about 2.0 (e.g., the loss weight w4may be about 1.0); a w2in a range from about 0.1 to about 1.0 (e.g., the loss weight w2may be about 0.5); a w3in a range from about 0.1 to about 1.0 (e.g., the loss weight w3may be about 0.5); a w4in a range from about 0.1 to about 0.05 (e.g., the loss weight w4may be about 0.1); a w5in a range from about 0.0 to about 0.2 (e.g., the loss weight w5may be about 0.05); and a w6in a range from about 0.0 to about 0.5 (e.g., the loss weight w6may be about 0.1).

[0102] Following initial optimization, the detection models may be scaled and at least periodically updated to maintain performance as new data arrive. Scaling may include adjusting loss weights based on validation feedback, fine-tuning conformal calibration sets using recent cohorts, updating ensemble feature-importance scores if new symptom features are introduced, and recalibrating prediction intervals to preserve coverage guarantees. These updates may be applied in a scheduled or continual manner to sustain accuracy and reliability under evolving data distributions.

[0103] As part of the training and evaluation framework, the detection system 100 may implement an uncertainty quantification. In some embodiments, the uncertainty quantification may be provided as a horizon-specific uncertainty quantification through conformal prediction to generate calibrated prediction sets and calibrated risk estimates for multiple prediction horizons. In some embodiments, the system may maintain separate calibration sets for each horizon: h G {3-year, 5-year, 7-year, thus providing calibration that reflects the differing levels of predictive uncertainty associated with short-term versus long-term forecasting.

[0104] For each prediction horizon h, the system may construct a calibration set Causing held-out validation data. For every validation example, the true outcome at horizon h, denoted yhmay be paired with the corresponding model prediction y^, and a nonconformity score may be computed as:n, (ft) (ft) iRi =l yt - yt l

[0105] The calibration set Ch= {R1,R2, —,Rn} may storeaH nonconformity scores for that horizon. For a new test patient with predicted risk ytestat horizon h, the system may construct a prediction interval with desired coverage probability 1 — a(e.g., 90% coverage for a = 0.1) as:[ytest—Ql-a(Ol)' Ytest T Ql-a(Ol)]

[0106] where q1-a(Cft) may denote the 1 — a quantile of the calibration set. In some embodiments, the interval endpoints may be clipped to [0* 1] to ensure that risk estimates remain within valid probability bounds.

[0107] Because each prediction horizon maintains its own calibration set and quantile thresholds, the detection system 100 is able to provide horizon-specific uncertainty quantification. Shorter-term prediction horizons (e.g., 3 years) may produce narrower, more certain intervals, while longer horizons (e.g., 7 years) may yield wider calibrated intervals reflecting increased long-term uncertainty. This separation advantageously ensures that risk estimates are communicated in a clinically meaningful and statistically reliable manner, tailored to each temporal horizon used by the system.

[0108] During the prediction phase (i.e., the inference phase), the detection system 100 may execute the two-stage pipeline, as described above, to generate PD diagnosis predictions in real-time or near real-time. In some embodiments, the prediction phase may predict a PDdiagnosis (e.g., by determining a likelihood of PD and / or a likelihood of a PD diagnosis) for a patient based at least in part on the input data for said patient. Generally, the PD diagnosis prediction provided by the detection system 100 may be based on a comprehensive analysis of the input data (e.g., patient signs and symptoms, medical records, notes, family history, and the like) and any available stored data in memory 120. In some instances, the PD diagnosis prediction may include weighing the input data, flagging changes and / or potentially alarming signs and symptoms, and comparing it to learned data sets that produced a positive PD diagnosis.

[0109] When determining a PD diagnosis prediction, the detection system 100 may recognize a negative trend in signs and may flag one or more symptoms as indicators for a potentially positive PD diagnosis. Thus, the detection system 100 may provide a positive PD diagnosis prediction to the patient. In contrast, the detection system 100 may recognize no correlation or warning trends in signs and symptoms and may not identify any concerning symptoms for a positive PD diagnosis. Thus, the detection system 100 may provide a negative PD diagnosis prediction. When supplying one or more users with the PD diagnosis prediction, the Detection system 100 may communicate to the user via at least one of the user devices 105, 110 and the communication interface 135. For example, the prediction may appear as, “Parkinson’s Disease POSITIVE” or “Parkinson’s Disease NEGATIVE,” based on the comprehensive analysis. Additionally, or alternatively, the prediction may appear in another language or visual format.

[0110] In some embodiments, the prediction generated by the detection system 100 may be provided to users via an interactive display, providing immediate and intuitive access to the diagnostic result. For example, a user may log in to a dedicated portal that directs them to an associated interface, where they may securely view their personalized prediction. In some embodiments, the results presented through the interface may not be limited to a single text output. Instead, the system may deliver a range of detailed information, including probability scores, risk levels, recommended follow-up actions, and explanatory notes configured to help users and clinicians interpret the prediction in context.[OHl] In various instances, the detection system 100 may deliver the prediction to the user in near real-time with respect to data entry. For example, the detection system 100 may display the PD diagnosis prediction nearly immediately after a user inputs the data (e.g., via the responses to the questionnaire). Alternatively, the detection system 100 may display the PD diagnosis predictions after a certain delay. In some instances, the delay may be dependent onthe completion of the data entered. For example, the detection system 100 may recognize missing data from the other of the one or more user devices 105, 110 or missing data such as the patient’s medical history or family history. In other instances, the delay may be dependent on other factors such as connectivity, other preset parameters (e.g., a preset configuration to delay the diagnosis until 24 hours), amongst others. In some instances, based at least in part on the predicted diagnosis and calibrated uncertainty, the detection system 100 may output a recommended clinical action selected from a previously defined constrained action set. In such embodiments, the constrained action set may include recommending additional monitoring, recommending confirmatory evaluation by a specialist, and / or recommending follow-up scheduling at a defined interval.

[0112] Following the prediction phase, the detection system 100 implements a reinforcement learning framework. In some embodiments, the reinforcement learning framework may be configured to continuously refine its action- sei ection policies based at least partially on the outcomes from verified diagnoses. This feedback loop may function separately from the inference process while still influencing the system’s actions in the future. To enhance accuracy of the detection process, the detection system 100 may include one or more reinforcement learning (“RL”) models 140. Within the RL framework, patient profiles, including their recorded non-motor signs and symptoms, may be treated as states, while selection among constrained output actions (e.g., additional monitoring, specialist referral, and / or follow-up scheduling at a defined interval) and / or selection of a prediction horizon may be treated as actions. In various instances, the RL framework may reward the detection system 100 for correct PD predictions, whereas errors or misdiagnoses may incur penalties.

[0113] In various instances, the detection system 100 may include one or more algorithms configured to reinforce learning and / or enhance action- sei ection. In some instances, the one or more algorithms may be model based. In other instances, the one or more algorithms may be model-free. For example, the detection system 100 may include a Q-learning algorithm configured to evaluate predictions and track accuracy for different decision-making scenarios. In some instances, the Q-leaming algorithm may reward points based on accurate diagnosis predictions and / or deduct points based on inaccurate diagnosis predictions. In some instances, a reward may be a +1 point, or an addition of one point, while a penalty may be a -1 point, or a deduction of one point. The reward may add points by any suitable integer and the penalty may deduct points by any suitable integer. In some instances, a number of points added via a reward may be the same as or different from a number of points deducted via a penalty. TheQ-learning algorithm may be configured to develop a Q-table for managing a score that is calculated based at least in part on the rewards and penalties from accurate and inaccurate diagnosis predictions made by the Detection system 100. Generally, over multiple training episodes, the Detection system 100 may learn to prioritize symptom patterns that strongly correlate with PD, thereby enhancing prediction accuracy and precision in diverse scenarios.

[0114] Turning to FIG. 3, a schematic of a workflow process 200 executable by the detection system 100 is provided. The process 200 may be implemented to provide predictions of a PD diagnosis. In various instances, the process 200 may be executable by one or more processors (not illustrated) of the detection system 100. The process 200 may include the step 205 of receiving input data. In some instances, receiving input data may come from user devices 105, 110 or any other remote device. Generally, the step 205 may include retrieving data such as non-motor signs and symptoms of PD.

[0115] The process 200 may also include the step 210 of executing data preprocessing techniques. In various instances, the step 210 may include aggregating the input data, the data stored in memory 120, and, if integrated, any data stored in a cloud. Additionally, data preprocessing at step 210 may include analyzing the data for completion. For example, the detection system 100 may verify if all necessary responses from the users (e.g., through the questionnaire) have been provided. In some instances, if the detection system 100 identifies missing data, then the detection system 100 may prompt a user to input said missing data.

[0116] During data preprocessing at step 210, the data sets may be preprocessed in a chronological order. For example, after a first data set completes preprocessing, the same first data set may immediately undergo data processing at step 215, while a second data set undergoes preprocessing. In such instances, the chronological order at which the data sets are preprocessed may be preset by a user. For example, when configuring the detection system 100, a user may set data preprocessing to always begin with data relating to a certain symptom. Alternatively, the chronological order may be based on the detection system 100 learning and training. For example, if the detection system 100 learns that one or more symptoms are weighted more (e.g., are larger indicators of PD), then the detection system 100 may preprocess the data sets relating to these symptoms prior to other data sets. In other instances, the detection system 100 may perform data processing for all data sets simultaneously. In various instances, step 210 produces normalized data sets.

[0117] The process 200 may also include the step in which the normalized data from step210 is processed. Data processing at this step may include comprehensively analyzing the normalized data using one or more detection models included in the detection system 100. The detection models of the detection system 100 may include one or more supervised learning models 130 and / or unsupervised learning models 135. In some instances, the detection system 100 may include model-agnostic methods in addition to or in the alternative to unsupervised learning models 135. In some instances, the detection system 100 may include a combination of supervised learning models 130, unsupervised learning models 135, and model-agnostic methods.

[0118] In various embodiments, the steps 215-219 may be substantially directed to executing the two-stage interference pipeline. In such embodiments, a step 215 may be substantially related to Stage 1 (i.e., feature weighing), wherein stage 215 may include applying ensemble feature importance weights. For example, applying the ensemble importance weight may be computed or identified via Random Forest and / or XGBoost to the normalized data, resulting in the generation of weighted feature vectors. Steps 216-219 of the process 200 may be substantially related to Stage 2 (i.e., the temporal neural network). In some embodiments, the step 216 may include processing the weighted feature vectors through the transformer encoder to generate temporal representations. Further, the step 216 may include computing onset time, delay, and disease probability predictions via the three parallel MLP heads. A step 217 may include applying delay-aware attention based on the inferred onset time. A step 218 may include marginalizing over onset uncertainty by sampling from the posterior distribution and computing weighted predictions. A step 219 may include generating a final PD risk prediction.

[0119] The detection models (e.g., the supervised and unsupervised learning models 135, 130) may cooperate to produce a PD diagnosis prediction based at least in part on the data input by the user at step 205. Data processing may include using the detection models to identify, highlight, and prioritize certain data sets. In various instances, due to the comprehensive analysis performed during data processing, the detection system 100 may identify one or more symptoms, or a specific combination of symptoms, which are prevalent in a positive PD diagnosis. For example, the detection system 100 may have previously learned that a specific non-motor symptom is relatively highly prevalent in positive PD diagnoses, therefore the detection system 100 may assign said specific non-motor symptom a higher weight or identify the specific non-motor symptom as a red flag. In another example, the Detection system may identify that a combination of symptoms (e.g., severe constipation, disrupted REM sleep, andcognitive decline) are strongly prevalent in a positive PD diagnosis. As such, during data processing, the detection system 100 may search for the presence of the identified combination of symptoms.

[0120] The detection system 100 may include detection models configured to capture and understand linear or non-linear relationships within the data and between multiple data sets to formulate nuanced PD diagnosis predictions. In some instances, the detection models may be configured to provide rationale behind diagnostic outcomes. For example, the detection models may be configured to provide details regarding the weighted symptoms. In some instances, the detection system 100 may include a detection model configured to perform additional dissection of the data sets. For example, a detection model may be included for grouping patients having similar signs and symptoms.

[0121] At step 220 of the process, the detection system 100 may determine a PD diagnosis prediction for the patient based at least in part on the comprehensive analysis performed at steps 215-219. The PD diagnosis prediction may include or indicate a likelihood of PD and / or a likelihood of a PD diagnosis. The PD diagnosis prediction may be positive, meaning a patient is predicted to have PD, or negative, meaning the patient, at the time, is not predicted to have PD.

[0122] The PD diagnosis prediction may then be provided to one or more users (e.g., the patient and a clinician). In some instances, the detection system 100 may provide a recommendation based at least in part on the results of the PD diagnosis prediction. For example, the detection system 100 may be configured to provide additional monitoring for patients that receive a positive PD diagnosis prediction. Further, if additional monitoring continues to predict a PD positive diagnosis and / or determines a trending decline in health (e.g., a worsening in symptoms), then the detection system 100 may recommend an appointment with a medical professional. Alternatively, the detection system 100 may provide additional monitoring for patients that may be trending towards a positive PD diagnosis prediction. In some instances, the Detection system 100 may determine and / or provide a risk level of developing PD to the patient and / or medical professional. In some instances, the detection system 100 may interface with another existing software system, such as an online patient portal, configured to facilitate access to health information and records, communication, and scheduling medical appointments.

[0123] In some embodiments, determining the PD diagnosis may include providing the PDdiagnosis prediction. For example, providing the diagnosis may include outputting or displaying whether the diagnosis prediction is positive (e.g., "Parkinson's Disease POSITIVE") or the diagnosis prediction is negative (e.g., "Parkinson's Disease NEGATIVE”) to a user device. A step 225 may include providing the uncertainty quantification (i.e., horizonspecific) confidence intervals generated via conformal prediction (e.g., for 3-year prediction a probability with a confidence interval of about 76% to about 96% interval at 90% coverage). A step 226 may include displaying the estimated onset time and diagnostic delay. A step 227 may include optionally providing a visualization of onset uncertainty showing the probability distribution over candidate onset times. A step 228 may include presenting feature importance scores indicating which symptoms most strongly influenced the prediction. A step 229 may include recommending a clinical action based on the predicted diagnosis and uncertainty.

[0124] The process 200 may include a step 230 directed to reinforcing data preprocessing and processing techniques. The step 230 may include comparing the PD diagnosis prediction to a verified PD diagnosis. In some instances, a verified PD diagnosis may be from a neurologist or physician that examines a patient and diagnoses the patient as either positive or negative for PD. In the process 200, the reinforcing data preprocessing and processing techniques may include rewarding or penalizing the detection system 100 for an accurate or inaccurate prediction, respectively, based on the comparison to the verified PD diagnosis.

[0125] The process 200 may also include the step 235 in which the detection models are trained based on the reinforcement at step 230. In some instances, step 235 may occur concurrently or in conjunction with any one of steps 205-229. That is, training the detection models for data preprocessing and processing may occur while data is input, preprocessed and / or processed and / or while predictions are being determined and provided. In other words, training detection models may be continuous or continual. For example, the detection system 100 may continuously or continually compare previous PD diagnosis predictions against verified PD diagnoses and train the detection models to ensure that future PD diagnosis predictions are minimally skewed by delays in updates to the process 200.

[0126] Turning to FIG. 4, an example training workflow or training method 300 is provided. The training method 300 may be executed by the detection system 100 during the process 200 (e.g., at step 235). In various instances, the training method 300 may be executable by one or more processors (not illustrated) of the detection system 100. In some instances, the training method 300 may be executed independently or separately from the process 200 such as when the detection system 100 is not actively providing a PD diagnosis prediction. In variousinstances, the training method 300 may be continuously or continually executed. In other instances, a user may determine when the training method 300 may be executed.

[0127] The training method 300 may be configured to optimizes the parameters of the two-stage inference pipeline (including both the ensemble feature weighting stage and the temporal neural network stage) as well as auxiliary components. According to various embodiments, the training method 300 may include the step 305 of receiving PD data and diagnoses. Here, the detection system 100 may receive information from patient profiles previously diagnosed with PD from a clinician. In some instances, the data received at step 305 may include longitudinal symptom data with corresponding timestamps; verified diagnosis labels indicating PD or healthy control; diagnosis times; optionally, estimated or retrospectively determined onset times when available; and demographic and medical history information.

[0128] The training method 300 may also include the step 310 of training the detection models to learn from the data received at step 305. One example of training may be identifying new signs, symptoms, and patterns that are relevant for a positive PD diagnosis. In some instances, training may include learning signs and symptoms particular to PD to avoid a misdiagnosis. In such instances, the detection system 100 may develop a criteria containing one or more signs and symptoms that are relatively prevalent in a true positive PD diagnosis. In other instances, other than solely learning from PD information, the detection system 100 may receive information relating to other diseases that may present similarly as PD. By learning and training on diseases similar to PD, the Detection system 100 may add limitations to the criteria or otherwise fine tune the criteria of signs and symptoms unique to PD, thus decreasing the chances of misdiagnoses. One benefit from this step 310 may be the reduction of false positives and false negatives when providing the PD diagnosis prediction.

[0129] In various embodiments, training at step 310 may comprise optimizing the combined loss function:Ljotal = W L-Classification + w2L_onset + w3L_delay + w4L_prior+ w5L_reconstruction + w6L_monotonicity

[0130] wherein the loss components may respectively penalize classification errors, onset time prediction errors, delay prediction errors, deviations from medical delay priors, failures to reconstruct symptom trajectories from inferred onset times, and non-monotonic relationships between symptom severity and predicted risk. In some embodiments, the training process may utilize an Adam optimizer with a learning rate of about le-4, batch size of about 32, gradientclipping with a maximum norm of about 1.0, early stopping with patience of about 20 epochs, and regularization techniques including dropout (rates 0.1-0.3), L2 weight decay (coefficient le-5), label smoothing, and data augmentation including temporal jittering, symptom noise injection, and sequence dropout.

[0131] The training method 300 may also include a step 315 in which, based on the training from step 310, the detection models can be scaled to increase accuracy. In such instances, scaling the detection models may include adjusting any scales (e.g., the weighing scale) used for forming a final PD diagnosis prediction. Scaling at step 315 may additionally include at least one of adjusting the loss weights based on validation performance, fine-tuning conformal prediction calibration sets using recent data, updating ensemble model (Random Forest and / or XGBoost) feature importance scores if new symptom features are added, and recalibrating prediction intervals to maintain coverage guarantees.

[0132] In various instances, the training method 300 may include the reinforcement learning models 140 and / or the one or more reinforcement learning algorithms such as the Q-leaming algorithm. For example, during training, the detection system 100 may select a patient with a verified PD diagnosis, allowing the detection system 100 to use said patient as a truth. Based on the patient’s data the detection system 100 may determine a PD diagnosis prediction. If the detection system 100 accurately predicts the verified PD diagnosis, then the detection system 100 may be rewarded via point(s) added to a score tracked by the Q-learning algorithm. On the other hand, if the detection system 100 inaccurately predicts the verified PD diagnosis, then the detection system 100 may be punished via point(s) deducted from the score tracked by the Q-leaming algorithm. Thus, the detection system 100 may train on various scenarios to enhance decision making. In some instances, the Q-leaming algorithm may update the Q-table iteratively, enabling the detection system 100 to learn from past decisions and continuously refine its accuracy across multiple training episodes. In other instances, verified PD diagnoses may be retroactively input for patients. In such instances, the detection system 100 may train from the retroactive diagnosis.

[0133] There are multiple benefits to detection system 100 including, but not limited to, providing a comprehensive analysis of patient’s overall health, monitoring non-motor symptoms associated with earlier stages of PD rather than delaying care until detrimental motor symptoms are presented, and facilitating access to care. The detection system 100 described herein provides a platform by which clinicians and other medical personnel may receive realtime updates to a patient’ s condition, thus permitting faster intervention and treatment. Becausethe detection system 100 allows for remote monitoring of patients, clinicians may have more time to allocate to other clinical visits. Further, regarding patient quality of life, the detection system 100 may alleviate stressors such as anxiety regarding a patient’s health, for example, by providing recommendations on when to schedule clinician visits. As such, the detection system 100 provides an outcome (e.g., a PD diagnosis prediction) that otherwise could not be provided, achieved, accomplished, or done without the use of the specifically trained algorithm, models, techniques, and approaches described herein.

[0134] Additional benefits and technical advantages provided by the present disclosure include, but are not limited to: a scalable and adaptable system for detecting and / or monitoring early symptoms of PD that can be implemented in various environments and use cases; enhancing accessibility to clinics and clinicians with limited resources via a readily launched application that can be accessed via existing user devices such as smart devices, phones, laptops, desktop PCs, etc.; enabling broader accessibility via integration with a cloud system and enabling remote predictions of PD and real-time data sharing; the provision of a smart device-compatible detection system that can extend its utility to home-based monitoring, empowering patients to track non-motor symptoms and receive preliminary assessments from the comfort of their homes; the provision of a highly detailed, modular, and scalable framework for the early detection and monitoring of PD; enabling the revelation of early signs and symptoms of PD using strategically trained models and non-motor signs and symptoms that would otherwise go undetected or ignored; bridging the gap between traditional diagnostic practices for PD and the availability of early stage information and data that can provide accurate, early, and reliable diagnostic predictions; the integration of advanced machine learning, explainability tools, and a user-friendly interface positions it as a transformative tool in the field of neurological diagnostics; addressing the fundamental technical challenge of learning from systematically delayed diagnostic labels through explicit modeling of temporal uncertainty; providing calibrated uncertainty quantification that distinguishes between model uncertainty, onset uncertainty, and horizon-specific prediction uncertainty; enabling early detection years before clinical diagnosis through identification of pre-diagnostic non-motor symptom patterns; and offering interpretable predictions through feature importance attribution that support clinical decision-making; among other advantages including those that are described and can be appreciated from the above description.

[0135] The systems and methods of this disclosure may be implemented in a variety of use cases. In one example, aspects of the present disclosure can be utilized by a primary carephysician (“PCP”) responsible for referring patients to specialists to receive further care. The systems and methods hereof can enable the PCP to refer to a neurologist patients exhibiting non-motor symptoms (e.g., sleep problems, constipation, anxiety, etc.) that may be manifest of PD, where traditionally the primary referral cases involved patients exhibiting motor symptoms. The systems and methods thereby provide an opportunity for a PCP to make an accurate PD diagnosis in a timely manner, particularly in patients presenting with only nonmotor symptoms, avoiding unnecessary referrals, and enabling patients in the earliest stages of PD to receive the specialized care that they need. Additionally, or alternatively, the systems and methods hereof can assist neurologists in the detection of PD at early stages thereof, and / or in differentiating between PD and other movement disorders that can manifest in similar motor and / or non-motor symptoms, such as Dementia with Lewy Bodies, Multiple System Atrophy, etc. In some instances, the systems and methods can be used by patients previously diagnosed with PD to track the progression of their disease and related symptoms. Additionally, or alternatively, the systems and methods can enable remote monitoring of patients diagnosed with PD and reduce costs and travel expenses associated with clinical visits by allowing such visits to be allocated to those patients with only severe changes in symptoms.

[0136] The various illustrative blocks and components described in connection with the disclosure herein may be implemented or performed using a general-purpose processor, a DSP, an Application-Specific Integrated Circuit (ASIC), a Central Processing Unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a Field Programmable Gate Array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor but, in the alternative, the processor may be any processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices (for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration). Any functions or operations described herein as being capable of being performed by a processor may be performed by multiple processors that, individually or collectively, are capable of performing the described functions or operations.

[0137] The functions described herein may be implemented using hardware, software executed by one or more processors, firmware, or any combination thereof. If implemented using software executed by multiple processors, the functions may be stored as or transmittedusing one or more instructions or code of a computer-readable medium. Other examples and implementations are within the scope of the disclosure and appended claims. For example, due to the nature of software, functions described herein may be implemented using software executed by one or more processors, hardware, controllers, firmware, hardwiring, circuitry, or combinations of any of these. Features implementing functions may also be physically located at various positions, including being distributed such that portions of functions are implemented at different physical locations.

[0138] Computer-readable media includes both non-transitory computer storage media and communication media including any medium that facilitates transfer of a computer program from one location to another. A non-transitory storage medium may be any available medium that may be accessed by a general-purpose or special-purpose computer. As used herein, the term “non-transitory” may be interpreted not as an eternal characteristic of a state, but as a characteristic of a state that will last for a period of time. The term “non-transitory” may specifically disavow fleeting characteristics such as characteristics of a particular carrier wave or signal or other forms that exist only transitorily in any place at any time. By way of example, and not limitation, non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), a cache, tape, electrically erasable programmable ROM (EEPROM), flash memory, compact disk (CD) ROM or other disk or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that may be used to carry or store desired program code means in the form of instructions or data structures and that may be accessed by a general-purpose or special-purpose computer or a general-purpose or specialpurpose processor. Also, any connection may be properly termed a computer-readable medium. For example, the software may be transmitted from a website, server, or other remote source using a wired technology such as a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), universal serial bus (USB), high-definition multimedia interface (HDMI), video graphics array (VGA), digital visual interface (DVI), thunderbolt cable, power cable, ribbon cable, integrated services digital network (ISDN), or wireless technologies such as wireless fidelity (Wi-Fi), Bluetooth, cellular network, near-field communication (NFC), Zigbee, long range (LoRa), infrared (IR), RFID, light fidelity (Li-Fi), satellite, ultra-wideband (UWB), millimeter wave (mm Wave), and microwave. The wired and or wireless technologies are included in the definition of computer-readable medium. Disk and disc, as used herein, include a compact disk (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk,and Blu-ray disc. Disks or discs may reproduce data magnetically or optically using lasers. Combinations of the above are also included within the scope of computer-readable media. Any functions or operations described herein as being capable of being performed by a memory may be performed by multiple memories that, individually or collectively, are capable of performing the described functions or operations.

[0139] Information and signals described herein may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0140] Machine learning includes one or more systems that may be adapted to generate, train, and execute one or more trained learning models, nodes, neural networks, gradient boosting algorithms, mutual information classifiers, random forest classifications, and other machine learning and artificial intelligence-based algorithms or models to process inputs, parameters, and other data elements. In some examples, the one or more trained learning models can include deep learning, machine learning, neural networks, computer vision, and similar advanced artificial intelligence-based technologies. In some examples, machine learning or artificial intelligence may use additional inputs or feedback loops to an iterative training process for enhanced data processing and improved outcomes.

[0141] Although the disclosure may describe components and functions that may be implemented in a particular example with reference to a particular standard or protocol, the disclosure is not limited to the standard or protocol. Other standards or protocols supporting similar functionality are considered equivalents thereof.

[0142] As used herein, including in the claims, “or” as used in a list of items (for example, a list of items prefaced by a phrase such as “at least one of’ or “one or more of’) indicates an inclusive list such that, for example, a list of at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Furthermore, “and / or” as used in a list of items indicates an inclusive list such that, for example, a list of at least one of A, B, and / or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Also, as used herein, the phrase “based on” shall not be construed as a reference to a closed set of conditions. For example, an example step that may be described as “based on condition A” may be based on both a condition A and a condition B. For example, as used herein, the phrase “based on” shall beconstrued in the same manner as the phrase “based at least in part on.”

[0143] As used herein, including in the claims, the article “a” before a noun may be open-ended and understood to refer to “at least one” of those nouns or “one or more” of those nouns. Thus, the terms “a,” “at least one,” “one or more,” and “at least one of one or more” may be interchangeable. For example, if a claim recites “a component” that performs one or more functions, each of the individual functions may be performed by a single component or by any combination of multiple components. Thus, the term “a component” having characteristics or performing functions may refer to “at least one of one or more components” having a particular characteristic or performing a particular function. Subsequent reference to a component introduced with the article “a” using the terms “the” or “said” may refer to any or all of the one or more components. For example, a component introduced with the article “a” may be understood to mean “one or more components,” and referring to “the component” subsequently in the claims may be understood to be equivalent to referring to “at least one of the one or more components.” Similarly, subsequent reference to a component introduced as “one or more components” using the terms “the” may refer to any or all of the one or more components. For example, referring to “the one or more components” subsequently in the claims may be understood to be equivalent to referring to “at least one of the one or more components.”

[0144] The terms “determine,” “determining,” “identify,” or “identifying” encompasses a variety of actions and, therefore, “determining” or “identifying” can include calculating, computing, processing, deriving, investigating, looking up (such as via looking up in a table, a database, or another data structure), receiving, ascertaining, and the like. Also, “determining” or “identifying” can include receiving (for example, receiving information), accessing (for example, accessing data stored in memory), retrieving, and the like. Also, “determining” or “identifying” can include resolving, obtaining, selecting, choosing, establishing, and other such similar actions.

[0145] In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If just the first reference label may be used in the specification, the description may be applicable to any one of the similar components having the same first reference label irrespective of the second reference label or other subsequent reference label.

[0146] The description set forth herein, in connection with the appended figures, describesexample configurations and does not represent all the examples that may be implemented or that are within the scope of the claims. The term “example” used herein means “serving as an example, instance, or illustration” and not “preferred” or “advantageous over other examples.” The detailed description includes specific details for the purpose of providing an understanding of the described techniques. These techniques, however, may be practiced without these specific details. In some figures, known structures and devices are shown in block diagram form in order to avoid obscuring the concepts of the described examples. The features of the various examples described herein may be combined in any suitable manner. It is contemplated that one or more features from one example may be incorporated into another example unless explicitly stated otherwise. The combinations of features from different examples are within the scope of the disclosure.

[0147] It will be appreciated by those skilled in the art that while the disclosure has been described above in connection with particular examples, the disclosure is not necessarily so limited, and that numerous other examples, uses, means for, modifications and departures from the examples, uses, and means for are intended to be encompassed by the claims attached hereto. The entire disclosure of each patent and publication cited herein is incorporated by reference, as if each such patent or publication were individually incorporated by reference herein. Various features and advantages of the disclosure are set forth in the following claims.

Claims

1. CLAIMS1. A system for generating a prediction indicative of Parkinson’s Disease (PD) risk, comprising:a memory storing instructions; anda processor configured to execute the instructions to:receive data entries for a user, each data entry comprising symptom data and an associated timestamp;analyze the symptom data to identify feature-importance weights; apply the feature-importance weights to the symptom data to generate weighted representations;process the weighted representations to interpret symptom progression over time;adjust the process using disease-relative timestamps derived from an inferred onset;generate inference outputs comprising at least one of an onset estimate, a diagnostic-delay estimate, and a risk value; anddetermine a diagnostic prediction based at least in part on the risk value.

2. The system of claim 1 wherein the symptom data includes non-motor symptoms of a user.

3. The system of claim 1, wherein the disease-relative timestamps are derived by transforming at least a portion of the associated timestamps using the inferred onset.

4. The system of claim 1, wherein the processor is configured to compute the feature-importance weights using at least one ensemble classifier, the at least one ensemble classifier being configured to rank features of the symptom data by predictive contribution and apply the feature-importance weights to emphasize higher-contribution features.

5. The system of claim 1, wherein the processor is further configured to:derive disease-relative positional information from the disease-relative timestamps; andincorporate the disease-relative positional information into an attention operation of a transformer encoder.

6. The system of claim 1, wherein the processor is configured to generate the onset estimate, the diagnostic-delay estimate, and the risk value based at least in part on multiple prediction heads that are trained jointly.

7. The system of claim 1, wherein the processor is configured to provide an uncertainty value, and wherein providing the uncertainty value comprises generating a horizon-specific calibrated prediction interval using conformal prediction with a separate calibration set for each of a plurality of prediction horizon.

8. The system of claim 1, wherein the processor is configured to select or recommend one or more clinical actions from a constrained action set.

9. The system of claim 8, wherein the constrained action set is continuously updated based at least in part on outcomes associated with prior predictions.

10. The system of claim 1, wherein the processor is configured to store, retrieve, and process temporal sequences, encoded representations, calibration artifacts, and model checkpoints using a database that supports longitudinal clinical records.

11. The system of claim 1, wherein the processor is configured to continuously refine at least one of the feature-importance weights, calibration sets, or model parameters based on newly acquired data and verified diagnoses to maintain or improve predictive performance over time.

12. A system for generating a prediction indicative of Parkinson’s Disease (PD) risk, comprising:a memory storing instructions; anda processor configured to execute the instructions to:receive data entries for a user, each data entry including non-motor symptom data and an associated timestamp;analyze the non-motor symptom data to identify feature-importance weights; apply the feature-importance weights to the non-motor symptom data to generate weighted representations;construct a disease-relative temporal frame by transforming at least a portion of the timestamps into disease-relative timestamps based at least in part on an inferred onset estimate;process the weighted representations to interpret symptom progression over time;generate inference outputs comprising at least one of an onset estimate, a diagnostic-delay estimate, and a risk value;determine a risk output based at least in part on the processed weight representations and inference outputs; andprovide a diagnostic prediction to a user device.

13. The system of claim 12, wherein the processor is configured to identify the feature-importance weights by an ensemble learning model that ranks symptom features by predictive contribution and applies the feature-importance weights to emphasize higher-contribution features.

14. The system of claim 12, wherein determining the risk output includes:sampling a plurality of candidate onset estimates;determining corresponding risk values for each candidate onset estimate of the plurality of candidate onset estimates; andaggregating the corresponding risk values .

15. The system of claim 12, wherein the processor is configured to detect a distribution shift from received symptom data and to trigger recalibration, retraining, or adjustment of decision thresholds responsive to the distribution shift.

16. The system of claim 12, wherein the processor is configured to select or recommend a clinical action from a constrained action set.

17. The system of claim 12, wherein the processor is configured to continuously refine at least one of the feature-importance weights, one or more temporal-encoder parameters, one or more prediction-head parameters for generating inference outputs, and one or more calibration artifacts based on the data entries and one or more verified diagnoses.

18. A method for generating a prediction indicative of a neurodegenerative disease comprising:receiving data entries for a user, each data entry including symptom data and an associated timestamp;analyzing the symptom data to identify feature-importance weights;applying the feature-importance weights to the symptom data to generate weighted feature vectors;encoding a sequence of the weighted feature vectors be processing the sequence with a transformer encoder comprising multi-head self-attention;adjusting the encoding with a delay-aware attention mechanism associated with the sequence based on an inferred disease onset, wherein adjusting includes transforming at least a portion of the associated timestamps into disease-relative timestamps using an onset estimate;generating inference outputs comprising at least one of an onset estimate, a diagnosticdelay estimate, and a risk value;determining a risk output based at least in part on the encoded sequence and the inference outputs; andproviding a diagnostic prediction.

19. The method of claim 18 further comprising:ranking symptom features of the symptom data;emphasizing higher-contribution features by applying corresponding featureimportance weights;sampling a plurality of candidate onset estimates;determining corresponding risk values for the plurality of candidate onset estimates; andaggregating the corresponding risk values to generate the risk output.

20. The method of claim 18 further comprising:continuously refining at least one of the feature-importance weights, one or more encoder parameters, one or more onset parameters, one or more diagnostic-delay parameters, and r one or more risk value parameters based on the symptom data and one or more verified diagnoses.