This invention discloses an aging prediction method,
system, and device based on multimodal
biosignal fusion. The invention constructs customized training paradigms for three modalities: ocular
medical imaging, behavior, and
gait. For the ocular modality, a self-masking-driven three-stage progressive training process is employed, sequentially achieving self-supervised pre-training, domain data fine-tuning, and classification task
adaptation. For the behavior modality,
unsupervised learning is used to extract behavioral
syllable features, establishing a quantitative correlation between behavior and aging degree. For the
gait modality, high-dimensional
gait features are extracted based on spatiotemporal feature quantification and structured statistical aggregation. This invention proposes a hierarchical cross-
modal fusion architecture, sequentially performing single-
modal feature encoding, image modality mean fusion, and cross-
modal feature enhancement, achieving collaborative reasoning between medical images and behavioral / gait representations. This significantly improves the accuracy, robustness, and clinical
interpretability of aging prediction, providing a structured new paradigm for cross-modal fusion in biomedical
multimodal data analysis.