Method, system and device for Parkinson's disease diagnosis based on depth map contrast learning

By using deep graph contrastive learning and the graph neural network framework GRFGraphNet, the problems of sensor failure and cross-device migration in pressure insoles in Parkinson's disease diagnosis were solved, achieving high-precision Parkinson's disease diagnosis and assessment.

CN121565429APending Publication Date: 2026-02-24SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511531464.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies for Parkinson's disease diagnosis using pressure insoles face challenges such as data reliability issues caused by sensor failure and noise disturbances, as well as difficulties in migrating across devices, which affect the robustness and scalability of the model.

Method used

We employ a deep graph contrastive learning approach, extracting frame-level embeddings and temporal features of plantar pressure signals through graph convolutional encoders and Transformer encoders. By combining graph contrastive learning and data augmentation techniques, we construct a graph neural network framework, GRFGraphNet, to improve the model's robustness to sensor failures and cross-device differences.

Benefits of technology

It improves the stability of the model in the face of sensor failure and noise disturbance, enhances cross-device transferability, and enables high-precision diagnosis and assessment of Parkinson's disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565429A_ABST
    Figure CN121565429A_ABST
Patent Text Reader

Abstract

The invention provides a method, system and device for Parkinson's disease diagnosis based on depth map contrast learning, and the method comprises the steps: carrying out the sliding window division of a plantar pressure signal, and obtaining a frame-level plantar pressure signal frame sequence; converting the frame-level plantar pressure signal frame sequence into plantar pressure diagram structure data; the method comprises the following steps: extracting depth feature representation based on plantar pressure diagram structure data, constructing positive and negative samples for comparative learning, performing feature extraction, and inputting Transform and a linear layer to output a classification result. According to the method, a node discarded data enhancement method is introduced, so that embedding representation of model learning has consistency under node disturbance; plantar pressure signals of different pressure insole devices are unified into the same region-level representation through region-level sub-graph data enhancement, the dependence of a model on the specific sensor number and spatial arrangement is weakened, and the transferability of cross-devices is improved; graph contrast learning is introduced into plantar pressure signal feature learning, and robustness of the model to input structure disturbance is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary fields of medical signal processing and artificial intelligence. Specifically, it relates to a method, system and device for Parkinson's disease diagnosis based on deep graph contrastive learning, and in particular a graph neural network framework (GRFGraphNet) with anatomical prior as the core and oriented towards multi-task assessment of PD. Background Technology

[0002] Parkinson's disease (PD) is a chronic and progressive neurological disorder characterized by prominent gait abnormalities, including tremor and postural instability. According to a 2019 report by the World Health Organization (WHO), more than 8.5 million people worldwide suffer from PD. Currently, the clinical diagnosis of PD mainly relies on methods such as the Unified Parkinson's Disease Rating Scale (UPDRS) and functional neuroimaging. However, these methods have significant limitations. The Unified Parkinson's Disease Rating Scale (UPDRS / MDS) is a relevant reference to this disease. The UPDRS (Universal Physical Response System) assesses motor and non-motor symptoms through interviews and physical examinations, covering multiple dimensions including activities of daily living, motor function tests, treatment complications, and pre- and post-treatment status. It is typically assessed by trained neurologists in both "on" and "off" states, quantifying symptom severity and changes over time. However, this method relies on the assessor's experience, exhibiting variability between and within assessors, and is highly subjective; assessments are time-consuming (a complete MDS UPDRS often requires 30–45 minutes), limiting follow-up frequency; it is insufficient in capturing fluctuations in daily life (motor complications, interleaving phenomena), and is affected by the timing of consultations; it has limited sensitivity to certain phenotypes such as postural instability and frozen gait; and its coverage of non-motor symptoms is still incomplete.

[0003] The Hoehn & Yahr (H&Y) staging system classifies Parkinson's disease (PD) into stages 1–5 based on postural instability and bilateral involvement. Clinically, it is used to rapidly stratify disease severity and predict prognosis. However, it has a limited scope, primarily reflecting axial symptoms and balance function, and cannot refine non-motor symptoms, cognitive and psychological burden; it is difficult to reflect subtle changes; there is significant heterogeneity within the same stage; and it is easily affected by comorbidities (such as joint disease and peripheral neuropathy), leading to overstaging or distortion of the staging.

[0004] Neuroimaging (structural and functional imaging: MRI, DaTSCAN / FP-CIT SPECT, PET, etc.) is mainly used clinically to rule out secondary Parkinson's syndrome and observe signs of substantia nigra degeneration through structural MRI. Dopamine transporter imaging (DaT-SPECT / FP-CIT) or PET is used to assess striatal dopaminergic terminal function. When necessary, glucose metabolism PET or receptor-ligand PET is used to characterize brain networks and receptor changes. However, these methods are expensive; some examinations involve ionizing radiation exposure; accessibility and waiting times limit routine clinical application; sensitivity and specificity for early or atypical phenotypes are not perfect, making it difficult to distinguish PD from the early stages of some atypical Parkinson's syndromes (such as MSA, PSP); image interpretation requires specialized experience, and there are inconsistencies in standards between centers; results are affected by medications, comorbidities, and motor status.

[0005] In contrast, pressure signals collected from the soles of the feet using pressure insoles provide a non-invasive, low-cost, and long-term follow-up approach for gait analysis, enabling refined assessment of patients with Parkinson's disease (PD). Pressure insoles deploy multiple pressure sensors at different locations on the sole of the foot to record the dynamic distribution of plantar pressure during movement. The pressure signals from different locations exhibit a coordinated dynamic interaction with the gait phase, containing physiological information related to skeletal muscle control and neural regulation, thus providing a reliable basis for the objective diagnosis and quantitative assessment of PD.

[0006] Plantar pressure data is typically collected using multi-point pressure sensors embedded in insoles. Each sensor channel corresponds to the instantaneous force magnitude and temporal changes in different anatomical regions of the foot (such as the calcaneus, metatarsals, and toes). Based on this type of multi-channel temporal data, data-driven machine learning methods are further introduced into traditional gait parameter analysis for the screening, classification, and severity assessment of Parkinson's disease (PD). The machine learning method first requires feature extraction from the plantar pressure signals, which can be broadly categorized into two types: signal features and gait features.

[0007] The signal characteristics are mainly derived from the original pressure waveform, covering time-domain statistics (mean, peak value, rise / fall time, pulse width, load-unload slope, coefficient of variation), frequency-domain indicators (power spectrum peak and bandwidth, dominant frequency energy distribution), and time-frequency and nonlinear measures (wavelet energy, spectral entropy, sample entropy, Lyapunov exponent, etc.). They mainly reflect the dynamic pattern, rhythmicity, and complexity of plantar force as the gait phase progresses, thus being more sensitive to minor changes in movement sluggishness, gait freeze, and posture control defects.

[0008] Gait characteristics are obtained through event detection and space-time parameter reconstruction, such as stride length, stride frequency, gait cycle and gait symmetry, stance / swing ratio, single / double support time, center of pressure (CoP) trajectory and its path length, velocity and ellipse area, heel-forefoot load transfer speed, left and right foot load distribution and turning parameters, etc. They mainly reflect the overall gait stability, rhythm and load transfer strategy, and have direct indicative significance for postural instability and fall risk.

[0009] The extracted features are input into machine learning models for the diagnosis and assessment of PD, mainly including SVM, RandomForest, Logistic Regression, and Gradient Boosting. However, traditional machine learning methods rely on feature engineering and time-frequency analysis, which not only require specialized knowledge but also introduce high computational complexity.

[0010] Deep learning, with its advantages in spatiotemporal feature representation, has become an important direction in Parkinson's gait analysis, mainly including two types of architectures: Convolutional Neural Networks (CNNs) and Temporal Neural Networks (TNs). CNNs focus on extracting local patterns of pressure signals, such as waveform shapes in the time domain. Temporal Neural Networks (RNNs, TCNs, Transformers, etc.) focus on the dynamic changes and long-term dependencies of pressure signals, and can model the rhythmic patterns of gait phases. In addition, hybrid architectures combine the local representations of CNNs with the dynamic modeling of temporal models. Generally, convolutions are first used to extract multi-scale local features, and then the temporal module integrates long-range dependencies to improve the ability to model PD gait patterns.

[0011] Most existing methods treat pressure signals as a gridded or sequential representation, treating multiple sensors as independent channels, neglecting the topological relationships of plantar sensors in anatomical space and the cooperative dynamics between regions. In real-world applications, pressure insoles inevitably face issues such as sensor failure and noise disturbances. Models lacking robust design are sensitive to single-point failures or local anomalies, easily leading to significant performance degradation. Furthermore, different brands of pressure insoles vary in the number and spatial arrangement of sensors. Models relying on a fixed number of channels and specific arrangement patterns are often difficult to transfer directly between devices, impacting clinical scalability and deployment efficiency.

[0012] In summary, although plantar pressure signals from pressure-sensitive insoles provide a non-invasive, low-cost, and longitudinally trackable objective quantitative method for gait analysis in Parkinson's disease, existing technological paths from traditional machine learning to deep learning still face key bottlenecks in real-world applications. The vulnerability at the sensor level is prominent: long-term use of insoles can lead to single-point or localized sensor failures (dead zones, drift, saturation, intermittent disconnections) and noise disturbances (mechanical shock, baseline drift caused by sweat / temperature, electromagnetic interference). Most models simply treat multiple sensors as independent channels, lacking robust mechanisms for missing / anomalies (such as redundant coding, anomaly detection and repair, dynamic channel reweighting), resulting in single-point failures causing feature distortion and a precipitous performance drop. Features and models are highly sensitive to data quality: time-domain / frequency-domain / time-frequency and nonlinear features are susceptible to noise. The amplified variance and deteriorated stability of the background mean that deep models may "learn" noise patterns as discrimination cues, leading to overfitting and extrapolation failures. In reality, different synchronicity tasks, rhythm fluctuations, and changes in wearing methods further exacerbate this problem. Cross-device migration remains a bottleneck for implementation: different brands of insoles have significant differences in the number, density, and spatial arrangement of sensors, and the channel naming, sampling rate, calibration curves, and measurement range are also inconsistent. Models that rely on fixed channel indices or specific permutation assumptions are difficult to reuse directly and require additional retraining, which increases deployment costs and limits multi-center expansion.

[0013] In summary, data reliability issues caused by sensor failure and noise disturbances, as well as the difficulties in transfer due to differences in cross-device topology, are the most prominent shortcomings of current methods in terms of clinical usability and scalability. There is an urgent need to develop systematic robust and transfer solutions at the model level and learning strategies. Summary of the Invention

[0014] To address the shortcomings of existing technologies, the present invention aims to provide a method, system, and apparatus for Parkinson's disease diagnosis based on depth map contrastive learning.

[0015] The present invention provides a system for Parkinson's disease diagnosis based on deep graph contrastive learning, comprising: a graph convolutional encoder, a Transformer encoder, and a linear layer.

[0016] The graph convolutional encoder extracts frame-level embedding codes, which are then input into the Transformer encoder. The Transformer encoder extracts temporal features and connects them to a linear layer; The linear layer outputs the classification probability.

[0017] Preferably, the graph convolutional encoder adopts a multi-level residual graph convolutional structure, uses two fully connected layers as feature mapping modules, and adopts NT-Xent as the contrastive loss function; By performing stacked graph convolution operations, combining local feature modeling with global information extraction of the spatial topological characteristics of the foot pressure map, graph-level embedding is obtained, and the positive sample distance is minimized by the projection head.

[0018] The Transformer encoder consists of a multi-layered, multi-head self-attention layer and a feedforward neural network.

[0019] The calculation of the self-attention layer includes querying. ,key Sum The weighted relationship.

[0020] The parameters are updated by minimizing a loss function, which uses cross-entropy loss:

[0021] in, Indicates the total number of samples; Indicates the total number of categories; Indicates the first Each sample in category The label on; Indicates the first Each sample was predicted by the model to be of category [class]. The probability of.

[0022] Preferably, the graph convolution operation formula of the graph convolution encoder is:

[0023] in, Represents the adjacency matrix; The degree matrix represents the adjacency matrix; The initial features of the nodes are represented as input; Indicates the node output features; This represents the learnable weight parameters.

[0024] Features are aggregated using the relative weights of neighboring nodes, and the node-level expression formula is as follows:

[0025] in, Represents a node The set of neighbors; Indicates weight; , Let represent the degree matrices of the adjacency matrices of nodes j and i, respectively. Represents the initial characteristics of node j; Represents a node The node output features.

[0026] Introduce residual connections between each network layer:

[0027] in, Indicates the first The input features of the layer; GCN(·) represents the output feature obtained through the convolution operation of the current layer graph.

[0028] The average-max pooling operation is applied to read out the features at each layer, and the feature representations of all layers are concatenated to output the graph embedding encoding.

[0029] According to the present invention, a training method for a system for Parkinson's disease diagnosis based on deep map contrastive learning is provided, for training the system for Parkinson's disease diagnosis based on deep map contrastive learning, comprising: Step A1: Perform node discarding and region-level subgraph structure enhancement on the preprocessed plantar pressure map frame by frame to obtain different but semantically consistent sample pairs; Step A2: Input the sample pairs into the graph convolutional encoder, perform recursive feature aggregation on the graph structure, normalize the neighborhood features with relative weights, and introduce residual connections between layers. Step A3: By maximizing the similarity of positive sample embeddings and minimizing the difference of negative sample embeddings, complete the graph contrastive learning of the graph convolutional encoder and save the graph convolutional encoder. Step A4: The preprocessed plantar pressure map is segmented into a plantar pressure frame sequence, input into the image convolutional encoder to obtain frame-level embedding encoding, stacked in chronological order and input into the Transformer encoder to capture temporal features and then input into a linear layer to complete the downstream fine-tuning task.

[0030] Preferably, in step A2, random sampling is used to obtain... A plantar pressure map, corresponding to... Each node discards the augmented graph and Each region-level subgraph augmentation map represents a node discarded augmentation map representation, which is treated as a positive sample pair with the region-level subgraph augmentation map representation from the same frame. Other... An augmented graph is represented as a negative sample pair.

[0031] In step A3, the cosine similarity function is used as the similarity measure:

[0032] in, They represent the first The latent space vectors of the point-drop augmentation map and the region-level submap augmentation map corresponding to the plantar pressure map.

[0033] No. The NT-Xent error for each plantar pressure map is:

[0034] in, Indicates temperature parameter; , These represent different plantar pressure chart ordinal numbers.

[0035] According to the present invention, a method for diagnosing Parkinson's disease based on depth map contrastive learning is provided, wherein the system for diagnosing Parkinson's disease based on depth map contrastive learning is used for diagnosis, including: Step S1: Perform sliding window segmentation on the plantar pressure signal to obtain a frame-level plantar pressure frame sequence; Step S2: Convert the frame-level plantar pressure frame sequence into a plantar pressure map; Step S3: Extract features from the gait time sequence in the plantar pressure map, generate deep feature representations, construct positive and negative sample pairs, train the graph convolutional encoder and save it; Step S4: The graph convolutional encoder extracts the graph embedding code, inputs it into the Transformer encoder and the linear layer, and outputs the classification probability.

[0036] Preferably, in step S1, the length of the plantar pressure signal is normalized to a fixed value, and the dataset S is represented as:

[0037] in, This indicates the number of plantar pressure signals in the dataset; Indicates the number of pressure sensor points; Represent the space of real numbers; This indicates a plantar pressure signal record; This indicates a fixed value.

[0038] Each sample Divided into Non-overlapping plantar pressure frames, each frame containing 10 data points constitute a time series:

[0039] in, , indicating from The number of complete plantar pressure signal frames extracted; Indicates the first One foot pressure frame.

[0040] In step S2, based on the plantar anatomical structure and the spatial distribution of each pressure sensor, the sensor is divided into three regions: forefoot, midfoot, and hindfoot. A connection method is constructed with full connectivity within the region and sparse connectivity between regions to form an adjacency matrix.

[0041] Each plantar pressure frame in the frame-level plantar pressure frame sequence is defined as a plantar pressure map: G=(V,E,A),|V|=n Where V represents the set of nodes; n represents the number of nodes, with each node corresponding to one sensor; E represents the set of edges; A represents the adjacency matrix of the plantar pressure map G.

[0042] Preferably, in step S3, a graph convolutional encoder is used to extract features from the gait time sequence to generate deep feature representations related to Parkinson's disease.

[0043] A graph contrast learning mechanism is introduced to augment the plantar pressure map data, resulting in an enhanced map. :

[0044] Wherein, G represents the plantar pressure map; This represents a predefined data augmentation distribution.

[0045] Augmented graphs will be created by performing data augmentation through node dropping and region-level subgraphs, respectively. , As sample pairs for graph contrastive learning; Perform graph comparison training and save the trained graph convolutional encoder.

[0046] Preferably, in step S4, the plantar pressure signal is divided into a plantar pressure frame sequence in a 1s window, and input into a graph convolutional encoder to obtain frame-level embedding encoding; The frame-level embedding codes are stacked in chronological order to form gait sequence features, which are then input into the Transformer encoder to capture phase transitions, rhythms, and long-range dependencies across frames, thereby obtaining global time series features. The global time series features are passed through a linear layer to output the classification probability.

[0047] According to the present invention, an apparatus includes the aforementioned system for diagnosing Parkinson's disease based on depth map contrastive learning, wherein the system for diagnosing Parkinson's disease based on depth map contrastive learning identifies Parkinson's disease based on input data and outputs classification results.

[0048] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention addresses the issues of sensor failure and noise disturbance that often arise in the practical use of pressure insoles that rely on data collection. It introduces a node discarding data augmentation method to ensure that the embedded representation learned by the model remains consistent under node perturbation.

[0049] 2. This invention addresses the difficulty of cross-device migration caused by differences in the number and arrangement of sensors in pressure insoles from different brands. By using region-level subgraph data augmentation, the plantar pressure signals of different pressure insole devices are unified into the same region-level representation, thereby reducing the model's dependence on the specific number and spatial arrangement of sensors and improving cross-device portability.

[0050] 3. This invention introduces graph contrast learning into plantar pressure signal feature learning, and enhances the robustness of the model to input structural perturbations through pre-training and fine-tuning. Attached Figure Description

[0051] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A schematic diagram of the system architecture for Parkinson's disease diagnosis based on deep map contrastive learning; Figure 2 This is a schematic diagram of a foot pressure time slice according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the plantar pressure map construction method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a data augmentation method according to an embodiment of the present invention. Detailed Implementation

[0052] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0053] Foot pressure signals from pressure-sensor insoles can objectively reflect the gait mechanics and rhythm abnormalities of Parkinson's disease (PD) patients. However, in real-world data acquisition, channel loss due to sensor damage, failure, or noise perturbation poses a significant challenge to the model's robustness, limiting its robustness and generalization ability in real-world scenarios. Furthermore, pressure-sensor insoles from different brands and devices exhibit significant differences in the number and arrangement of sensors, preventing fixed-grid convolutional neural networks from generalizing across devices and further exacerbating the difficulty of cross-device transfer.

[0054] Therefore, this invention provides a training method for a system of Parkinson's disease diagnosis based on deep map contrastive learning, so as to... Figure 1 For example, including: Data preprocessing includes: temporal segmentation of plantar pressure signals and construction of plantar pressure maps based on anatomical priors. The plantar pressure signals obtained from the original gait acquisition system are preprocessed, dividing the continuous signal into multiple 1-second plantar pressure signal frames. Based on the plantar anatomy and sensor layout, the plantar pressure frames are further represented as plantar pressure map structure data with node and edge relationships, where nodes represent force-bearing areas and edges reflect the correlation weights between areas.

[0055] Step S1: Perform sliding window segmentation on the plantar pressure signal to obtain a frame-level plantar pressure frame sequence; Specifically, the raw plantar pressure signal is preprocessed to construct a frame-level map incorporating anatomical priors as a unified representation.

[0056] Continuous plantar pressure signals are divided into non-overlapping frames in 1-second windows, each frame representing a gait phase. Each frame is then used to construct an anatomically based graph as input for subsequent models. The graph consists of three main regions: forefoot, midfoot, and hindfoot. Fully connected regions are used within each region to represent local synergistic forces, while sparse, physiologically directional connections are used between regions to characterize the propulsion pathway (connecting only the nearest node pairs between adjacent regions, with the direction from back to front). Each sensor is a graph node, and the node's characteristic is the pressure time series of the corresponding sensor within that window.

[0057] by Figure 2 For example, plantar pressure signal time sequence segmentation: In order to model the temporal dynamic characteristics of plantar pressure signal and uniformly process samples of different signal lengths, the original signal is divided into a structured time slice sequence.

[0058] The plantar pressure signals of all samples were normalized to a fixed length using a padding operation. This is to facilitate unified processing. The dataset can then be represented as:

[0059] in, This indicates the number of plantar pressure signal records in the dataset. The number of pressure sensor points. Represent the real number space, each sample This represents a plantar pressure signal record.

[0060] Based on the preset time window length Each sample Divided into Non-overlapping frames, each containing Each data point corresponds to a specific phase in the gait cycle, resulting in a time series composed of plantar pressure frames:

[0061] in, Representative from The number of complete plantar pressure signal frames extracted. Representing the One foot pressure frame.

[0062] This segmentation not only extracts local features of the signal at each gait phase but also preserves the overall periodic dynamic features of the plantar pressure signal across multiple slice sequences. Furthermore, a reasonable window size definition is crucial. It can effectively balance the preservation of signal details with computational overhead. After time slicing, each generated plantar pressure frame will be input into the subsequent plantar pressure map construction module to characterize the spatial topological relationship between plantar pressure sensors.

[0063] Step S2: Based on the plantar anatomical structure and the spatial distribution of each pressure sensor, the frame-level plantar pressure frame sequence is converted into structural data of the plantar pressure map that incorporates anatomical knowledge. by Figure 3 For example, the plantar pressure graph construction module uses a graph structure to represent plantar pressure signals.

[0064] The plantar pressure signal is divided into multiple 1-second plantar pressure signal frames. Each frame is used to construct a directed graph that incorporates anatomical knowledge, i.e., a plantar pressure map. An original signal is 100 seconds long, which is divided into 100 1-second frames, each of which forms a directed graph.

[0065] Specifically, each plantar pressure frame is defined as a directed graph G=(V,E,A), where V represents the set of vertices (nodes), each node corresponds to a sensor, the number of nodes in the plantar pressure graph is n, |V|=n, E represents the set of edges, and A represents the adjacency matrix of the plantar pressure graph G, which is used to quantize the weight of the edges.

[0066] To construct a reasonable adjacency matrix, based on the anatomical structure of the foot, the sensor is divided into three regions: forefoot, midfoot, and hindfoot. A connection method is constructed that is fully connected within the region and sparsely connected between regions to simultaneously capture the local and global dynamic characteristics of the foot.

[0067] The fully connected region means that any two sensor nodes within the same region are connected, based on the division of labor among the functional areas of the foot during walking: the forefoot is responsible for propulsion output, the midfoot is responsible for stability regulation, and the hindfoot provides ground support. Considering that plantar pressure signals at different locations within the same region are usually similar in temporal morphology, but have cooperative and complementary relationships in amplitude and temporal phase, the fully connected region can capture the dynamic correlation and cooperative patterns between local sensors, ensuring the complete representation of features within the region.

[0068] The sparse connectivity between regions connects only the pair of sensors with the closest physical location in two regions with directed edges, and sets the direction according to the force transmission sequence (hind foot → midfoot → forefoot). This not only models key interaction information between regions, such as the dynamic coordination mode between the forefoot and midfoot during walking, but also effectively reduces the computational burden. The weight of each edge is the reciprocal of the physical distance between the two connected sensors.

[0069] Step S3: Based on the stress map structure data, a graph neural network is used to extract features from the gait time sequence to generate a deep feature representation related to Parkinson's disease.

[0070] Step S4: Introduce a graph contrastive learning mechanism. Through various graph structure data augmentation methods such as node dropping and region-level subgraph, construct positive and negative samples for contrastive learning. This improves the model's robustness to structural perturbations such as missing nodes and arrangement differences, resulting in a trained graph convolutional encoder.

[0071] For the obtained plantar pressure maps, a graph contrastive learning pre-training phase is performed. Specifically, a deep graph contrastive learning model is constructed, or in other words, a graph neural network framework GRFGraphNet centered on anatomical priors and geared towards multi-task evaluation in plantar pressure (PD). GRFGraphNet is used to learn robust plantar pressure signal representations, thereby improving the ability to adapt to sensor deficiencies and cross-device differences. After contrastive learning, the pre-trained encoder is fine-tuned together with the downstream task classifier for specific targets.

[0072] Specifically, structure-aware self-supervised graph contrastive learning is performed at the frame level to obtain an encoder robust to sensor layout differences and channel missingness.

[0073] After obtaining the frame-level maps, robust feature representations to structural perturbations are learned under unlabeled conditions. To address two common defects in real-world data acquisition, two types of structure-aware enhancement are constructed for each frame-level map. Figure 4 For example, one is node drop: randomly remove a node and its associated edge from the plantar pressure map to simulate sensor failure, so that the learned representation is consistent under node perturbation; the other is subgraph: according to the partition of the plantar pressure map, average pooling is performed on all nodes in each anatomical region (forefoot, midfoot, hindfoot) to obtain three regional nodes, and two unidirectional edges across regions are retained.

[0074] Each node's features represent the overall pressure pattern of its respective region. Two directed edges are constructed between these three region-level nodes based on the prior knowledge of force transmission in gait: hindfoot to midfoot, and midfoot to forefoot. The prior for this enhancement is that within the same anatomical region, the region-level nodes obtained by averaging and pooling sensor signals can represent the gait pattern of that region, preserving the main transmission relationships between regions with minimal structural complexity.

[0075] Unifying multi-sensor inputs from different devices into the same region-level representation reduces the model's dependence on the specific number and spatial arrangement of sensors, thereby improving cross-device portability.

[0076] The resulting regional directed graph reduces the dependence on specific sensor layouts; that is, foot pressure insoles with different numbers and arrangements of sensors can all be uniformly mapped to a graph representation consisting of three nodes, thereby improving the generalization ability across device inputs. Two types of enhancements are defined as "homogeneous cross-enhancement contrast," where the two enhanced views within the same frame are positive sample pairs, and the remaining samples within the batch are negative samples.

[0077] For the obtained preprocessed plantar pressure map data, targeted data augmentation was first performed using strategies such as node dropping and region-level subgraphs to obtain different but semantically consistent sample pairs. This is a further structural optimization of the original preprocessed data, enabling the model to better capture plantar spatial relationships and gait features.

[0078] In more preferred examples, data augmentation is performed on plantar pressure maps: contrastive learning enables representation learning by maximizing feature consistency across different augmented views, where representation invariance is obtained through augmented views relevant to the data or task. Data augmentation for a graph can be represented as: given a plantar pressure map... Enhanced image It can be represented as ,in, This represents a predefined data augmentation distribution.

[0079] Given plantar pressure map After data augmentation, two related augmented graphs were obtained. , As positive sample pairs in comparative studies.

[0080] A multi-level residual graph convolutional (MR-GCN) is constructed as the encoder. The enhanced samples are input into the multi-level residual graph convolutional encoder, which recursively aggregates features of the graph structure, normalizes neighborhood features with relative weights, and introduces residual connections between layers to enhance the stability of feature representation.

[0081] MR-GCN further explores the spatial topological characteristics of plantar pressure data by combining local feature modeling and global information extraction through stacked graph convolution operations. Two views are encoded using multi-level residual graph convolution with shared weights to obtain graph-level embeddings, which are then minimized by a projection head to minimize the positive sample distance.

[0082] Using the enhanced graph data described above, a graph convolutional encoder is trained during the contrastive learning phase. The encoder employs a multi-level residual graph convolutional structure, relying on which the two enhanced views are compared to obtain a representation robust to graph structure perturbations (i.e., obtaining a representation similar to the trained graph encoder).

[0083] By maximizing the similarity of positive sample embeddings and minimizing the dissimilarity of negative sample embeddings, adaptive learning of unsupervised graph features is achieved, thereby completing the contrastive learning pre-training of the encoder. Through this process, a pre-trained graph convolutional encoder model with robust gait representation capabilities can be obtained.

[0084] Specifically, the graph convolutional encoder: To characterize the dynamic correlation features of plantar pressure maps in different neighborhood ranges, a graph convolution with multiple residual connections is used as the encoder for embedding the plantar pressure map. The basic operation formula of graph convolution is shown below:

[0085] in, It is an adjacency matrix; Let be the degree matrix of the adjacency matrix, and its diagonal elements be Input features This represents the initial characteristics of the node, i.e., the plantar pressure signal. These are learnable weight parameters.

[0086] Normalization aggregates features based on the relative weights of neighboring nodes, facilitating standardized feature propagation and learning. Its node-level expression is as follows:

[0087] in, Represents a node The neighbor set, weight It is determined by the physical distance between the sensors (taken as the reciprocal of the distance). , These are the degree matrices of the adjacency matrices of different nodes; Represents the initial characteristics of node j; Represents a node The node output features.

[0088] Meanwhile, to effectively alleviate the potential feature degradation and gradient vanishing problems in deep models, residual connections are introduced between each network layer to achieve direct fusion of cross-layer features. The operation of residual connections can be expressed by the following formula:

[0089] in, It is the first The input features of the layer, and GCN(·) represents the output features obtained by convolution operation on the current layer graph.

[0090] To integrate the output information of multi-layer graph convolution, average-max pooling is applied to read out the features in each layer, and the feature representations of all layers are concatenated to form the final output. This captures the correlation of nodes in local regions and the interaction characteristics of nodes across regions in plantar pressure data, and comprehensively reflects the spatial and dynamic relationships between plantar sensors.

[0091] By extracting the global average features and local maximum features from each layer's output and concatenating them, the model can more meticulously represent the dynamic change patterns and key regional characteristics of the foot. The fusion of multi-layer features and residual connections effectively enhance feature representation capabilities, reduce information loss and overfitting risks in deep models, and provide more accurate and reliable feature inputs for gait analysis and abnormal pattern detection.

[0092] During the graph contrast learning phase, a single plantar pressure frame is used as input, rather than a complete sequence of plantar pressure frames.

[0093] From the enhanced graph using a graph convolution encoder and Obtain the preliminary representation vector and Then, the projection head (a feature mapping module composed of fully connected layers) maps to another latent space to obtain... and This is used to calculate the contrast error. GRFGraphNet uses two fully connected layers as the projection head, and the contrast loss function is NT-Xent.

[0094] In more preferred examples, during graph contrastive learning, i.e., the pre-training process, random sampling is performed from the dataset to obtain data containing... A plantar pressure map, corresponding to... Each node discards the augmented graph and Each region-level subgraph augmentation map is used. The representation of the node-dropped augmentation map and the representation of the region-level subgraph augmentation map within the same graph are considered positive sample pairs, while the corresponding representation is used with other regions. Each augmented image is represented as a negative sample pair. A cosine similarity function is used as the similarity measure.

[0095] in, They represent the first The latent space vectors of the point-drop augmentation map and the region-level submap augmentation map corresponding to the plantar pressure map.

[0096] No. The NT-Xent error of a plantar pressure map can be defined as:

[0097] in, The temperature parameter is represented (set to 0.2). This operation essentially maximizes the mutual information between the representations of the two types of augmented maps.

[0098] The encoder performs recursive feature aggregation on the graph structure, normalizes and aggregates features by using the relative weights of neighboring nodes, and introduces residual connections between each network layer to enhance feature representation and stabilize deep representation learning.

[0099] In the Parkinson's disease diagnosis task, during the construction of positive and negative sample pairs and the optimization of the embedding representation, the graph embedding encoding is first obtained by relying on the trained multi-level residual graph convolution. Then, the graph embedding encoding is input into the Transformer encoder and the fully connected layer for classification. The graph convolution encoder works together with the Transformer encoder and the linear layer to achieve adaptive learning of unsupervised graph features by maximizing the similarity of positive samples and minimizing the difference of negative samples, thus realizing robust gait embedding representation.

[0100] Downstream task fine-tuning: Supervised fine-tuning at the sequence level to capture phase transition and rhythm features.

[0101] After contrastive learning, GRFGraphNet is fine-tuned for specific downstream tasks. In this fine-tuning phase, the input is a downstream labeled dataset (Parkinson's disease diagnosis / staging task), specifically the frame-level plantar pressure data map obtained in step S1. The complete record is then further divided into plantar pressure frame sequences within a 1-second window. Each frame first obtains a frame-level embedding through a pre-trained MR-GCN; these embeddings are then stacked chronologically and input into a Transformer encoder to capture temporal features, namely phase transitions, rhythms, and long-range dependencies across frames; finally, a linear layer produces a task-specific output.

[0102] Specifically, the input is a complete sequence of plantar pressure frames obtained from slices of the subject's original records. After feature extraction by a graph convolutional encoder, the plantar pressure maps of all frames are concatenated in chronological order to form a unified time series representation, which is then input into a Transformer encoder to model the dynamic dependencies and gait transition patterns between time steps.

[0103] The Transformer encoder consists of multiple multi-head self-attention layers and a feedforward neural network. Its core is the multi-head self-attention mechanism, which can learn the global dependencies of time series data. The computation of the self-attention layer includes querying... ,key Sum The weighted relationship is given by the following formula:

[0104] in, Scaling factor , , Each feature is obtained from the input sequence features through linear transformation. By dynamically learning the weighted relationships between different time steps, the global dependency of the time series is modeled. The global time series features extracted from the Transformer encoder are input into the classification module, and through a multi-layer fully connected structure, the final output is the probability distribution of the classification.

[0105] During model training, parameters are updated by minimizing a loss function. The loss function used is cross-entropy loss.

[0106] in, Represents the total number of samples. Represents the total number of categories. Representing the Each sample in category The label on it (one-hot encoded, 1 for the correct class, 0 for the rest). Representing the Each sample was predicted by the model to be of category [class]. The probability of.

[0107] The effectiveness of GRFGraphNet was validated on four downstream tasks: health-PD binary classification, Hoehn-Yahr staging, gait phase recognition (standing / walking / turning), and frozen gait (FoG) detection, covering key aspects such as screening, staging, and intervention triggering.

[0108] Step S5: Use the graph convolutional encoder optimized by contrastive learning to extract features from the input samples, and input them into the Transformer and fully connected layers to achieve automatic diagnosis of Parkinson's disease.

[0109] The fully connected layer includes a linear layer and an activation function.

[0110] Specifically, the frame-level plantar pressure data obtained in step S1 is input into a pre-trained graph convolutional encoder to extract frame-level embedding features. These frame-level embeddings are stacked in chronological order to form gait sequence features, which are then input into a Transformer encoder to capture temporal dependencies and cross-frame dynamic patterns. The temporal features extracted by the Transformer encoder are then classified and output through a linear layer, and the model ultimately provides a diagnosis of Parkinson's disease.

[0111] By integrating anatomically aligned frame-level graph representation, application-oriented self-supervised contrastive learning, and temporal modeling in the fine-tuning stage, a technical path for integrated assessment of Parkinson's disease (PD) was formed. The entire process forms an end-to-end model training and inference flow from plantar pressure data to disease diagnosis results. The pre-training mechanism of contrastive learning significantly improves the model's robustness to node missingness, permutation differences, and sampling perturbations, achieving high-precision identification of Parkinson's disease.

[0112] In further optimized examples, to evaluate the applicability of GRFGraphNet for Parkinson's disease (PD) diagnosis under different acquisition conditions, subjects were independently trained and tested on the PhysioNet and WearGaitPD datasets. Differences in sensor density and task paradigms between the two datasets resulted in different signal-to-noise ratios and gait rhythm distributions, thus rigorously testing its cross-device and cross-scenario applicability. To ensure objective evaluation, each dataset was randomly divided into a 90% training set and a 10% independent test set at the subject level. The model was trained and fine-tuned using the entire training set, and the test set was used to evaluate model performance.

[0113] On the PhysioNet dataset, GRFGraphNet achieved an accuracy of 0.935 and a recall of 1.00 for Parkinson's disease samples, with the main errors stemming from a small number of healthy controls being classified as positive. Considering the imbalanced distribution of classes (30% healthy controls and 70% Parkinson's), the results demonstrate high specificity while maintaining high sensitivity. Furthermore, GRFGraphNet exhibits a high AUC, indicating good ranking ability across different thresholds. Comparisons with existing methods show that GRFGraphNet significantly outperforms all other approaches. Compared to the suboptimal 1D Convnet, its ACC and MCC are improved by 3.5% and 10.9%, respectively, demonstrating GRFGraphNet's effectiveness and robustness in Parkinson's disease diagnosis. In contrast, methods relying on phase space reconstruction (PSR) and machine learning generally achieved lower results, with the highest PSR ELA achieving an MCC of only 0.381.

[0114] Table 1. Comparison results with other existing methods on the PhysioNet dataset:

[0115] Similar to its results on the PhysioNet dataset, GRFGraphNet also achieved no missed detections of Parkinson's patients on the WearGaitPD dataset, achieving an AUC of 0.857. In comparisons with existing methods, GRFGraphNet outperformed all other metrics except Precision. Specifically, GRFGraphNet's ACC was 0.833 and MCC was 0.683, representing improvements of 24.9% and 39.9% respectively compared to the second-best methods. Notably, the method with the highest Precision had a maximum Recall of only 0.429, indicating that GRFGraphNet has a superior balance between sensitivity and overall discriminative power. Furthermore, all methods showed a performance decline compared to the PhysioNet dataset, with machine learning methods even predicting the opposite direction to reality (MCC < 0). In contrast, GRFGraphNet exhibited a smaller performance decline under the same conditions, demonstrating greater applicability. Overall, GRFGraphNet models gait based on graph structures and achieves leading performance under different acquisition devices and data conditions, demonstrating good cross-device and cross-scenario applicability.

[0116] Table 2. Comparison results with other existing methods on the WearGaitPD dataset:

[0117] In further preferred examples, to evaluate the reliability of GRFGraphNet in practical applications, two perturbation testing protocols were designed: a single-sensor node random loss test and a noise perturbation test. The former randomly deletes one node from each test sample and correspondingly removes its connecting edges when constructing the plantar pressure anatomy map; the latter replaces the pressure signal of each sensor node with Gaussian noise sequentially, while leaving other nodes unchanged. Using Parkinson's disease diagnosis as the target task, without any fine-tuning for the perturbation test, the GRFGraphNet trained on the original signal was directly used for testing to evaluate its reliability and robustness under sensor damage and noise interference scenarios.

[0118] The random loss of a single sensor node significantly alters the pressure trajectory in key gait phases. Taking the random removal of a node from the left foot as an example, the peak value of the perturbed pressure signal is significantly reduced compared to the original signal, and the peak phase shifts. GRFGraphNet's metrics on both datasets remain consistent with the original data, showing no decline. Specifically, the ACC is 0.935 on PhysioNet and 0.833 on WearGaitPD.

[0119] Because GRFGraphNet employs contrastive learning with node dropout augmentation graphs during the pre-training phase, the model's representation remains consistent under node perturbations, ensuring stable performance even in the event of sensor failures.

[0120] Table 3. Results of random single-sensor node loss experiments on the PhysioNet dataset.

[0121] Table 4. Results of random single-sensor node loss experiments on the WearGaitPD dataset.

[0122] In graph learning, noisy features can contaminate neighboring nodes and amplify errors during message passing and information aggregation, thus weakening the discriminative power of the representation. Taking replacing a node on the left foot with noise as an example, the pressure signal waveform exhibits numerous irregular spikes. From a frequency domain perspective, the energy of the signal above 10 Hz is significantly increased after noise perturbation. Given that the plantar pressure signal generated by walking is mainly concentrated below 10 Hz, these newly added high-frequency components pose a significant challenge to the model's robust representation and feature aggregation. In the noise perturbation test, statistics were performed on each sensor node separately. The results show that although noise perturbation causes varying degrees of performance degradation, GRFGraphNet still maintains reliable classification performance on both datasets. Specifically, the AUC on the PhysioNet dataset is 0.859, a decrease of only 8.9% compared to the baseline; the AUC on WearGaitPD is 0.859, a decrease of only 2.2% compared to the baseline.

[0123] Table 5. Results of noise perturbation experiments on the PhysioNet digital dataset

[0124] Table 6. Results of noise perturbation experiments on the WearGaitPD digital set

[0125] In more advanced examples, GRFGraphNet unifies the input with "region-level subgraphs" that incorporate anatomical knowledge and reduces its dependence on sensor layout through graph contrastive learning. A Parkinson's disease diagnosis model trained on the PhysioNet dataset was tested on the WearGaitPD dataset to evaluate its cross-device transferability. The evaluation results show that, without fine-tuning (Test), the model achieves an ACC of 0.583 and an MCC of 0.378. After one round of fine-tuning using only 10 samples, the model's ACC improved to 0.667 and MCC to 0.488, outperforming other deep learning methods trained directly on the WearGaitPD dataset. GRFGraphNet demonstrates strong cross-device transferability, significantly improving performance with only a small number of target domain samples.

[0126] Table 7. Results of cross-device migration experiment

[0127] A system for Parkinson's disease diagnosis based on deep graph contrastive learning according to the present invention includes: a graph convolutional encoder, a Transformer encoder, and a linear layer; The graph convolutional encoder extracts frame-level embedding codes, which are then input into the Transformer encoder. The Transformer encoder extracts temporal features and connects them to a linear layer; The linear layer outputs the classification probability.

[0128] According to the present invention, a training method for a system for Parkinson's disease diagnosis based on deep map contrastive learning is provided, for training the system for Parkinson's disease diagnosis based on deep map contrastive learning, comprising: Step A1: Perform node discarding and region-level subgraph structure enhancement on the preprocessed plantar pressure map frame by frame to obtain different but semantically consistent sample pairs; Step A2: Input the sample pairs into the graph convolutional encoder, perform recursive feature aggregation on the graph structure, normalize the neighborhood features with relative weights, and introduce residual connections between layers. Step A3: By maximizing the similarity of positive sample embeddings and minimizing the difference of negative sample embeddings, complete the graph contrastive learning of the graph convolutional encoder and save the graph convolutional encoder. Step A4: The preprocessed plantar pressure map is segmented into a plantar pressure frame sequence, input into the image convolutional encoder to obtain frame-level embedding encoding, stacked in chronological order and input into the Transformer encoder to capture temporal features and then input into a linear layer to complete the downstream fine-tuning task.

[0129] According to the present invention, an apparatus includes the aforementioned system for diagnosing Parkinson's disease based on depth map contrastive learning, wherein the system for diagnosing Parkinson's disease based on depth map contrastive learning identifies Parkinson's disease based on input data and outputs classification results.

[0130] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function as logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0131] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A system for Parkinson's disease diagnosis based on deep map contrastive learning, characterized in that, include: Graph convolutional encoder, Transformer encoder, and linear layer; The graph convolutional encoder extracts frame-level embedding codes, which are then input into the Transformer encoder. The Transformer encoder extracts temporal features and connects them to a linear layer; The linear layer outputs the classification probability.

2. The system for Parkinson's disease diagnosis based on depth map contrastive learning according to claim 1, characterized in that, The graph convolutional encoder adopts a multi-level residual graph convolutional structure, uses two fully connected layers as the feature mapping module, and uses NT-Xent as the contrastive loss function. By performing stacked graph convolution operations, combining local feature modeling and global information extraction of the spatial topological characteristics of the foot pressure map, graph-level embedding is obtained, and the positive sample distance is minimized by the projection head. The Transformer encoder consists of a multi-layered, multi-head self-attention layer and a feedforward neural network; The calculation of the self-attention layer includes querying. ,key Sum Weighted relationships; The parameters are updated by minimizing a loss function, which uses cross-entropy loss: in, Indicates the total number of samples; Indicates the total number of categories; Indicates the first Each sample in category The label on; Indicates the first Each sample was predicted by the model to be of category [class]. The probability of.

3. The system for Parkinson's disease diagnosis based on deep map contrastive learning according to claim 1, characterized in that, The graph convolution operation formula of the graph convolution encoder is: in, Represents the adjacency matrix; The degree matrix represents the adjacency matrix; The initial features of the nodes are represented as input; Indicates the node output features; Indicates the learnable weight parameters; Features are aggregated using the relative weights of neighboring nodes, and the node-level expression formula is as follows: in, Represents a node The set of neighbors; Indicates weight; , Let represent the degree matrices of the adjacency matrices of nodes j and i, respectively. Represents the initial characteristics of node j; Represents a node The node output features; Introduce residual connections between each network layer: in, Indicates the first The input features of the layer; GCN(·) represents the output feature obtained through the convolution operation of the current layer graph; The average-max pooling operation is applied to read out the features at each layer, and the feature representations of all layers are concatenated to output the graph embedding encoding.

4. A training method for a system for Parkinson's disease diagnosis based on deep map contrastive learning, used to train the system for Parkinson's disease diagnosis based on deep map contrastive learning as described in any one of claims 1-3, characterized in that, include: Step A1: Perform node discarding and region-level subgraph structure enhancement on the preprocessed plantar pressure map frame by frame to obtain different but semantically consistent sample pairs; Step A2: Input the sample pairs into the graph convolutional encoder, perform recursive feature aggregation on the graph structure, normalize the neighborhood features with relative weights, and introduce residual connections between layers. Step A3: By maximizing the similarity of positive sample embeddings and minimizing the difference of negative sample embeddings, complete the graph contrastive learning of the graph convolutional encoder and save the graph convolutional encoder. Step A4: The preprocessed plantar pressure map is segmented into a plantar pressure frame sequence, input into the image convolutional encoder to obtain frame-level embedding encoding, stacked in chronological order and input into the Transformer encoder to capture temporal features and then input into a linear layer to complete the downstream fine-tuning task.

5. The training method for the system of Parkinson's disease diagnosis based on deep map contrastive learning according to claim 4, characterized in that, In step A2, random sampling is obtained. A plantar pressure map, corresponding to... Each node discards the augmented graph and Each region-level subgraph augmentation map represents a node discarded augmentation map representation, which is treated as a positive sample pair with the region-level subgraph augmentation map representation from the same frame. Other... Each augmentation graph is represented as a negative sample pair; In step A3, the cosine similarity function is used as the similarity measure: in, They represent the first The latent space vectors of the point-drop augmentation map and the region-level submap augmentation map corresponding to the plantar pressure map; No. The NT-Xent error for each plantar pressure map is: in, Indicates temperature parameter; , These represent different plantar pressure chart ordinal numbers.

6. A method for diagnosing Parkinson's disease based on deep map contrastive learning, comprising using the system for diagnosing Parkinson's disease based on deep map contrastive learning as described in any one of claims 1-3, characterized in that, include: Step S1: Perform sliding window segmentation on the plantar pressure signal to obtain a frame-level plantar pressure frame sequence; Step S2: Convert the frame-level plantar pressure frame sequence into a plantar pressure map; Step S3: Extract features from the gait time sequence in the plantar pressure map, generate deep feature representations, construct positive and negative sample pairs, train the graph convolutional encoder and save it; Step S4: The graph convolutional encoder extracts the graph embedding code, inputs it into the Transformer encoder and the linear layer, and outputs the classification probability.

7. The method for Parkinson's disease diagnosis based on depth map contrastive learning according to claim 6, characterized in that, In step S1, the length of the plantar pressure signal is normalized to a fixed value, and the dataset S is represented as follows: in, This indicates the number of plantar pressure signals in the dataset; Indicates the number of pressure sensor points; Represent the space of real numbers; This indicates a plantar pressure signal record; Indicates a fixed value; Each sample Divided into Non-overlapping plantar pressure frames, each frame containing 10 data points constitute a time series: in, , indicating from The number of complete plantar pressure signal frames extracted; Indicates the first One plantar pressure frame; In step S2, based on the plantar anatomical structure and the spatial distribution of each pressure sensor, the sensor is divided into three regions: forefoot, midfoot, and hindfoot. A connection method of fully connected regions and sparsely connected regions is constructed to form an adjacency matrix. Each plantar pressure frame in the frame-level plantar pressure frame sequence is defined as a plantar pressure map: G=(V,E,A),|V|=n Where V represents the set of nodes; n represents the number of nodes, with each node corresponding to one sensor; E represents the set of edges; A represents the adjacency matrix of the plantar pressure map G.

8. The method for Parkinson's disease diagnosis based on depth map contrastive learning according to claim 6, characterized in that, In step S3, a graph convolutional encoder is used to extract features from the gait time sequence to generate deep feature representations related to Parkinson's disease. A graph contrast learning mechanism is introduced to augment the plantar pressure map data, resulting in an enhanced map. : Wherein, G represents the plantar pressure map; This represents a predefined data augmentation distribution; Augmented graphs will be created by performing data augmentation through node dropping and region-level subgraphs, respectively. , As sample pairs for graph contrastive learning; Perform graph comparison training and save the trained graph convolutional encoder.

9. The method for Parkinson's disease diagnosis based on depth map contrastive learning according to claim 6, characterized in that, In step S4, the plantar pressure signal is divided into a plantar pressure frame sequence in a 1s window and input into a graph convolutional encoder to obtain frame-level embedding encoding. The frame-level embedding codes are stacked in chronological order to form gait sequence features, which are then input into the Transformer encoder to capture phase transitions, rhythms, and long-range dependencies across frames, thereby obtaining global time series features. The global time series features are passed through a linear layer to output the classification probability.

10. An apparatus comprising the system for Parkinson's disease diagnosis based on depth map contrastive learning as described in any one of claims 1-3, characterized in that, The system for Parkinson's disease diagnosis based on deep map contrastive learning identifies Parkinson's disease based on input data and outputs classification results.