Distributed rehabilitation data analysis method based on federal learning

By employing federated learning and adversarial data generation technologies, the contradiction between privacy protection and collaborative efficiency in distributed rehabilitation data analysis is resolved. This enables the comprehensiveness and accuracy of rare disease case feature representations to be improved without compromising patient privacy, and supports comprehensive medical data analysis through secure collaboration among multiple institutions.

CN121768631AActive Publication Date: 2026-03-31中国人民解放军总医院第八医学中心

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In a distributed environment, rehabilitation data analysis among medical institutions faces a contradiction between data privacy protection and collaborative efficiency. In particular, the scarcity of data in rare cases exacerbates the difficulty of analysis, making it difficult to achieve high-quality data supplementation and collaborative analysis.

Method used

We employ a federated learning approach, anonymizing patient feature data through a noise-adding privacy protection mechanism, training the data using machine learning to generate global model parameters, combining adversarial data generation techniques to generate supplementary data locally, and integrating real data for multi-party collaborative analysis. We optimize the dataset by adjusting the weights of the discrimination components, ensuring a balanced distribution of data while protecting privacy.

Benefits of technology

It has improved the comprehensiveness and accuracy of rare disease case feature representation without compromising patient privacy, supported comprehensive medical data analysis through secure multi-institutional collaboration, and improved the accuracy of diagnosis and treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768631A_ABST
    Figure CN121768631A_ABST
Patent Text Reader

Abstract

The invention provides a distributed rehabilitation data analysis method based on federal learning, and the method comprises the steps: obtaining an encrypted rehabilitation medical data abstract, and obtaining an anonymized patient feature data set suitable for multi-party cooperation; performing learning model training on the anonymized patient feature data set in the local environment to obtain a shared global model parameter set; simulating distribution characteristics of rare disease cases based on parameters, performing adversarial data generation processing on the global model parameter set to obtain a patient information supplementary data set, and fusing real patient rehabilitation data to obtain a comprehensive rare disease case characteristic representation set; according to the comprehensive rare disease case feature representation set, a patient data distribution equilibrium index is calculated, a final distributed patient rehabilitation medical data analysis result is obtained, the problem of data insufficiency under privacy protection is effectively solved, the comprehensiveness and accuracy of rare disease case feature representation are improved, and the patient rehabilitation medical data analysis efficiency is improved. And comprehensive medical data analysis of multi-mechanism safety cooperation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a distributed rehabilitation data analysis method based on federated learning. Background Technology

[0002] In the healthcare field, rehabilitation data analytics can provide doctors with data for disease diagnosis, thereby helping to improve patient recovery outcomes and optimize treatment plans. Especially in a distributed environment, cross-institutional collaborative analysis can provide a more comprehensive perspective on rare cases and complex rehabilitation needs. However, research and application in this field face many challenges, urgently requiring innovative methods to overcome existing bottlenecks.

[0003] Currently, most rehabilitation data analysis methods do not perform ideally in distributed environments, primarily due to the conflict between data privacy protection and collaborative efficiency. Traditional methods often require direct sharing of patient information between institutions, which not only raises privacy concerns but may also lead to low analytical efficiency due to the complexity of data transmission and integration. More importantly, this approach struggles to address situations with uneven data distribution and insufficient sample sizes across different institutions, especially in rare rehabilitation scenarios where data scarcity further exacerbates the difficulty of analysis.

[0004] The core technical challenges facing distributed rehabilitation data analysis are becoming increasingly apparent. The first is how to achieve collaborative data value extraction without directly sharing raw information. This means ensuring the participation of all institutions in the analysis process while protecting privacy. This requirement further raises another crucial issue: how to generate sufficiently high-quality supplementary information during collaboration to compensate for the data deficiencies of some institutions.

[0005] Therefore, how to ensure the privacy and security of various institutions in a distributed environment while simultaneously enhancing analytical capabilities through collaborative generation of high-quality supplementary information has become a key issue that this research urgently needs to address. Summary of the Invention

[0006] This invention provides a distributed rehabilitation data analysis method based on federated learning, which mainly includes: Extract the encrypted rehabilitation medical data summary to obtain an anonymized patient feature dataset suitable for multi-party collaboration; The anonymized patient feature dataset in the local environment is used to train a learning model to obtain a shared global model parameter set; Based on the distribution characteristics of rare disease cases simulated by parameters, adversarial data generation processing is performed on the global model parameter set to obtain a supplementary dataset of patient information, and real patient rehabilitation data is fused to obtain a comprehensive set of rare disease case feature representations. Based on the comprehensive rare disease case feature representation set, the patient data distribution equilibrium index is calculated to obtain the final distributed patient rehabilitation medical data analysis results.

[0007] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a disease rehabilitation medical data analysis method based on privacy protection and federated learning. It addresses the logical problems of high risk of patient data privacy leakage among medical institutions, incomplete analysis due to insufficient rare disease case data, and difficulty in achieving balanced data distribution in multi-party collaboration. The method obtains patient feature datasets anonymized using a noise-adding privacy protection mechanism from various medical institutions, conducts federated machine learning training to securely aggregate gradient information to obtain a global model parameter set, and then uses adversarial data generation technology to generate preliminary synthetic patient data locally. An optimized supplementary dataset is obtained by adjusting the weights of the discriminant component. Distributed multi-party collaborative analysis is then performed by integrating local real data. If the coverage is insufficient, generation control instructions are extracted to expand the subset dataset. Finally, the local model is updated to calculate the equilibrium index, yielding distributed analysis results. This method effectively solves the problem of insufficient data under privacy protection, improves the comprehensiveness and accuracy of rare disease case feature representation, and enables comprehensive medical data analysis through secure multi-institutional collaboration. Attached Figure Description

[0008] Figure 1 This is a flowchart of a distributed rehabilitation data analysis method based on federated learning according to the present invention.

[0009] Figure 2 This is a model structure diagram of a distributed rehabilitation data analysis method based on federated learning according to the present invention. Figure 3 This is a schematic diagram of a distributed rehabilitation data analysis method based on federated learning according to the present invention.

[0010] Figure 4 This is another schematic diagram of a distributed rehabilitation data analysis method based on federated learning according to the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. The described embodiments are merely some embodiments of the present invention.

[0012] like Figure 1-4 This embodiment of a distributed rehabilitation data analysis method based on federated learning may specifically include: S101. Obtain the encrypted rehabilitation medical data summary to obtain an anonymized patient feature dataset suitable for multi-party collaboration.

[0013] The abstract anonymizes the original patient recovery records using a noise-adding privacy protection mechanism.

[0014] First, the rehabilitation data summary from medical institutions needs to be encrypted. The specific encryption process is as follows: 1. The patient records were anonymized using a noise-adding mechanism to obtain preliminary anonymization features.

[0015] 2. Based on the initial anonymity features, a differential privacy algorithm is used to add noise to obtain a noise-added encrypted digest.

[0016] 3. Extract patient features suitable for multi-party collaboration from the encrypted digest to obtain an anonymized patient feature dataset, and perform multi-party collaborative data sharing on the anonymized patient feature dataset to obtain fused features after collaborative data sharing.

[0017] For example, in one implementation, the process of obtaining encrypted rehabilitation medical data summaries from various medical institutions first involves a secure data transmission mechanism. Medical institutions process patient rehabilitation records, such as progress of physical therapy and functional recovery indicators, into summary form using encryption algorithms.

[0018] Specifically, these summaries employ symmetric encryption to ensure that data is not accessed without authorization during transmission. Additionally, secure network channels need to be established between medical institutions, such as using a Virtual Private Network (VPN) to transmit these encrypted summaries, thereby achieving initial data protection. This approach is suitable for multi-institutional collaboration in the field of rehabilitation medicine, such as multiple hospitals sharing fracture recovery data for joint research. Furthermore, the summaries anonymize the original patient rehabilitation records using a noise-added privacy protection mechanism. Noise-added privacy protection is a differential privacy-based technique that introduces random noise into the data to obscure individual information while maintaining overall statistical properties.

[0019] For example, when processing raw patient rehabilitation records, firstly, key features such as the patient's age group, rehabilitation duration, and recovery score need to be extracted. These key features are then summarized from the records, and Laplace noise, generated according to a preset privacy budget parameter, is introduced and added to the feature values.

[0020] For example, for a recovery score, if the original value is 80, a noise value drawn from a Laplace distribution, such as ±2, is added to obtain an anonymized value of 82. This mechanism ensures that even if an attacker obtains the data, it is difficult to reverse engineer the original information of a specific patient. In rehabilitation medicine scenarios, this is suitable for processing recovery records of chronic diseases, such as daily activity data of arthritis patients, balancing privacy and data availability through noise addition.

[0021] It's important to note that the noise addition process involves several key steps. First, a privacy budget is defined, which controls the amount of noise; for example, a smaller budget results in more noise and stronger privacy protection. Next, feature extraction is performed on the original patient recovery records. This involves separating non-sensitive statistical indicators such as average recovery rate from the records, while sensitive information such as patient names is directly removed. Finally, the noise is applied to the extracted features.

[0022] Specifically, for numerical features, the Laplace mechanism is used to calculate the noise value, which follows a distribution with a mean of 0, and the scale parameter is determined by the privacy budget.

[0023] For example, when processing stroke rehabilitation data, noise is added to patients' motor function scores to ensure that the statistical distribution of the anonymized dataset is similar to the original, but individual data is effectively obscured. This detailed process is particularly important in rehabilitation medicine collaborations because it allows multiple institutions to share data for model training without violating privacy regulations. Through this mechanism, the processed data summary can support multi-party analysis, such as predicting rehabilitation success rates, without exposing personal details. In practical applications, this method enables efficient data utilization while providing strong privacy protection.

[0024] In one possible implementation, anonymization is further combined with data aggregation techniques. The original patient recovery records are first aggregated into group-level data, for example, by summing the recovery curves of multiple patients into an average curve, and then noise is added. This reduces the impact of noise on individual data points, improving the accuracy of the anonymized dataset.

[0025] For example, when processing spinal injury rehabilitation records, walking ability indicators of multiple patients are aggregated to form a population statistic, and then noise is added to ensure that the anonymized feature dataset is suitable for multi-party collaboration, such as joint clinical trials.

[0026] In one embodiment, diverse treatments are provided for different types of rehabilitation.

[0027] For example, in cardiac rehabilitation data, noise addition focuses on heart rate recovery indicators, and the scale of the added noise is adjusted according to data sensitivity to ensure that the anonymized dataset supports cardiac health research across institutions. This implementation demonstrates the versatility of the technology, covering multiple scenarios within the same domain.

[0028] Specifically, the principle of noise-added privacy protection mechanisms further involves the sensitivity calculation of query functions. When processing raw records, the sensitivity of features is first evaluated, that is, the maximum impact of a single data change on the overall output, and then the noise scale is determined accordingly.

[0029] For example, in joint replacement rehabilitation data, sensitivity calculations target the recovery angle index, ensuring that added noise is sufficient to mask individual differences. This interpretation helps in understanding how mechanisms achieve balance in medical data.

[0030] In one embodiment, the encrypted digest of rehabilitation medical data, once acquired, is directly input into the anonymization process. The healthcare institution uses public-key cryptography to generate the digest, for example, encrypting patient functional independence measurement data, before transmission. This combination ensures a complete privacy chain from source to processing, providing end-to-end protection in rehabilitation collaboration.

[0031] For example, in processing neurorehabilitation data, anonymized patient feature datasets can be used for multi-party collaborative trend analysis, such as identifying common recovery bottlenecks, without disclosing personal information. The generation process of such datasets, through a noise-addition mechanism, ensures a balance between data availability and privacy, thus providing reliable technical support for practical medical research. Furthermore, the application of this technology in the field of rehabilitation medicine enables secure data sharing, such as multiple clinics jointly optimizing treatment protocols, while anonymization ensures compliance. This objective effect is achieved through a detailed mechanism, supporting a wide range of collaborative scenarios.

[0032] S102. Train the learning model on the anonymized patient feature dataset in the local environment to obtain a shared global model parameter set.

[0033] This invention employs a federated machine learning model training approach, updating model parameters within the local environments of each medical institution and securely aggregating local gradient information. Specifically, by anonymizing patient data, local feature datasets are obtained from each medical institution to acquire initial model parameters. Based on these initial model parameters, a stochastic gradient descent algorithm is used in the local environment. Using the local feature dataset and the initial model parameters as input values, parameter adjustment values ​​are iteratively calculated to update the parameters and obtain local gradient information.

[0034] Then, the uploaded content is obtained from the local gradient information. If the gradient information exceeds a preset threshold, compression is performed using quantization encoding, where quantization encoding maps gradient values ​​to a finite-bit representation, resulting in an optimized gradient set. Using this optimized gradient set, a secure aggregation method is employed on the central server. This method involves adding random noise to each uploaded gradient using a noise-adding mechanism and then averaging the results to obtain global model parameters. Based on these global model parameters, a shared set is distributed to each medical institution, resulting in an updated global model parameter set.

[0035] For example, in one implementation, understanding the specific implementation of the anonymization process is crucial for anonymizing a patient feature dataset. An anonymized patient feature dataset refers to transforming raw medical data into a set of features that do not contain personally identifiable information by removing or replacing personal patient information such as name, ID number, and address. In medical institutions, patient features may include numerical indicators such as age group, symptom description, and test results. This data is processed using differential privacy technology to ensure that individual patient information cannot be reverse-engineered.

[0036] Specifically, anonymization can be achieved by adding noise or generalization methods, such as converting precise age into an age range, thereby protecting privacy in federated learning. Furthermore, the core of training federated machine learning models lies in distributed collaboration mechanisms. Federated learning is a framework that allows multiple participants to collaboratively train a model without sharing the original data. It is particularly suitable for the medical field because patient data is highly sensitive and cannot be transferred across institutions.

[0037] In scenarios involving multiple hospitals, each hospital trains its model using a local anonymized dataset and updates its local parameters without exposing the data itself. The goal of this approach is to improve the overall model's accuracy in disease prediction while complying with data protection regulations. Local updates refer to each institution running machine learning algorithms, such as neural networks, within its own computing environment, iteratively training on its local dataset. Specifically, after model initialization, each institution calculates its local gradient—the derivative of the model parameters with respect to the loss function—and updates the parameters using a backpropagation algorithm.

[0038] For example, a hospital might use a patient feature dataset to train a model for cancer risk assessment, generating gradient information after several local iterations. This local training reduces the risk of data leakage and allows the institution to adjust the learning rate according to the characteristics of its own data.

[0039] Preferably, secure aggregation of gradient information is a key step in federated learning. Gradient information consists of parameter change vectors computed during model training. These vectors do not contain the original data but must be uploaded encrypted to prevent theft. Secure aggregation refers to performing weighted averaging or other fusion operations on gradients from various institutions on a central server or distributed nodes, while using encryption techniques such as homomorphic encryption or secure multi-party computation to ensure that the aggregation process does not reveal details of individual gradients. In a medical scenario, for example, multiple clinics may participate in training a shared diabetes diagnostic model. Each institution uploads encrypted gradients, and the server aggregates them to generate a global update. The principle behind this mechanism is to prevent any party from inferring others' data through noise injection or secret sharing, thereby achieving model optimization under privacy protection. The business value of this process is that medical institutions can benefit from global data without sharing sensitive patient information, thereby improving diagnostic accuracy and treatment planning.

[0040] It should be noted that there can be several variations of the specific implementation of secure aggregation.

[0041] In one embodiment, the FedAvg algorithm is used as a basis, where gradients uploaded by various institutions are aggregated by average. The specific process includes: the server collecting encrypted gradients, decrypting them, calculating the average, and then broadcasting the average back to each institution to update their local models.

[0042] For example, when processing cardiovascular disease datasets, the gradients of one hospital reflect local patient characteristics, while the aggregated global model integrates epidemiological patterns from multiple locations. This aggregation ensures strong model generalization ability, making it applicable to medical practices in different regions.

[0043] For example, in another implementation, differential privacy is introduced into the aggregation process to further enhance security. Differential privacy achieves quantified protection of privacy by adding Gaussian noise to the gradient.

[0044] Specifically, the noise level is set according to a privacy budget to ensure that the aggregation results are statistically indistinguishable from individual contributions. Medical institutions add noise to the gradients before uploading, eliminating the need for additional decryption during server aggregation. This approach is particularly effective when dealing with rare disease datasets because it allows participation from small sample institutions without concerns about privacy exposure due to data scarcity. Based on the above steps, a shared global set of model parameters is further obtained. This global parameter set consists of updated model weights and biases after aggregation, and these parameters are distributed back to each institution for local model synchronization.

[0045] For example, after several rounds of federated training, the global model can be used to predict the risk of patient relapse, and hospitals can optimize their treatment pathways accordingly. The principle behind this sharing mechanism is iterative optimization, which eventually converges to performance superior to a single-institution model.

[0046] In one embodiment, the heterogeneity of data across different healthcare institutions, such as uneven data distribution, is considered. To address this, a personalized federated learning variant is employed, where a global model serves as the foundation, and each institution fine-tunes its local parameters. Specifically, after aggregating the global gradient, institutions further optimize using their local data to adapt to regional healthcare disparities. For example, a city hospital's data might be biased towards chronic diseases, while a rural institution's data might be biased towards acute diseases. In this way, the global parameter set supports diverse applications.

[0047] Understandably, this federated training process can be extended to more scenarios, such as multimodal data integration. When processing patient features from images and text, each institution extracts feature vectors locally, and then the classifier is trained in a federated manner. The aggregated global parameters improve the model's diagnostic ability for complex diseases without requiring centralized data storage. Finally, in implementation, the benefits are reflected in improved model accuracy and privacy compliance. Through a federated approach, models trained collaboratively by medical institutions outperform isolated models on the test set while meeting data protection requirements. This method is applicable to various subfields of medicine, such as infectious disease prediction, ensuring the versatility of the technical solution.

[0048] S103. Based on the parameter simulation of the distribution characteristics of rare disease cases, adversarial data generation processing is performed on the global model parameter set to obtain a supplementary patient information dataset, and real patient rehabilitation data is fused to obtain a comprehensive rare disease case feature representation set. Specifically, the method of simulating the distribution characteristics of rare disease cases based on parameters and performing adversarial data generation processing on the global model parameter set to obtain a supplementary patient information dataset includes: 1. Based on the parameter simulation of the distribution characteristics of rare disease cases, adversarial data generation technology is applied to the global model parameter set to generate preliminary synthetic patient data in the local environment.

[0049] By using a global model parameter set, distribution characteristic parameters of rare disease cases are obtained to obtain a preliminary simulated distribution. For this preliminary simulated distribution, an adversarial generative technique is employed. This technique includes a generator and a discriminator. The generator produces synthetic data based on the distribution characteristic parameters, and the discriminator evaluates the realism. Synthetic patient data is then generated in a medical institution environment to determine supplementary patient information.

[0050] From the supplementary patient information, a validation index for the disease case distribution is obtained. This index is calculated by comparing the statistical difference between the synthetic data and the true distribution. If the validation index meets a preset threshold, a preliminary supplementary dataset is obtained. Based on the preliminary supplementary dataset, a privacy protection mechanism for local data generation is simulated. This mechanism adds noise-protected data using a differential privacy method to obtain enhanced distribution features. Using these enhanced distribution features, expanded patient information is generated, resulting in the final rare disease simulation dataset.

[0051] For example, in one implementation, the global model parameter set refers to shared parameters obtained from federated learning or distributed training, which encapsulate the statistical characteristics of historical patient data from multiple medical institutions without containing actual sensitive information.

[0052] Specifically, the global model parameter set includes weight vectors and bias terms, used to initialize the local generative model. In this way, local healthcare institutions can use these parameters to guide the generation of synthetic data without transmitting the original data, thereby protecting patient privacy. Furthermore, the adversarial data generation technology is based on the principles of generative adversarial networks (GANs), where the generator module is responsible for creating synthetic patient data, and the discriminator module evaluates the authenticity of this data.

[0053] It's worth noting that this technology is particularly suitable for the medical field because it can generate high-quality supplementary data without accessing external databases, thus reducing the risk of data breaches. In practice, medical personnel first load global parameters into a local server, then initiate an adversarial training loop, typically requiring several iterations to converge the model.

[0054] Preferably, the generation technology is based on parameter simulation of the distribution characteristics of rare disease cases, where the distribution characteristics include statistical models of disease occurrence probability, symptom correlation, and patient demographics.

[0055] For example, for a rare genetic disease like cystic fibrosis, the distribution model embedded in the parameter set simulates its low incidence and high variability. The process involves extracting a prior distribution from global parameters, such as using a Gaussian mixture model to represent the symptom vector distribution, and then the generator samples new cases based on this. This simulation ensures the diversity and representativeness of the synthetic data, avoiding the scarcity of rare cases in real-world data. In a local environment, this simulation step is achieved by adjusting parameter weights, such as increasing the probability of rare symptom generation to match known epidemiological data, thus obtaining a more balanced supplementary dataset.

[0056] Specifically, the process of generating preliminary synthetic patient data begins with data preparation: the medical institution loads a global set of model parameters and prepares local noise input vectors. Subsequently, the generator network uses these parameters to initialize its layer structure; for example, convolutional layers are used to process symptom sequence data. The discriminator employs a multilayer perceptron architecture to perform binary classification on the generated data. Through multiple forward and backward propagations, the model is progressively optimized to ensure that synthetic patient information, such as blood test results and imaging features, conforms to medical standards. The advantage of this local deployment is that it generates data in real time, supporting research or diagnostic assistance within the hospital without the need for external connections.

[0057] To simulate the distribution characteristics of rare disease cases, the process understandably begins by analyzing statistical indicators in global parameters, such as mean and variance, to describe the potential distribution of the disease. The generator then transforms these indicators to generate specific patient records; for example, a synthetic case might include a 25-year-old patient with a specific gene mutation and respiratory symptoms. Furthermore, to enhance accuracy, medical institutions can introduce local adjustment mechanisms, such as fine-tuning parameters based on their own historical cases, thus making the simulation more closely reflect regional characteristics. A key aspect of this step is distribution matching: the generator aims to minimize the difference between the synthetic distribution and the target rare disease distribution, typically quantified and optimized using a loss function.

[0058] Specifically, the steps to obtain a preliminary supplementary dataset of patient information involve post-processing the generated synthetic data, such as denoising and validation.

[0059] The specific process includes: after the generator outputs a batch of synthetic records, the discriminator performs a final evaluation and selects samples with high confidence. Subsequently, these samples are organized into a structured dataset containing fields such as patient ID, disease type, and treatment history.

[0060] In one embodiment, for the simulation of rare neurodegenerative diseases, the dataset may generate hundreds of virtual cases to supplement the lack of real data, thereby enabling the training of downstream diagnostic models. The generation of this supplementary dataset not only increases the amount of data but also maintains the realism of the distribution, helping medical institutions conduct more reliable rare disease research. In another implementation, the generation technology can flexibly adjust the simulation parameters for different types of rare diseases, such as rare cancers or immunodeficiency diseases.

[0061] For example, for rare cancers, the distribution features focus on the combined distribution of tumor markers and genetic variations, and the generator produces a synthetic record that includes imaging data and biomarkers.

[0062] It's important to note that this adjustment is achieved by modifying a subset of global parameters, such as the parameter vectors for specific disease modules, thus demonstrating the technology's versatility. In a local hospital environment, this approach allows doctors to generate targeted datasets based on their needs, supporting personalized medical applications. Furthermore, the logical flow of the entire generation process begins with parameter loading, then proceeds to the adversarial training phase, and finally outputs a supplementary dataset. To ensure consistency, medical institutions can monitor training metrics, such as the decreasing trend of the generator's loss value, to determine the iteration stopping point.

[0063] For example, in an implementation at a general hospital, the dataset generated by this technology was used to supplement the electronic medical record system, significantly improving the robustness of rare disease diagnostic models without introducing privacy risks.

[0064] Preferably, when generating preliminary synthetic patient data, diverse noise inputs can be introduced to enhance variability, for example, using uniformly distributed noise combined with global parameters to simulate different patient groups.

[0065] Specifically, for cases of rare diseases in the elderly, noise can be biased towards an older age distribution, thereby generating more realistic supplementary data. This optional feature expands the applicability of the technology, supporting data augmentation in various medical scenarios.

[0066] Understandably, the preliminary patient information supplementary dataset obtained through the above methods can be directly applied to local analytical tasks, such as training disease prediction models, thereby achieving data self-sufficiency within medical institutions. This approach not only covers the simulation of rare diseases but also ensures the safety and efficiency of the generation process.

[0067] 2. Compare the supplementary patient information dataset with the local real patient rehabilitation data. If the similarity in distribution between the preliminary supplementary patient information dataset and the local real patient rehabilitation data does not reach a preset threshold, then adjust the weight of the discrimination component of the adversarial data generation technology through the data optimization processing step to obtain an optimized supplementary patient information dataset.

[0068] Specifically, by supplementing the preliminary patient information dataset with local real patient rehabilitation data, a distribution similarity index is obtained. This index is obtained by calculating the statistical difference between the two. If the distribution similarity index does not reach a preset threshold, then optimization requirements are determined.

[0069] To address the optimization requirements, the current weights are obtained from the discriminant component of the adversarial data generation technology. The weights of the discriminant component are adjusted by increasing the evaluation strength of the discriminant component on the authenticity of the synthesized data, resulting in the adjusted weights.

[0070] Based on the adjusted weights, extended synthetic patient data is generated, and rehabilitation distribution characteristics are obtained to obtain enhanced patient information supplementation. Simulating the privacy mechanism of local medical institutions, an optimized patient information supplementation dataset is obtained.

[0071] For example, in one implementation, the distribution similarity assessment of the preliminary supplementary patient information dataset is based on statistical metrics. Specifically, the local medical institution first extracts key statistical features from the preliminary synthetic dataset, such as the mean and variance of the age distribution, and a histogram of symptom frequencies. These features are then compared with corresponding features from local real patient rehabilitation data, using methods such as the Kolmogorov-Smirnov test to quantify distribution differences. This test calculates the maximum deviation of the cumulative distribution functions of the two distributions; if the deviation exceeds a preset threshold, such as 0.1, the similarity is deemed insufficient. This assessment process ensures the statistical reliability of the synthetic data and supports subsequent optimization. Further, if the similarity does not reach the preset threshold, the data optimization process begins. The core of this process is adjusting the weights of the discriminant components in the adversarial data generation technique.

[0072] It's important to note that adversarial data generation techniques are typically based on generative adversarial networks (GANs), where the discriminant component is responsible for distinguishing synthetic data from real data. Weight adjustments involve modifying the neural network layer parameters of the discriminator, for example, fine-tuning the weight vectors of fully connected layers using gradient descent to enhance their sensitivity to distributional differences. In a local healthcare environment, this adjustment is performed on a dedicated server to avoid external data transmission.

[0073] In scenarios involving rare disease rehabilitation data, the optimization process begins with analyzing distribution mismatches. For example, if the distribution of rehabilitation duration in the synthetic dataset deviates from the real data, the weights of the discriminant component are adjusted to place greater emphasis on duration-related features. The specific process includes: first, calculating the discriminant's loss function to identify the input features causing the bias; then, applying a weight decay strategy to progressively update the weights, enabling the discriminant to better capture patterns from the real rehabilitation data. Through several iterations, this adjustment can improve the similarity to above a threshold.

[0074] Preferably, the data optimization process can incorporate a local feedback mechanism to adapt to the needs of different medical institutions.

[0075] In one possible implementation, hospital technicians monitor the optimization progress, using visualization tools to display the changing curves of distribution similarity. If the initial similarity is 0.05, below the threshold of 0.15, the weight adjustments are strengthened by increasing the number of training epochs for the discriminator. This mechanism ensures the targeted nature of the optimization; for example, for rehabilitation data of rare immune diseases, the adjustment focuses on matching the distribution of symptom recovery probabilities.

[0076] Understandably, after weight adjustment, synthetic data is regenerated to form an optimized supplementary dataset of patient information.

[0077] In one embodiment, the process for optimizing rehabilitation data for neurodegenerative diseases emphasizes the weighting of multimodal features.

[0078] For example, the weights of the discriminant component are increased to focus on the joint distribution of imaging data and physiological indicators. This adjustment improves the similarity of the optimized dataset from an initial low value, compensating for the deficiencies in real-world data and supporting research applications within the hospital. Furthermore, the logical flow of this optimization process begins with similarity assessment, followed by weight adjustment, and finally outputs the optimized dataset. To ensure consistency, medical institutions can record parameter changes for each adjustment for subsequent auditing. This method is applicable to various rare disease scenarios, such as the generation of rehabilitation data for hereditary diseases, demonstrating the versatility of the technology.

[0079] The fusion of real patient rehabilitation data yields a comprehensive set of rare disease case feature representations, specifically including: By integrating the optimized supplementary patient information dataset with real patient rehabilitation data held by local medical institutions, a distributed, multi-party collaborative data analysis is performed. The analysis process aggregates results from multiple institutions to obtain a comprehensive set of rare disease case feature representations.

[0080] 1. Obtain the real patient rehabilitation data from local medical institutions, and perform preliminary matching with the optimized supplementary patient information dataset to determine the matched joint dataset.

[0081] 2. For the aforementioned joint dataset, a federated learning algorithm is used to perform distributed, multi-party collaborative data analysis to obtain the local analysis results of each institution. The input to the federated learning algorithm is the patient rehabilitation indicators and supplementary information within the joint dataset, and the output is the model parameters of each institution based on gradient averaging.

[0082] 3. Aggregate the local analysis results of each mechanism to generate an intermediate aggregated model.

[0083] 4. Based on the intermediate aggregation model, extract rare disease-related indicators. By calculating the similarity scores between the indicators and historical cases, if an indicator exceeds a preset threshold, adjust the extraction parameters, which include similarity weights and case filtering ranges, to obtain an adjusted set of indicators. Using this adjusted set of indicators, integrate case details from multiple institutions to generate a comprehensive set of rare disease case feature representations.

[0084] For example, in one implementation, a distributed, multi-party collaborative data analysis process begins by fusing an optimized supplemental dataset of patient information with real patient rehabilitation data held by a local healthcare institution. This fusion first ensures format compatibility between the two datasets, for example, aligning patient symptom descriptions in the supplemental dataset with treatment records in the rehabilitation data to form a unified patient information framework. The optimized supplemental dataset typically includes rare disease ancillary data extracted from public databases, such as genetic variation information, while the local real patient rehabilitation data includes recovery indicators from actual cases, such as symptom relief time and relapse rate. This fusion provides a more comprehensive rare disease data foundation, supporting subsequent analysis.

[0085] Specifically, distributed multi-party collaborative data analysis refers to joint calculations performed by multiple medical institutions without directly sharing raw data. For example...

[0086] In one possible implementation, institutions use a federated learning framework, where each local institution trains a local model based on its own data and then uploads only the model parameters, not the raw data. The core of this framework is protecting patient privacy through differential privacy mechanisms, ensuring data integrity during collaboration. The analysis process focuses on extracting key features of rare diseases, such as disease progression patterns and treatment response variables, and improves the accuracy and comprehensiveness of the analysis by calculating average feature values ​​through multi-party collaboration. Furthermore, the analysis process aggregates results from multiple institutions to obtain a comprehensive set of rare disease case feature representations. The aggregation step includes collecting the local analysis outputs from each institution; for example, one institution might calculate symptom clustering results for a specific rare disease, while another institution provides statistics on rehabilitation pathways. These outputs are then merged into a global result using a secure multi-party computation protocol, such as a secret-sharing method. This aggregation avoids the risks of data centralization and is applicable to different rare disease scenarios in the medical field, such as neurodegenerative diseases or inherited metabolic disorders, ensuring that the feature representation set covers a wide range of case variations.

[0087] For example, in a scenario targeting a rare genetic disease, the fusion process first cleanses the supplementary dataset, removing redundant gene sequence information, and then matches it with local rehabilitation data, such as associating the patient's genotype with functional recovery data during the rehabilitation period. Distributed analysis is then performed collaboratively by multiple hospitals, with each hospital processing symptom evolution data from local cases and calculating local feature vectors, such as disease severity scores. Through parameter aggregation, a comprehensive feature representation set is generated to predict potential relapse risk. This approach, in rare disease research, enables privacy protection through data sharing and improves the generalization ability of diagnostic models.

[0088] It's important to note that the principle of distributed, multi-party collaborative data analysis lies in distributing computational tasks. Each institution independently performs feature extraction steps, such as using statistical models to calculate the average recovery period for cases, without needing to transmit sensitive personal health records. When aggregating results from multiple institutions, a weighted average strategy is employed, where the weights are based on the proportion of data volume from each institution, ensuring the representativeness of the comprehensive feature representation set.

[0089] In one embodiment, for the analysis of rare immunodeficiency diseases, the fusion dataset includes supplemental information on global cases, while local data covers regional rehabilitation tracking. Feature sets, such as immune response pattern vectors, are generated collaboratively to support cross-institutional research collaborations.

[0090] Preferably, in another implementation, the analysis process can be extended to more institutions, such as collaborations involving international medical alliances. The fusion step emphasizes data standardization, such as unifying rare disease classification codes, and then performing distributed analysis to aggregate the results into a multidimensional feature representation set. This set includes not only numerical features such as the age distribution of onset, but also categorical features such as treatment effectiveness labels. In this way, the technical solution demonstrates versatility in the medical field, applicable to various types of rare diseases, rather than being limited to a single disease.

[0091] Understandably, once a comprehensive set of rare disease case feature representations is obtained, this set can be used for further medical decision support. The feature set is fed into a predictive model to analyze potential disease progression trends, thus providing a reference for clinicians. This process ensures the logical coherence of the analysis, a complete chain from data fusion to result output.

[0092] Homomorphic encryption allows addition operations to be performed in an encrypted state, generating a comprehensive set of features that does not expose the original contributions. In medical collaborations, this technology aims to balance data utilization with privacy protection, providing a more reliable foundation for research on rare diseases.

[0093] In addressing rare neurological diseases, the fusion and analysis process can refine feature extraction steps, such as extracting neurological function scores from rehabilitation data and combining them with brain imaging information from supplementary datasets. Through distributed collaboration, the aggregated results form a set of feature representations for identifying disease subtypes. This implementation demonstrates the flexibility of the technical solution, supporting applications across various rare disease scenarios.

[0094] S104. Calculate the patient data distribution equilibrium index based on the comprehensive rare disease case characteristic representation set to obtain the final distributed patient rehabilitation medical data analysis results.

[0095] Specifically, it includes: First, for the comprehensive rare disease case feature representation set, if the feature representation does not meet the preset standard for coverage of the data-insufficient area, then control instructions are extracted from the shared global model parameter set to obtain an expanded supplementary subset of patient information.

[0096] For the rare disease features, a coverage assessment value is obtained from the data-deficient areas. This assessment value is calculated by determining the coverage ratio of the feature representations within the data-deficient areas to the comprehensive case set. If the coverage assessment value is lower than a preset threshold, shared parameter extraction results are extracted from the global model parameters to obtain generation control instructions. Based on these instructions, the expansion range for supplementing patient information is determined. A pre-established generative network is used to obtain supplementary feature values ​​from the shared parameter extraction results. This generative network is constructed using a variational autoencoder, with the input being the shared parameter extraction results combined with the generation control instructions, and the output being the supplementary feature values, thus obtaining the subset expansion basis. Using this subset expansion basis, the feature representation optimization requirements for the comprehensive case set are determined. These requirements are obtained by comparing the matching degree between the subset expansion basis and the comprehensive case set, acquiring the information generation fusion path, and determining the expanded supplementary patient information subset. The fused rare disease features are obtained from the expanded supplementary patient information subset. Updated coverage assessment values ​​are determined for the data-deficient areas. These updated values ​​are obtained by recalculating the coverage ratio of the fused rare disease features to the data-deficient areas, resulting in an optimized feature representation set. The feature representation optimization set is used to determine the completeness of information generation fusion for the extended subset. The completeness of information generation fusion is obtained by examining the integration consistency between the feature representation optimization set and the extended subset, thus obtaining the final supplementary subset of patient information.

[0097] For example, in one implementation, for a comprehensive set of rare disease case feature representations, these features first need to be collected and represented.

[0098] For example, the rare disease case feature representation set may include multi-dimensional data such as patient symptom descriptions, genetic information, imaging data, and laboratory indicators. These data are standardized to form vectors or matrices to ensure consistency.

[0099] For example, when dealing with a rare genetic disease, feature representations may encompass elements such as gene mutation type, clinical manifestations, and age of onset, forming a multidimensional feature space. Furthermore, determining whether the feature representations cover regions with insufficient data to the extent that a preset standard is met is a crucial step.

[0100] Specifically, insufficient data regions refer to sparsely distributed or missing portions of the feature space. For example, in rare disease data, there may be insufficient sample sizes for certain age groups or regions. Coverage can be assessed by calculating the density distribution of the feature space, such as using clustering algorithms to analyze the distances between sample points. If the average distance exceeds a preset threshold, it is considered not to meet the standard. This judgment process helps identify data gaps and ensures that subsequent data supplementation is targeted.

[0101] In one possible implementation, the preset criteria can be dynamically adjusted according to the disease type. For example, for diseases with extremely low incidence rates, the threshold can be set more strictly to cover more potential variations.

[0102] It's important to note that if the coverage doesn't meet the preset criteria, generation control instructions are extracted from a shared global model parameter set. This shared global model parameter set refers to the aggregated model parameters obtained through collaborative training by multiple medical institutions in a distributed medical data environment. These parameters do not contain raw patient data; they only share abstract model weights to protect privacy. Extracting generation control instructions involves selecting a subset of these parameters relevant to areas with insufficient data. For example, this can be achieved by querying the model's embedding layer parameters to generate instruction sequences targeting specific feature dimensions. These instructions can then guide the generative model to produce simulated data to fill in the gaps.

[0103] In one embodiment, the specific implementation of obtaining the expanded supplementary subset of patient information is as follows: Based on the extracted generation control instructions, supplementary data similar to the original feature representation is generated using frameworks such as generative adversarial networks or variational autoencoders.

[0104] For example, for a rare neurological disease, instructions might specify generating patient information with specific combinations of symptoms, such as virtual cases incorporating age and genetic factors, thereby expanding the subset dataset. This supplementary subset, when merged with the original dataset, can improve the robustness of model training.

[0105] Preferably, in practical applications, implementation in multiple scenarios can be considered.

[0106] For example, in the management of rare disease databases in hospitals, if feature coverage is insufficient, the system automatically extracts instructions from global parameters to generate supplementary data for research and analysis. This approach ensures the universality of data expansion without introducing interference from external domains. Further extending this approach, in another implementation, the extraction of control instructions can be refined to the selection of parameter subsets.

[0107] Specifically, the global model parameter set may include weights at multiple levels, such as shallow layers for basic features and deep layers for complex patterns. During extraction, the feature types in areas with insufficient data are first analyzed, and then corresponding parameters are matched to generate instruction sequences.

[0108] For example, for rare diseases with insufficient image data, parameters relevant to image recognition can be extracted to generate a supplementary subset of virtual image features. This refinement improves the accuracy of the generated dataset.

[0109] Understandably, this technical solution can effectively address the data scarcity problem when applied to rare disease diagnostic assistance systems. By supplementing the dataset with subsets, the model's predictive accuracy is improved; for example, in predicting disease progression, it covers more variation scenarios, thus providing more comprehensive support for clinical decision-making.

[0110] For example, when dealing with rare metabolic diseases, the feature representation set might include biochemical indicators and family history. If the coverage is insufficient, extract instructions generate supplementary subsets that cover data with rare variant types. This implementation demonstrates the flexibility of the approach. In a preferred implementation, the entire process is integrated on a cloud platform, allowing healthcare institutions to share parameter sets and achieve real-time data expansion.

[0111] For example, instructions are extracted immediately after coverage is determined, and a subset of data is generated for model updates. This integration enhances the system's usability.

[0112] Secondly, based on the expanded patient information, a supplementary subset is added to update the local data analysis model. After integrating the subset, the model recalculates the patient data distribution balance index to obtain the final distributed patient rehabilitation medical data analysis results.

[0113] By expanding patient information, supplementary subsets are obtained from the expanded patient information, and the rehabilitation characteristics of these subsets are determined. Using these supplementary subsets, the local data analysis model is updated, resulting in an updated model version. For the updated model version, the subsets are integrated, and the patient data distribution equilibrium index is recalculated. If the distribution equilibrium index exceeds a preset threshold, the data distribution is adjusted to obtain a balanced dataset. Based on the balanced dataset, distributed patient rehabilitation data analysis is performed to obtain the final distributed patient rehabilitation data analysis results.

[0114] For example, in one implementation, the process of expanding the patient information supplementary subset first extracts basic patient data from an existing rehabilitation medicine database, including information such as age, type of rehabilitation, and treatment history. Subsequently, the subset is expanded by supplementing it with data from external sources, such as hospital records or motion metrics monitored by wearable devices.

[0115] Specifically, this supplement aims to fill data gaps and ensure that the subset dataset covers multiple rehabilitation scenarios, such as joint recovery data in physical rehabilitation or cognitive training records in neurorehabilitation. In this way, the subset dataset can more comprehensively reflect the patient's rehabilitation process and avoid bias caused by a single source.

[0116] Preferably, when updating the local data analysis model, the supplemented subset is input into the local model.

[0117] It's important to note that the local data analysis model refers to an analytical framework deployed on a healthcare institution's server to process distributed patient data. This model may employ machine learning algorithms, such as random forests or neural networks, to analyze recovery progress. The update process involves fusing the new subset of data with the original model parameters, for example, by adjusting weights through incremental learning methods, thereby adapting the model to expanded information. This update helps improve the model's accuracy in predicting patient recovery trends. Furthermore, after integrating the subset of data, recalculating the patient data distribution equilibrium index becomes a crucial step.

[0118] Specifically, the patient data distribution balance index is a quantitative tool used to assess the degree of balance of data across different categories, such as the recovery phase or patient groups. For example...

[0119] In one possible implementation, the sample size distribution within the subset is first statistically analyzed, and then an equilibrium metric is calculated, such as using the Shannon entropy formula to measure diversity, without involving specific numerical calculations. The process includes identifying data skew; for example, if there is an abundance of physical rehabilitation data and a shortage of neurorehabilitation data, the distribution is adjusted through weighted sampling.

[0120] For example, in a neurorehabilitation scenario, if a subset of the dataset shows an uneven distribution of cognitive training data, virtual sample generation can be introduced to balance the distribution, thereby ensuring the fairness of the model analysis. This computational process is detailed as follows: First, the data is categorized, such as by rehabilitation type; second, the sample proportion for each category is calculated; then, a balancing algorithm is applied to adjust the proportions, for example, by oversampling the minority class or undersampling the majority class; finally, the adjusted distribution is verified to meet a preset threshold. This detailed calculation helps maintain data consistency in a distributed system, thereby supporting accurate rehabilitation medical decisions. In practical applications, this balancing can bring technical benefits, such as reducing model bias and improving the reliability of analysis for rare rehabilitation cases.

[0121] For example, in the implementation of physical rehabilitation, the integrated subset may include joint flexibility measurement data. After recalculating the equilibrium index, the model can output a more balanced distribution result.

[0122] Understandably, the final distributed patient rehabilitation medical data analysis results are obtained through the comprehensive output of the above steps.

[0123] Specifically, the results include patient recovery prediction reports, such as estimated recovery time or potential risk assessments. These results are distributed across multiple healthcare institutions, enabling shared analysis. In another embodiment, for patients recovering from chronic diseases, extended subsets of data can supplement long-term tracking data, such as blood pressure variability records, and the updated model can calculate equilibrium indicators to ensure data coverage of urban and rural patient population differences. Furthermore, this approach emphasizes data privacy in a distributed environment by integrating subsets through a federated learning framework without transmitting the original data, thereby protecting patient information.

[0124] For example, in neurorehabilitation analysis, the final results can generate personalized treatment recommendations, based on the calculation of balanced indicators, to ensure that the recommendations are applicable to diverse patient populations.

[0125] Preferably, the overall process can be extended to community rehabilitation centers, incorporating mobile application data when supplementing subsets, and updating the model for real-time analysis.

[0126] It should be noted that these implementation methods are all limited to the field of patient rehabilitation medicine, demonstrating the versatility of the technical solutions in different types of rehabilitation.

[0127] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the concept of this application. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A distributed rehabilitation data analysis method based on federated learning, characterized in that, The method includes: Extract the encrypted rehabilitation medical data summary to obtain an anonymized patient feature dataset suitable for multi-party collaboration; The anonymized patient feature dataset in the local environment is used to train a learning model to obtain a shared global model parameter set; Based on the distribution characteristics of rare disease cases simulated by parameters, adversarial data generation processing is performed on the global model parameter set to obtain a supplementary dataset of patient information, and real patient rehabilitation data is fused to obtain a comprehensive set of rare disease case feature representations. Based on the comprehensive rare disease case feature representation set, the patient data distribution equilibrium index is calculated to obtain the final distributed patient rehabilitation medical data analysis results.

2. The distributed rehabilitation data analysis method based on federated learning according to claim 1, characterized in that, The process involves extracting encrypted rehabilitation medical data summaries to obtain an anonymized patient feature dataset suitable for multi-party collaboration, including: Rehabilitation data summaries are obtained from various medical institutions, the rehabilitation data summaries are encrypted, and the patient records are anonymized using a noise addition mechanism to obtain preliminary anonymization features; Based on the initial anonymity features, a differential privacy algorithm is used to add noise to obtain a noisy encrypted digest. Patient features suitable for multi-party collaboration are extracted from the encrypted digest to obtain an anonymized patient feature dataset. Multi-party collaborative data sharing is then performed on the anonymized patient feature dataset to obtain fused features after collaborative data sharing.

3. The distributed rehabilitation data analysis method based on federated learning according to claim 1, characterized in that, The process of training a learning model on the anonymized patient feature dataset in the local environment to obtain a shared global model parameter set includes: Local feature datasets are obtained from various medical institutions, and federated machine learning models are trained to obtain initial model parameters. By iteratively calculating the parameter adjustment values, the initial model parameters are updated using the stochastic gradient descent algorithm in the local environment to obtain local gradient information; If the gradient information exceeds a preset threshold, it is compressed using quantization encoding to obtain an optimized gradient set. Random noise is added to each gradient information using a noise-adding mechanism, and the average value is calculated. This average value is then distributed and shared among various medical institutions to obtain an updated global model parameter set.

4. The distributed rehabilitation data analysis method based on federated learning according to claim 1, characterized in that, The method involves simulating the distribution characteristics of rare disease cases based on parameters, and then performing adversarial data generation processing on the global model parameter set to obtain a supplementary patient information dataset, including: Based on the distribution characteristics of rare disease cases simulated by parameters, adversarial data generation technology is applied to the global model parameter set to generate preliminary synthetic patient data in the local environment. The supplementary patient information dataset is compared with local real patient recovery data. If the similarity in distribution between the preliminary supplementary patient information dataset and the local real patient recovery data does not reach a preset threshold, the weight of the discrimination component of the adversarial data generation technology is adjusted through the data optimization processing step to obtain an optimized supplementary patient information dataset.

5. A distributed rehabilitation data analysis method based on federated learning according to claim 4, characterized in that, The process involves comparing the supplementary patient information dataset with local real patient recovery data. If the similarity in distribution between the initial supplementary patient information dataset and the local real patient recovery data does not reach a preset threshold, the weights of the discriminant component in the adversarial data generation technology are adjusted through a data optimization processing step to obtain an optimized supplementary patient information dataset, including: By supplementing the preliminary patient information dataset with local real patient rehabilitation data, a distribution similarity index is obtained. If the distribution similarity index does not reach a preset threshold, optimization requirements are determined. To address the aforementioned optimization requirements, the current weights are obtained from the discriminative component of the adversarial data generation technology to arrive at the adjusted weights. Based on the adjusted weights, extended synthetic patient data is generated, and rehabilitation distribution characteristics are obtained to obtain enhanced patient information supplementation. Simulating the privacy mechanism of local medical institutions, an optimized patient information supplementation dataset is obtained.

6. A distributed rehabilitation data analysis method based on federated learning according to claim 1 or 5, characterized in that, The fusion of real patient rehabilitation data yields a comprehensive set of rare disease case feature representations, including: The real patient recovery data is initially matched with the supplementary patient information dataset to determine the matched joint dataset; For the joint dataset, a federated learning algorithm is used to perform distributed multi-party collaborative data analysis to obtain the local analysis results of each institution; The local analysis results of the various mechanisms are aggregated to generate an intermediate aggregated model; Based on the intermediate aggregation model, rare disease-related indicators are extracted, and the similarity scores between the indicators and historical cases are calculated. If an indicator exceeds a preset threshold, the extraction parameters are adjusted to obtain an adjusted set of indicators, generating a comprehensive set of rare disease case feature representations. The extraction parameters include: similarity weight and case filtering range.

7. A distributed rehabilitation data analysis method based on federated learning according to claim 6, characterized in that, The process involves calculating a patient data distribution equilibrium index based on a comprehensive set of rare disease case characteristics to obtain the final distributed patient rehabilitation medical data analysis results, including: For the comprehensive rare disease case feature representation set, if the feature representation does not meet the preset standard for coverage of the data-insufficient area, then the generation control instructions are extracted from the shared global model parameter set to obtain an expanded supplementary subset of patient information. Based on the patient information, a supplementary subset is added, and the local data analysis model is updated. After integrating the subset, the model recalculates the patient data distribution balance index to obtain the final distributed patient rehabilitation medical data analysis results.

8. A distributed rehabilitation data analysis method based on federated learning according to claim 7, characterized in that, If the feature representations for the comprehensive rare disease case feature representation set do not meet a preset standard in terms of coverage of areas with insufficient data, then control instructions are extracted from the shared global model parameter set to obtain an expanded supplementary subset of patient information, including: For the characteristics of the rare disease, a coverage assessment value is obtained from the data-insufficient area. If the coverage assessment value is lower than a preset standard threshold, the shared parameter extraction result is extracted from the global model parameters to obtain the generation control command. The expansion range is determined based on the generation control instructions, the basis for the expansion of the subset dataset is obtained, and the matching degree with the comprehensive case set is obtained to obtain the feature representation optimization requirements; Determine the feature representation optimization requirements, obtain information to generate fusion paths, and determine the expanded patient information supplementary subset; The fused rare disease features are obtained from the expanded patient information supplementary subset dataset. The coverage evaluation update value is determined for the data-insufficient areas to obtain the optimized feature representation set. By examining the integration consistency between the feature representation optimization set and the supplementary patient information subset, the fusion completeness is obtained, and the final supplementary patient information subset is acquired.

9. A distributed rehabilitation data analysis method based on federated learning according to claim 7 or 8, characterized in that, The process involves supplementing the data with a subset of patient information, updating the local data analysis model, and then recalculating the patient data distribution equilibrium index after integrating the subset of data to obtain the final distributed patient rehabilitation medical data analysis results, including: The extended patient information is used to supplement the sub-dataset, the rehabilitation medical characteristics of the sub-dataset are determined, and the local data analysis model is updated to obtain the updated model version. For the updated model version, the subset datasets are integrated, and the patient data distribution equilibrium index is recalculated. If the distribution balance index exceeds the preset threshold, the data distribution is adjusted to obtain a balanced data set. Based on the balanced dataset, distributed patient rehabilitation medical data analysis is performed to obtain the final distributed patient rehabilitation medical data analysis results.

Citation Information

Patent Citations

  • Medical image data set making method based on federated learning and generative adversarial network

    CN115588487A

  • Heart disease auxiliary diagnosis system based on federal learning

    CN116403700A

  • Federal learning TMI model based on generative adversarial network enhancement

    CN118484828A

  • Large-span bridge structure health monitoring abnormity identification method based on image classification

    CN119785128A

  • Method for processing federal learning data heterogeneity based on conditional generative adversarial network

    CN119885222A

Cited By

  • Inspection and detection visual data management method and system

    CN122045448A