Federal learning-based non-alcoholic fatty liver disease data sharing analysis method

Through standardized processing and federated learning algorithms to evaluate the data contribution rate and design an incentive allocation mechanism, the difficulties of privacy protection and model collaboration in data sharing analysis of non-alcoholic lipid hepatitis disease are solved, fair contribution evaluation and model optimization are achieved, and the efficiency and model performance of data sharing are improved.

CN120299597APending Publication Date: 2025-07-11THE SECOND HOSPITAL OF TIANJIN MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417508.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the data sharing analysis of non-alcoholic fatty liver disease, it is difficult to achieve effective data sharing and model collaboration under the premise of protecting privacy, and the lack of fair incentive mechanisms leads to data quality and quantitative heterogeneity issues, affecting the contribution assessment and willingness of cooperation of the model.

Method used

Through standardized processing and federal average algorithm training data, gradient contribution rate and sample diversity characteristics are calculated, entropy evaluation method and linear weighted fusion are adopted, incentive allocation mechanism is designed, data input of low-contributing institutions is optimized, and global model iteratively is updated.

Benefits of technology

Quantitative evaluation and fair incentives for contributions to various institutions have been achieved, the overall performance and robustness of the model have been improved, and the effective implementation of multi-party collaboration and data sharing analysis has been promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299597A_ABST
    Figure CN120299597A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning-based non-alcoholic fatty liver disease data sharing analysis method, which comprises the following steps: firstly, obtaining quality indexes and quantity statistics of non-alcoholic fatty liver disease data of each mechanism, and carrying out cleaning and normalization standardization processing to obtain a data set with a uniform format; thirdly, performing initial training on the local model of each mechanism by adopting a federated average algorithm to obtain a global model initial parameter, and evaluating the initial precision of the global model initial parameter on data of each mechanism; according to the initial precision, gradient contribution of data of each mechanism to the global model is calculated, the contribution rate is determined, and weighting adjustment is carried out in combination with data quality. And further analyzing the lifting weight of the data of each mechanism to the model diversity, comprehensively calculating the contribution score of each mechanism, and allocating the excitation resources according to the contribution score. For a mechanism with low contribution, input is adjusted by increasing the sampling rate or introducing synthetic data, a global model is iteratively updated, model parameters and precision are optimized, and efficient and safe data sharing and analysis are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical data analysis, and particularly relates to a method for sharing and analyzing non-alcoholic fatty liver disease data based on federated learning. Background Art

[0002] As a globally prevalent disease, the research on non-alcoholic fatty liver disease is of crucial significance in the medical field. With the change of lifestyle and the rising obesity rate, the incidence of non-alcoholic fatty liver disease continues to climb, seriously threatening human health. Data-driven analysis methods have shown great potential in the diagnosis, prediction, and treatment of non-alcoholic fatty liver disease. Especially through multi-party collaborative data sharing, the generalization ability and accuracy of the model can be significantly improved. However, data privacy protection and the issue of interest distribution among medical institutions make data sharing an urgent bottleneck to be broken through. Federated learning has become a key technical path to solve this problem because it can achieve distributed collaborative modeling while protecting privacy.

[0003] Currently, traditional methods for sharing and analyzing non-alcoholic fatty liver disease data mostly rely on centralized data integration or local independent modeling. The former requires uploading raw data to a central server, which is prone to privacy leakage risks and has high data transmission costs; the latter avoids data concentration, but due to the isolation of data from each institution, the performance of the model is limited by the sample size and diversity, and it is difficult to cope with the complexity of the disease. These limitations lead to poor performance of existing solutions in practical applications and make it difficult to balance the requirements of privacy protection and model performance.

[0004] Although federated learning provides a framework for distributed collaboration, it still faces core challenges in the research of non-alcoholic fatty liver disease. First, the heterogeneity of data quality and quantity makes it difficult to quantify the contributions of each institution to the model, which may lead to some institutions being marginalized due to insufficient data. Second, the attribution problem of model accuracy improvement is complex, and there is no unified standard for evaluating the contributions of all parties, which affects the fairness of cooperation. If these technical factors cannot be solved, they will hinder the design of incentive mechanisms in the federated learning ecosystem, thereby weakening the willingness of participants to cooperate.

[0005] Therefore, how to design a scientific contribution rate evaluation mechanism in the non-alcoholic fatty liver disease federated learning model to quantify the roles of each participant in data quality, quantity, and model accuracy improvement, and accordingly establish a fair incentive system has become a key issue to promote the effective implementation of data sharing and analysis. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method for sharing and analyzing non-alcoholic fatty liver disease data based on federated learning, including:

[0007] Obtain the data quality indicators and data quantity statistics related to non-alcoholic fatty liver disease uploaded by each institution, and clean and normalize the original data through a preset standardized processing process to obtain a dataset in a unified format;

[0008] According to the dataset in the unified format, use the federated averaging algorithm to initially train the local models of each institution to obtain the preliminary parameters of the global model, and judge the preliminary accuracy performance of the global model on the data of each institution;

[0009] According to the preliminary accuracy performance, obtain the gradient contribution of the data of each institution to the global model, and determine the specific contribution rate of each institution in model optimization by calculating the magnitude and direction of the gradient vector;

[0010] Analyze the relationship between the gradient contribution and data quality. If the data quality of a certain institution is lower than the preset threshold, then weight-adjust the contribution rate to obtain the adjusted contribution rate;

[0011] According to the adjusted contribution rate, obtain the data quantity and sample diversity characteristics of each institution, and use the entropy-based evaluation method to judge the role weight of the data of each institution in improving the diversity of the global model;

[0012] According to the role weight and the adjusted contribution rate, use the linear weighted fusion method to calculate the comprehensive contribution score of each institution in improving model accuracy, and obtain the final contribution quantification result;

[0013] Through the comprehensive contribution score, use the preset incentive allocation algorithm to allocate resources or benefits to each institution to obtain the output result of the incentive system in distributed collaboration;

[0014] According to the output result of the incentive system, if it is judged that the comprehensive contribution score of a certain institution is lower than the average value, then adjust the data input of the institution by increasing the local data sampling rate or introducing external synthetic data, and track and judge whether the subsequent contribution is improved;

[0015] According to the adjusted data input, use the federated learning framework to iteratively update the global model to obtain the optimized model parameters and accuracy performance, and realize real-time data sharing and analysis.

[0016] Preferably, the process of obtaining the data quality indicators and data quantity statistics related to non-alcoholic fatty liver disease uploaded by each institution, and cleaning and normalizing the original data through a preset standardized processing process to obtain a dataset in a unified format includes:

[0017] Obtain the original non-alcoholic fatty liver disease-related data uploaded by each institution, the corresponding quality indicators and quantity statistics, and judge whether the data quality reaches the preset threshold. If it is lower than the threshold, perform preliminary screening on the original non-alcoholic fatty liver disease-related data to obtain the first data set;

[0018] Perform data cleaning on the first data set through a preset standardization process, remove outliers and duplicates, and obtain the second data set;

[0019] Process the second data set using a data normalization method to unify the data range and format, and obtain the third data set;

[0020] Extract feature variables from the third data set, combine with quality indicators and quantity statistics, and judge whether the features are suitable for distributed modeling. If suitable, generate the fourth data set;

[0021] Apply the distributed modeling algorithm K-means to the fourth data set for clustering analysis to obtain the clustering result;

[0022] According to the clustering result and quality indicators, judge whether the output of each institution's local model is stable. If it is not stable, adjust the normalization parameters and regenerate the fifth data set;

[0023] Perform distributed modeling again through the fifth data set, using the support vector machine algorithm, to obtain the data set in the final unified format.

[0024] Preferably, according to the data set in the unified format, the process of initially training each institution's local model using the federated averaging algorithm to obtain the preliminary parameters of the global model and judging the preliminary accuracy performance of the global model on the data of each institution includes:

[0025] Perform initial training on the local model using the federated averaging algorithm to obtain the training results of each institution;

[0026] Extract the preliminary parameters from the training results of each institution to determine the initial state of the global model;

[0027] Perform running inference on the institution data based on the global model to obtain the preliminary accuracy performance;

[0028] If the accuracy performance is lower than the preset threshold, perform a new round of training on the local model to update the preliminary parameters;

[0029] Adjust the global model according to the updated preliminary parameters, and judge the change trend of the accuracy performance of the data of each institution;

[0030] Optimize the federated averaging algorithm through trend analysis to obtain the final global model parameters.

[0031] Preferably, according to the preliminary accuracy performance, obtaining the gradient contribution of each institution's data to the global model, and determining the specific contribution rate of each institution in model optimization by calculating the magnitude and direction of the gradient vector includes:

[0032] Collecting the data information uploaded by each institution through a preset interface to obtain the initial value of the gradient contribution;

[0033] By calculating the vector norm of the initial value of the gradient contribution and using Euclidean norm processing, the norm result is obtained;

[0034] According to the norm result, calculate the direction of the gradient vector and use the arctangent function to determine the direction angle;

[0035] According to the direction angle and the norm result, obtain the gradient distribution characteristics of the global model to get the optimization reference value;

[0036] If the optimization reference value exceeds the preset threshold, adjust the global model by the gradient descent method to obtain updated parameters;

[0037] According to the updated parameters, calculate the contribution rate value of each institution's data to model optimization, and use the weighted average method to determine the final contribution rate;

[0038] Through the final contribution rate, obtain the gradient influence ranking of each institution in the global model to get the optimization priority.

[0039] Preferably, analyzing the relationship between the gradient contribution and the data quality. If the data quality of a certain institution is lower than the preset threshold, the process of weighted adjustment of the contribution rate to obtain the adjusted contribution rate includes:

[0040] Obtain the data quality parameter of the institution's data, judge whether the data quality is lower than the preset threshold. If so, use the weighted adjustment method to process the contribution rate to obtain the preliminary adjustment value;

[0041] According to the preliminary adjustment value and in combination with the gradient contribution data, determine the adjusted contribution rate;

[0042] Judge whether the adjusted contribution rate meets the requirements of the preset threshold through quality analysis to obtain the verification result;

[0043] According to the verification result, obtain the correlation coefficient between the gradient contribution and the data quality, and generate a coefficient matrix;

[0044] Adopt the linear regression algorithm to extract the key influencing factors from the coefficient matrix to obtain the final index;

[0045] Generate result data through the final index to complete the calculation of the adjusted contribution rate.

[0046] Preferably, according to the adjusted contribution rate, obtaining the data quantity and sample diversity characteristics of each institution, and adopting an evaluation method based on entropy value to determine the weight of the role of the data of each institution in enhancing the diversity of the global model includes the following steps:

[0047] Obtain the data quantity and sample characteristics submitted by each institution, and obtain the preliminary distribution result through statistical analysis;

[0048] Process the preliminary distribution result by using the entropy value calculation method to determine the diversity index of the data of each institution;

[0049] Compare the existing characteristics of the global model through the diversity index to judge the incremental impact of the data of each institution on the model diversity;

[0050] Obtain the incremental impact result, and calculate the weight of the role of the data of each institution by using the weighted average method;

[0051] If the weight exceeds the preset threshold, include the corresponding institution data in the model training data set;

[0052] Run the random forest algorithm based on the updated training data set to obtain an optimized global model;

[0053] According to the optimized global model, output the final contribution rate adjustment result of the data of each institution.

[0054] Preferably, according to the weight and the adjusted contribution rate, adopting a linear weighted fusion method to calculate the comprehensive contribution score of each institution in improving the model accuracy and obtaining the final contribution quantification result includes the following steps:

[0055] Obtain the diversity weight data submitted by each institution, and use statistical tools to calculate the initial weight value to obtain the diversity role weight;

[0056] Process the initial contribution rate data through a preset adjustment rule to obtain the adjusted contribution rate;

[0057] Adopt a linear weighted fusion method to combine the diversity role weight and the adjusted contribution rate to calculate the comprehensive contribution score of each institution;

[0058] If the comprehensive contribution score exceeds the preset threshold, it is judged that the institution has a significant role in improving the model accuracy, and the preliminary quantification result is determined;

[0059] Sort according to the comprehensive contribution score to obtain the priority sequence of the institution's role, and determine the contribution distribution of accuracy improvement;

[0060] Process the priority sequence and the comprehensive contribution score through regression analysis to obtain the final quantification result;

[0061] For the final quantization result, hierarchical clustering analysis is used to divide the institutional contribution levels, and the hierarchical contribution calculation results are obtained.

[0062] Preferably, through the comprehensive contribution score, the process of using a preset incentive allocation algorithm to allocate resources or benefits to each institution and obtaining the output result of the incentive system in distributed collaboration includes:

[0063] Through the comprehensive contribution score, the data input is standardized using a preset algorithm to obtain the normalized contribution value;

[0064] According to the normalized contribution value and combined with the number of institutions, calculate the resource allocation ratio for each institution and determine the preliminary allocation plan;

[0065] If the resource allocation ratio exceeds the preset threshold, optimize the calculation process through the adjustment algorithm rules to obtain the balanced allocation ratio;

[0066] Using the balanced allocation ratio and combined with the distributed structure, generate the income allocation data for each institution;

[0067] According to the income allocation data and the requirements of the incentive mechanism, adjust the system balance parameters to determine the final incentive output result.

[0068] Preferably, according to the output result of the incentive system, if it is determined that the comprehensive contribution score of an institution is lower than the average value, the process of adjusting the data input of the institution by increasing the local data sampling rate or introducing external synthetic data and tracking and judging whether the subsequent contribution is improved includes:

[0069] According to the output result of the incentive system, calculate the comprehensive contribution scores of each institution, determine the average value through mean calculation, and obtain the list of institutions with scores lower than the average value;

[0070] Extract the local data from the institutions with scores lower than the average value and analyze the sampling rate, adjust the data input by increasing the sampling frequency, and generate the frequency adjustment data set;

[0071] According to the frequency adjustment data set, if the local data is insufficient, introduce external synthetic data for supplementation to obtain the supplemented data set;

[0072] Rerun the incentive system through the supplemented data set, obtain the new output result, and calculate the adjusted comprehensive contribution score;

[0073] If the adjusted comprehensive contribution score is lower than the average value, analyze the correlation between the data input and the subsequent contribution through the random forest algorithm to determine the key influencing features;

[0074] According to the key influencing features, use the gradient boosting algorithm to optimize the weight allocation of the data input and generate the weight allocation data set;

[0075] Run the incentive system again through the said weight allocation, obtain the final output result, and determine whether the subsequent contribution is improved.

[0076] Preferably, according to the adjusted data input, use the federated learning framework to iteratively update the global model to obtain optimized model parameters and accuracy performance. The process of realizing real-time data sharing and analysis includes:

[0077] Obtain the adjusted data input through the federated learning framework, iteratively update the global model, and obtain preliminary model parameters;

[0078] Extract features from the said preliminary model parameters, use the gradient descent algorithm to optimize the model parameters, and obtain optimized model parameters;

[0079] Evaluate the said optimized model parameters, calculate the accuracy performance, and determine the performance indicators of the current model;

[0080] If the accuracy performance is lower than the preset threshold, re-adjust the data input through the federated learning framework and iteratively update the global model;

[0081] According to the adjusted data input and the updated global model, obtain new model parameters and accuracy performance;

[0082] Integrate the calculation results of all parties through the data sharing mechanism to realize real-time data sharing and analysis.

[0083] Compared with the prior art, the present invention has the following advantages and technical effects:

[0084] The present invention discloses a method for quantifying federated learning contributions based on distributed collaboration. The method first standardizes and cleans the original data uploaded by multiple institutions, and then uses the federated averaging algorithm for preliminary model training. By analyzing the gradient contributions of the data of each institution to the global model, combining data quality and sample diversity, the entropy value evaluation method is used to calculate the comprehensive contribution score. Based on this score, the present invention designs a corresponding incentive allocation mechanism, and optimizes the performance of low-contribution institutions by adjusting the data input. Finally, the global model is iteratively updated to realize the closed-loop of data sharing and analysis. This method not only effectively quantifies the contributions of all participating parties in federated learning, but also establishes a fair and reasonable incentive mechanism, promotes multi-party collaboration, and improves the overall performance and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0086] Figure 1Schematic flowchart of the method according to the embodiments of the present invention. Detailed implementation manners

[0087] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0088] It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0089] As Figure 1 shown, a non-alcoholic fatty liver disease data sharing and analysis method based on federated learning is provided in this embodiment, including:

[0090] Obtain the non-alcoholic fatty liver disease-related data quality indicators and data quantity statistics uploaded by each institution, and clean and normalize the original data through a preset standardization process to obtain a dataset in a unified format;

[0091] According to the dataset in the unified format, initially train the local models of each institution using the federated averaging algorithm to obtain the preliminary parameters of the global model, and judge the preliminary accuracy performance of the global model on the data of each institution;

[0092] According to the preliminary accuracy performance, obtain the gradient contributions of the data of each institution to the global model, and determine the specific contribution rates of each institution in the model optimization by calculating the magnitude and direction of the gradient vectors;

[0093] Analyze the relationship between the gradient contributions and the data quality. If the data quality of a certain institution is lower than the preset threshold, the contribution rate is weighted and adjusted to obtain the adjusted contribution rate;

[0094] According to the adjusted contribution rates, obtain the data quantity and sample diversity characteristics of each institution, and use an entropy-based evaluation method to judge the weight of the role of the data of each institution in enhancing the diversity of the global model;

[0095] According to the role weights and the adjusted contribution rates, use the linear weighted fusion method to calculate the comprehensive contribution scores of each institution in the model accuracy improvement, and obtain the final contribution quantification result;

[0096] Through the comprehensive contribution scores, use a preset incentive allocation algorithm to allocate resources or benefits to each institution to obtain the output result of the incentive system in the distributed cooperation;

[0097] According to the output results of the incentive system, if it is judged that the comprehensive contribution score of a certain institution is lower than the average value, then adjust the data input of the institution by increasing the local data sampling rate or introducing external synthetic data, and track and judge whether the subsequent contribution is improved;

[0098] According to the adjusted data input, use the federated learning framework to iteratively update the global model to obtain optimized model parameters and accuracy performance, and achieve real-time data sharing and analysis.

[0099] Furthermore, the process of obtaining the quality indicators and data quantity statistics related to non-alcoholic fatty liver disease uploaded by each institution, and cleaning and normalizing the original data through a preset standardization process to obtain a dataset in a unified format includes:

[0100] Obtain the original non-alcoholic fatty liver disease-related data uploaded by each institution and the corresponding quality indicators and quantity statistics, judge whether the data quality reaches the preset threshold, if it is lower than the threshold, then preliminarily screen the original non-alcoholic fatty liver disease-related data to obtain the first dataset;

[0101] Clean the first dataset through a preset standardization process to remove outliers and duplicates to obtain the second dataset;

[0102] Process the second dataset using a data normalization method to unify the data range and format to obtain the third dataset;

[0103] Extract feature variables from the third dataset, combine the quality indicators and quantity statistics, and judge whether the features are suitable for distributed modeling. If they are suitable, generate the fourth dataset;

[0104] For the fourth dataset, apply the distributed modeling algorithm K-means for clustering analysis to obtain the clustering results;

[0105] According to the clustering results and quality indicators, judge whether the output of each institution's local model is stable. If it is not stable, then adjust the normalization parameters and regenerate the fifth dataset;

[0106] Perform distributed modeling again through the fifth dataset, using the support vector machine algorithm, to obtain the final dataset in a unified format.

[0107] Specifically, first, obtain data quality metrics and data quantity statistics from various institutions through a data interface. For example, the medical visit data obtained from a hospital contains 1 million records, each record has 10 fields, and the proportion of missing values is 3%. Then, use a data quality assessment algorithm, such as a rule-based method, to perform a preliminary analysis of the data, identify outliers and duplicate data. For example, use the Z-score method to detect that 5% of the outliers exist in the fields of medical visit records. Next, adopt a preset standardized processing process to clean the original data, including filling missing values, handling outliers, and unifying data formats. For example, use the mean filling method to fill the missing values. Then, through a normalization algorithm, such as min-max normalization, convert the data to the range of 0 to 1 to ensure that the numerical ranges of all fields are consistent. Finally, store the cleaned and normalized data set in a unified format, such as a CSV file, for subsequent distributed modeling use.

[0108] Furthermore, based on the data set in the unified format, the process of initially training the local models of each institution using the Federated Averaging algorithm to obtain the preliminary parameters of the global model and judging the preliminary accuracy performance of the global model on the data of each institution includes:

[0109] Use the Federated Averaging algorithm to initially train the local models to obtain the training results of each institution;

[0110] Extract the preliminary parameters from the training results of each institution to determine the initial state of the global model;

[0111] Perform running inference on the institutional data based on the global model to obtain the preliminary accuracy performance;

[0112] If the accuracy performance is lower than the preset threshold, perform a new round of training on the local models to update the preliminary parameters;

[0113] Adjust the global model according to the updated preliminary parameters and judge the changing trend of the accuracy performance of the data of each institution;

[0114] Optimize the Federated Averaging algorithm through trend analysis to obtain the final global model parameters.

[0115] Specifically, when initially training the local models of each institution using the Federated Averaging algorithm, first divide the data of each institution into a training set and a test set with proportions of 80% and 20% respectively. Each institution uses local data to train the model, and the Stochastic Gradient Descent algorithm is adopted with a learning rate set to 0.1 and the number of iterations to 100. After training, each institution uploads the local model parameters to the central server, and the central server calculates the global model parameters using the weighted average method, and the weights are allocated according to the data volume of each institution.

[0116] For example, if the data volumes of institutions A, B, and C are 10,000, 8,000, and 12,000 respectively, then the weights are 33, 27, and 40 respectively. After the global model parameters are updated, they are distributed to each institution for synchronous update of the local model. Subsequently, each institution evaluates the preliminary accuracy performance of the model on the local test set. Suppose the accuracies of the models of institutions A, B, and C on the test set are 85%, 88%, and 83% respectively, then the average accuracy of the global model is 85.3%. By analyzing the accuracy differences of the models of each institution, the impact of the heterogeneity of data distribution on the model performance can be preliminarily judged, providing a basis for subsequent federated learning optimization.

[0117] Furthermore, the process of obtaining the gradient contribution of each institution's data to the global model according to the preliminary accuracy performance and determining the specific contribution rate of each institution in model optimization by calculating the magnitude and direction of the gradient vector includes:

[0118] Collect the data information uploaded by each institution through a preset interface to obtain the initial value of the gradient contribution;

[0119] By calculating the vector norm of the initial value of the gradient contribution and using the Euclidean norm for processing, the norm result is obtained;

[0120] According to the norm result, calculate the direction of the gradient vector and use the arctangent function to determine the direction angle;

[0121] According to the direction angle and the norm result, obtain the gradient distribution characteristics of the global model to get the optimization reference value;

[0122] If the optimization reference value exceeds the preset threshold, then adjust the global model through the gradient descent method to obtain the updated parameters;

[0123] According to the updated parameters, calculate the contribution rate value of each institution's data to model optimization and use the weighted average method to determine the final contribution rate;

[0124] Through the final contribution rate, obtain the gradient impact ranking of each institution in the global model to get the optimization priority.

[0125] Furthermore, analyze the relationship between the gradient contribution and the data quality. If the data quality of a certain institution is lower than the preset threshold, the process of weighted adjustment of the contribution rate to obtain the adjusted contribution rate includes:

[0126] Obtain the data quality parameter of the institution's data, judge whether the data quality is lower than the preset threshold. If so, use the weighted adjustment method to process the contribution rate to get the preliminary adjustment value;

[0127] According to the preliminary adjustment value, combined with the gradient contribution data, determine the adjusted contribution rate;

[0128] Judge whether the adjusted contribution rate meets the preset threshold requirements through quality analysis to obtain the verification result;

[0129] Obtain the correlation coefficient between gradient contribution and data quality according to the verification result, and generate a coefficient matrix;

[0130] Adopt a linear regression algorithm to extract key influencing factors from the coefficient matrix to obtain the final index;

[0131] Generate result data through the final index to complete the calculation of the adjusted contribution rate.

[0132] Specifically, when analyzing the relationship between gradient contribution and data quality, it is first necessary to set a preset threshold for data quality. For example, set the comprehensive score of data integrity and accuracy to 8. Calculate the data quality score of each institution through the data quality evaluation algorithm.

[0133] For example, the data quality score of an institution is 7.5, which is lower than the preset threshold. Next, adopt a weighted adjustment algorithm to adjust the gradient contribution rate of this institution. Assume that the original contribution rate of this institution is 3. According to the difference between the data quality score and the threshold (7.5 - 8 = -0.5), set the adjustment coefficient to 9. By multiplying the original contribution rate by the adjustment coefficient, the adjusted contribution rate index is obtained as 27. This process ensures that when the data quality is low, the contribution rate of this institution can be reasonably reduced, thus more accurately reflecting its actual contribution. In addition, by introducing a dynamic adjustment mechanism for data quality and contribution rate, the accuracy and reliability of the overall analysis result can be further improved, providing a more scientific basis for subsequent decision-making.

[0134] Furthermore, according to the adjusted contribution rate, obtaining the data quantity and sample diversity characteristics of each institution, and adopting an entropy-based evaluation method to judge the process of the weight of the data of each institution on the improvement of the global model diversity includes:

[0135] Obtain the data quantity and sample characteristics submitted by each institution, and obtain the preliminary distribution result through statistical analysis;

[0136] Adopt an entropy calculation method to process the preliminary distribution result to determine the diversity index of the data of each institution;

[0137] Compare the diversity index with the existing characteristics of the global model to judge the incremental impact of the data of each institution on the model diversity;

[0138] Obtain the incremental impact result, and adopt a weighted average method to calculate the weight of the data of each institution;

[0139] If the weight exceeds the preset threshold, include the data of the corresponding institution in the model training dataset;

[0140] Run the random forest algorithm based on the updated training dataset to obtain an optimized global model;

[0141] Output the final contribution rate adjustment results of each institution's data according to the optimized global model.

[0142] Specifically, under the adjusted contribution rate index framework, first obtain the data quantity of each institution. For example, institution A provides 10,000 pieces of data, institution B provides 8,000 pieces of data, and institution C provides 12,000 pieces of data. Then, calculate the sample diversity characteristics of each institution's data, using an entropy-based evaluation method. The specific algorithm is to calculate the information entropy of each feature. For example, for feature 1 of institution A, its information entropy is 2, and for feature 2 it is 5; for feature 1 of institution B, the information entropy is 3, and for feature 2 it is 4; for feature 1 of institution C, the information entropy is 1, and for feature 2 it is 6. Then, according to the data quantity and sample diversity characteristics of each institution, calculate its weight for enhancing the diversity of the global model, using the weighted average method. For example, the weight of institution A is 35, the weight of institution B is 30, and the weight of institution C is 35. Finally, by comparing the weights of each institution, it is concluded that the data of institution C has the greatest effect on enhancing the diversity of the global model, followed by institution A and institution B. This process ensures the reasonable distribution of each institution's data in the global model and improves the diversity and accuracy of the model.

[0143] Furthermore, according to the weight of the effect and the adjusted contribution rate, using the linear weighted fusion method, the process of calculating the comprehensive contribution score of each institution in improving the model accuracy and obtaining the final contribution quantification result includes:

[0144] Obtain the diversity weight data submitted by each institution, use statistical tools to calculate the initial weight value, and obtain the diversity effect weight;

[0145] Process the initial contribution rate data through a preset adjustment rule to obtain the adjusted contribution rate;

[0146] Use the linear weighted fusion method to combine the diversity effect weight and the adjusted contribution rate, and calculate the comprehensive contribution score of each institution;

[0147] If the comprehensive contribution score exceeds the preset threshold, then judge that the institution has a significant effect on improving the model accuracy and determine the preliminary quantification result;

[0148] According to the ranking of the comprehensive contribution scores, obtain the priority sequence of the institution's role and determine the contribution distribution of the accuracy improvement;

[0149] Process the priority sequence and the comprehensive contribution score through regression analysis to obtain the final quantification result;

[0150] For the final quantization results, cluster analysis is used to divide the institutional contribution levels and obtain the hierarchical contribution calculation results.

[0151] Specifically, by analyzing the contributions of each institution to the improvement of model accuracy, the diversity effect weights of each institution are first calculated using the principal component analysis method. Assume that the weights of institutions A, B, and C are 4, 3, and 3 respectively. Then, according to the performance of each institution in model training, the adjusted contribution rates are calculated. Assume that the contribution rates of institutions A, B, and C are 5, 3, and 2 respectively. Using the linear weighted fusion method, the diversity effect weights are multiplied by the adjusted contribution rates and summed to obtain the comprehensive contribution scores of each institution.

[0152] For example, the score of institution A is 4×5 = 20, the score of institution B is 3×3 = 9, and the score of institution C is 3×2 = 6. Finally, according to the comprehensive contribution scores, the quantization results of each institution in the improvement of model accuracy are determined. The contribution quantization values of institutions A, B, and C are 20, 9, and 6 respectively. Through this process, the contributions of each institution in the improvement of model accuracy can be clearly quantified, providing a basis for subsequent model optimization.

[0153] Furthermore, through the comprehensive contribution scores, using a preset incentive allocation algorithm to allocate resources or benefits to each institution, the process of obtaining the output results of the incentive system in distributed collaboration includes:

[0154] Through the comprehensive contribution scores, using a preset algorithm to standardize the data input to obtain the normalized contribution values;

[0155] According to the normalized contribution values and combined with the number of institutions, calculate the resource allocation ratio of each institution to determine the preliminary allocation plan;

[0156] If the resource allocation ratio exceeds the preset threshold, optimize the calculation process through adjusting the algorithm rules to obtain the balanced allocation ratio;

[0157] Using the balanced allocation ratio and combined with the distributed structure, generate the benefit allocation data of each institution;

[0158] According to the benefit allocation data and the requirements of the incentive mechanism, adjust the system balance parameters to determine the final incentive output results.

[0159] Specifically, in distributed collaboration, first evaluate each institution through the comprehensive contribution scores and use a preset incentive allocation algorithm to allocate resources or benefits.

[0160] For example, assume there are three institutions A, B, and C, with their comprehensive contribution scores being 80, 90, and 70 respectively. Using the linear weighted algorithm, multiply the comprehensive contribution scores by the preset weight coefficients. Assuming the weight coefficients are 0.4, 0.3, and 0.3 respectively, the weighted score of institution A is 80 * 0.4 = 32, the weighted score of institution B is 90 * 0.3 = 27, and the weighted score of institution C is 70 * 0.3 = 21. Next, based on the weighted scores, resources or benefits are allocated. Assuming the total resources are 100 units, institution A is allocated 40 units of resources, institution B is allocated 37.5 units of resources, and institution C is allocated 22.5 units of resources. In this way, the fairness and incentive of resource or benefit allocation are ensured, further promoting the active participation and efficient cooperation of each institution in distributed collaboration.

[0161] Furthermore, according to the output result of the incentive system, if it is determined that the comprehensive contribution score of a certain institution is lower than the average value, then by increasing the local data sampling rate or introducing external synthetic data, adjusting the data input of the institution, and the process of tracking and judging whether the subsequent contribution is improved includes:

[0162] According to the output result of the incentive system, calculate the comprehensive contribution scores of each institution, determine the average value through mean calculation, and obtain a list of institutions with scores lower than the average value;

[0163] Extract the local data from the institutions with scores lower than the average value and analyze the sampling rate. Adjust the data input by increasing the sampling frequency to generate a frequency-adjusted data set;

[0164] According to the frequency-adjusted data set, if the local data is insufficient, introduce external synthetic data for supplementation to obtain a supplemented data set;

[0165] Run the incentive system again through the supplemented data set, obtain the new output result, and calculate the adjusted comprehensive contribution score;

[0166] If the adjusted comprehensive contribution score is lower than the average value, then analyze the correlation between the data input and the subsequent contribution through the random forest algorithm to determine the key influencing features;

[0167] According to the key influencing features, use the gradient boosting algorithm to optimize the weight allocation of the data input to generate a weight allocation data set;

[0168] Run the incentive system again through the weight allocation to obtain the final output result, and judge whether the subsequent contribution is improved.

[0169] Specifically, after obtaining the output result of the incentive system, the system first calculates the average value of the comprehensive contribution scores of all institutions. Assume the average value is 75 points. If the score of a certain institution is 60 points, which is lower than the average value, the system activates the data adjustment mechanism. First, the system analyzes the historical data distribution of this institution through algorithms and finds that its local data sampling rate is 10%, which is much lower than the average sampling rate of 30% of other institutions. Therefore, the system gradually increases the local data sampling rate to 25% and introduces external synthetic data to supplement its data input.

[0170] For example, use a generative adversarial network (GAN) to generate 1000 synthetic data to ensure data diversity. Next, the system re-evaluates the comprehensive contribution score of this institution. Using a weighted average algorithm, the local data and synthetic data are weighted at a ratio of 7:3, and the new score is 68 points. The system further analyzes and finds that the quality of the synthetic data affects the score improvement. Therefore, it optimizes the synthetic data generation algorithm and adds a data verification step to ensure the authenticity of the synthetic data. After re-evaluation, the score is increased to 72 points, approaching the average value. The system continuously monitors the contribution changes of this institution. If the score still does not meet the standard, the above process will be repeated until its contribution reaches or exceeds the average value.

[0171] Furthermore, according to the adjusted data input, using the federated learning framework to iteratively update the global model, the process of obtaining optimized model parameters and accuracy performance for real-time data sharing and analysis includes:

[0172] Obtain the adjusted data input through the federated learning framework, iteratively update the global model, and obtain the preliminary model parameters;

[0173] Extract features from the preliminary model parameters, use the gradient descent algorithm to optimize the model parameters, and obtain the optimized model parameters;

[0174] Evaluate the optimized model parameters, calculate the accuracy performance, and determine the performance indicators of the current model;

[0175] If the accuracy performance is lower than the preset threshold, readjust the data input through the federated learning framework and iteratively update the global model;

[0176] According to the adjusted data input and the updated global model, obtain the new model parameters and accuracy performance;

[0177] Integrate the calculation results of all parties through the data sharing mechanism to achieve real-time data sharing and analysis.

[0178] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for sharing and analyzing non-alcoholic fatty liver disease data based on federated learning, characterized in that Including: Obtain the data quality indicators and data quantity statistics related to non-alcoholic fatty liver disease uploaded by each institution, and clean and normalize the original data through a preset standardized processing flow to obtain a dataset in a unified format; According to the dataset in the unified format, use the Federated Averaging algorithm to initially train the local models of each institution to obtain the preliminary parameters of the global model, and judge the preliminary accuracy performance of the global model on the data of each institution; According to the preliminary accuracy performance, obtain the gradient contribution of the data of each institution to the global model, and determine the specific contribution rate of each institution in model optimization by calculating the magnitude and direction of the gradient vector; Analyze the relationship between the gradient contribution and the data quality. If the data quality of a certain institution is lower than the preset threshold, then weight-adjust the contribution rate to obtain the adjusted contribution rate; According to the adjusted contribution rate, obtain the data quantity and sample diversity characteristics of each institution, and use the entropy-based evaluation method to judge the role weight of the data of each institution in improving the diversity of the global model; According to the role weight and the adjusted contribution rate, use the linear weighted fusion method to calculate the comprehensive contribution score of each institution in improving the model accuracy, and obtain the final contribution quantification result; Through the comprehensive contribution score, use the preset incentive allocation algorithm to allocate resources or benefits to each institution to obtain the output result of the incentive system in distributed cooperation; According to the output result of the incentive system, if it is judged that the comprehensive contribution score of a certain institution is lower than the average value, then adjust the data input of the institution by increasing the local data sampling rate or introducing external synthetic data, and track and judge whether the subsequent contribution is improved; According to the adjusted data input, use the federated learning framework to iteratively update the global model to obtain the optimized model parameters and accuracy performance, and realize real-time data sharing and analysis.

2. The method according to claim 1, wherein The process of obtaining the data quality indicators and data quantity statistics related to non-alcoholic fatty liver disease uploaded by each institution, and cleaning and normalizing the original data through a preset standardized processing flow to obtain a dataset in a unified format includes: Obtain the original non-alcoholic fatty liver disease-related data uploaded by each institution and the corresponding quality indicators and quantity statistics, judge whether the data quality reaches the preset threshold, and if it is lower than the threshold, perform preliminary screening on the original non-alcoholic fatty liver disease-related data to obtain a first dataset; Clean the first dataset through a preset standardized processing flow to remove outliers and duplicates to obtain a second dataset; Process the second dataset using a data normalization method to unify the data range and format to obtain a third dataset; Extract feature variables from the third dataset, combine the quality indicators and quantity statistics, and judge whether the features are suitable for distributed modeling. If they are suitable, generate a fourth dataset; For the fourth dataset, apply the distributed modeling algorithm K-means for clustering analysis to obtain the clustering result; Based on the clustering results and quality metrics, determine whether the outputs of the local models of each institution are stable. If not, adjust the normalization parameters and regenerate the fifth dataset. Execute distributed modeling again using the fifth dataset, and adopt the support vector machine algorithm to obtain the dataset in the final unified format.

3. The method according to claim 1, wherein The process of initially training the local models of each institution using the federated averaging algorithm based on the dataset in the unified format to obtain the preliminary parameters of the global model, and determining the preliminary accuracy performance of the global model on the data of each institution includes: Adopt the federated averaging algorithm to initially train the local models to obtain the training results of each institution; Extract the preliminary parameters from the training results of each institution to determine the initial state of the global model; Perform running inference on the institutional data based on the global model to obtain the preliminary accuracy performance; If the accuracy performance is lower than the preset threshold, conduct a new round of training on the local models and update the preliminary parameters; Adjust the global model according to the updated preliminary parameters and judge the change trend of the accuracy performance of the data of each institution; Optimize the federated averaging algorithm through trend analysis to obtain the final global model parameters.

4. The method according to claim 1, wherein The process of obtaining the gradient contribution of the data of each institution to the global model based on the preliminary accuracy performance, and determining the specific contribution rate of each institution in model optimization by calculating the magnitude and direction of the gradient vector includes: Collect the data information uploaded by each institution through a preset interface to obtain the initial value of the gradient contribution; Calculate the vector norm of the initial value of the gradient contribution and process it using the Euclidean norm to obtain the norm result; Calculate the direction of the gradient vector according to the norm result, and use the arctangent function to determine the direction angle; Obtain the gradient distribution characteristics of the global model based on the direction angle and norm result to obtain the optimization reference value; If the optimization reference value exceeds the preset threshold, adjust the global model through the gradient descent method to obtain the updated parameters; Calculate the contribution rate value of the data of each institution to model optimization according to the updated parameters, and adopt the weighted average method to determine the final contribution rate; Obtain the gradient impact ranking of each institution in the global model through the final contribution rate to obtain the optimization priority.

5. The method according to claim 1, wherein Analyze the relationship between the gradient contribution and data quality. If the data quality of a certain institution is lower than the preset threshold, perform weighted adjustment on the contribution rate to obtain the adjusted contribution rate, and the process includes: Obtain the data quality parameter of the institutional data, judge whether the data quality is lower than the preset threshold. If so, use the weighted adjustment method to process the contribution rate to obtain the preliminary adjustment value; Determine the adjusted contribution rate according to the preliminary adjustment value in combination with the gradient contribution data; Judge whether the adjusted contribution rate meets the requirements of the preset threshold through quality analysis to obtain the verification result; Obtain the correlation coefficient between the gradient contribution and data quality according to the verification result and generate a coefficient matrix; Adopt the linear regression algorithm to extract the key influencing factors from the coefficient matrix to obtain the final index. Generate result data through the final metrics to complete the calculation of the adjusted contribution rate.

6. The method according to claim 1, wherein The process of obtaining the data quantity and sample diversity characteristics of each institution according to the adjusted contribution rate and using the evaluation method based on entropy value to judge the contribution weight of each institution's data to the improvement of the global model diversity includes: Obtain the data quantity and sample characteristics submitted by each institution, and obtain the preliminary distribution result through statistical analysis; Process the preliminary distribution result by using the entropy value calculation method to determine the diversity index of each institution's data; Compare the existing characteristics of the global model through the diversity index to judge the incremental impact of each institution's data on the model diversity; Obtain the incremental impact result, and use the weighted average method to calculate the contribution weight of each institution's data; If the contribution weight exceeds the preset threshold, include the corresponding institution's data in the model training data set; Run the random forest algorithm based on the updated training data set to obtain the optimized global model; Output the final contribution rate adjustment result of each institution's data according to the optimized global model.

7. The method according to claim 1, wherein The process of calculating the comprehensive contribution score of each institution in the improvement of the model accuracy and obtaining the final contribution quantification result by using the linear weighted fusion method according to the contribution weight and the adjusted contribution rate includes: Obtain the diversity weight data submitted by each institution, use statistical tools to calculate the initial weight value, and obtain the diversity contribution weight; Process the initial contribution rate data through the preset adjustment rule to obtain the adjusted contribution rate; Use the linear weighted fusion method to combine the diversity contribution weight and the adjusted contribution rate to calculate the comprehensive contribution score of each institution; If the comprehensive contribution score exceeds the preset threshold, judge that the institution has a significant effect on the improvement of the model accuracy and determine the preliminary quantification result; Sort according to the comprehensive contribution score, obtain the priority sequence of the institution's role, and determine the contribution distribution of the accuracy improvement; Process the priority sequence and the comprehensive contribution score through regression analysis to obtain the final quantification result; For the final quantification result, use cluster analysis to divide the institution contribution levels to obtain the hierarchical contribution calculation result.

8. The method according to claim 1, wherein The process of allocating resources or benefits to each institution through the comprehensive contribution score by using the preset incentive allocation algorithm to obtain the output result of the incentive system in the distributed cooperation includes: Through the comprehensive contribution score, use the preset algorithm to standardize the data input to obtain the normalized contribution value; According to the normalized contribution value and combined with the number of institutions, calculate the resource allocation ratio of each institution to determine the preliminary allocation plan; If the resource allocation ratio exceeds the preset threshold, optimize the calculation process through the adjusted algorithm rule to obtain the balanced allocation ratio; Use the balanced allocation ratio and combine with the distributed structure to generate the income allocation data of each institution; According to the income allocation data and the requirements of the incentive mechanism, adjust the system balance parameters to determine the final incentive output result.

9. The method according to claim 1, wherein According to the output result of the incentive system, if it is judged that the comprehensive contribution score of a certain institution is lower than the average value, the process of adjusting the data input of the institution by increasing the local data sampling rate or introducing external synthetic data and tracking and judging whether the subsequent contribution is improved includes: According to the output result of the incentive system, calculate the comprehensive contribution scores of each institution, determine the average value through mean calculation, and obtain a list of institutions with scores lower than the average value; Extract the local data from the institutions with scores lower than the average value and analyze the sampling rate. Adjust the data input by increasing the sampling frequency to generate a frequency-adjusted data set; According to the frequency-adjusted data set, if the local data is insufficient, introduce external synthetic data for supplementation to obtain a supplemented data set; Run the incentive system again through the supplemented data set, obtain the new output result, and calculate the adjusted comprehensive contribution score; If the adjusted comprehensive contribution score is lower than the average value, analyze the correlation between the data input and the subsequent contribution through the random forest algorithm to determine the key influencing features; According to the key influencing features, use the gradient boosting algorithm to optimize the weight allocation of the data input to generate a weight allocation data set; Run the incentive system again through the weight allocation, obtain the final output result, and judge whether the subsequent contribution is improved.

10. The method according to claim 1, wherein According to the adjusted data input, the process of iteratively updating the global model using the federated learning framework to obtain optimized model parameters and accuracy performance and realizing real-time data sharing and analysis includes: Obtain the adjusted data input through the federated learning framework and iteratively update the global model to obtain preliminary model parameters; Extract features from the preliminary model parameters and use the gradient descent algorithm to optimize the model parameters to obtain optimized model parameters; Evaluate the optimized model parameters, calculate the accuracy performance, and determine the performance indicators of the current model; If the accuracy performance is lower than the preset threshold, readjust the data input through the federated learning framework and iteratively update the global model; According to the adjusted data input and the updated global model, obtain new model parameters and accuracy performance; Integrate the calculation results of all parties through the data sharing mechanism to realize real-time data sharing and analysis.