Federal learning-based cross-institution medical data modeling method, apparatus and device, and medium thereof
By acquiring anonymized metadata datasets, calculating scores for each dimension, dynamically adjusting weight values, and establishing an institutional ranking and queuing mechanism, the problem of unreasonable institutional selection in federated learning is solved, improving the efficiency and security of cross-institutional medical data modeling, and realizing the rational allocation of resources and the improvement of model performance.
Patent Information
- Application Number
- CN202511720577.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing federated learning methods suffer from problems such as unreasonable institution selection, low resource allocation efficiency, insufficient data security and privacy protection, and inability to balance multi-dimensional indicators in cross-institutional medical data modeling, which affect model training effectiveness and compliance.
By acquiring anonymized metadata, calculating score vectors for each dimension, determining dynamic weight values based on the performance evaluation results of the federated learning model, dynamically adjusting the weights and calculating the comprehensive score of the institution, establishing an institution ranking and queuing mechanism, rationally allocating resources, and prioritizing data quality and model requirements.
This enables the rational allocation of resources under limited network capacity, balancing the data scale advantages of large institutions with the specialized characteristics of small institutions, improving the performance and data utilization efficiency of federated learning models, and ensuring data security and compliance.
Smart Images

Figure CN121583558A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of federated learning model training, and in particular relates to a method, apparatus, device and medium for cross-institutional medical data modeling based on federated learning. Background Technology
[0002] With the deepening application of federated learning technology in the medical field, cross-institutional collaborative modeling of medical data has become an important way to improve the performance of medical artificial intelligence models. This technology enables medical institutions to jointly train high-quality medical diagnostic models through distributed training, without leaving their local facilities. In traditional technologies, federated learning networks typically use a first-come, first-served or random selection method to determine participating institutions, and some systems also select institutions based on simple single indicators such as the amount of data. This approach is simple to implement, has low computational overhead, and can quickly complete institution selection. However, current federated learning methods have significant limitations. Single-condition ranking cannot comprehensively evaluate the data value of institutions, which may lead to the loss of key case data and affect the model training effect; static selection mechanisms are difficult to adapt to the dynamically changing data needs during model training, resulting in low resource allocation efficiency; the lack of comprehensive consideration of key factors such as data security and privacy protection may bring compliance risks; in terms of fairness, traditional methods also have difficulty balancing the complex relationship between multiple dimensions of indicators such as the number of cases, specialty characteristics, and data scarcity, resulting in the inability to fully realize the value of medical data and ultimately affecting the overall performance of federated learning models. Summary of the Invention
[0003] Therefore, it is necessary to provide a federated learning-based cross-institutional medical data modeling method, device, equipment, and medium that can reasonably allocate access resources, ensure fair and efficient resource allocation, and take into account the actual needs of data modeling, in order to address the above-mentioned technical problems.
[0004] Firstly, this application provides a method for cross-institutional medical data modeling based on federated learning, including:
[0005] Obtain the de-identified meta-dataset; and calculate the score vectors for each dimension based on the de-identified meta-dataset; the dimension score vectors include data security score, privacy protection score, ethical review score, number of cases score, data format score, specialty characteristics score, and data scarcity score;
[0006] Based on the performance evaluation results of the federated learning model, the weight values of the score vectors of each dimension are determined to obtain dynamic weight values;
[0007] Based on the score vectors of each dimension and the dynamic weight values, the comprehensive score of each institution is calculated; and the institutions are ranked based on the comprehensive scores to obtain a ranking list.
[0008] Starting from the top of the ranking list, a predetermined number of institutions are selected to obtain the list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data that participated in the training of the federated learning model in this round.
[0009] Furthermore, after selecting a predetermined number of institutions from the top of the ranking list to obtain the list of selected institutions, it also includes:
[0010] Based on the ranking list and the list of selected institutions, institutions not selected are assigned to queues of different priorities to obtain the complete queuing status.
[0011] Based on the new performance evaluation results, the queuing score of the organization in the complete queuing state is calculated using the queuing score algorithm; and the complete queuing state is updated based on the queuing score to obtain the updated queuing state.
[0012] Based on the updated queuing status, institutions are selected to participate in the new round of training of the federated learning model, resulting in a list of participating institutions.
[0013] Furthermore, based on the ranking list and the list of selected institutions, institutions not selected are assigned to queues of different priorities to obtain the complete queuing status, including:
[0014] By comparing the ranking list and the list of selected institutions, the selected institutions are identified, and the list of non-selected institutions is obtained.
[0015] Based on the ranking list, the allocation threshold for each queue is calculated using the following formula:
[0016]
[0017]
[0018] in, The priority queue score threshold, The threshold for the queue conditions to be improved is N, where N is the total number of institutions not selected. Preset the ratio for the priority queue. Ranked The overall score of the institution. To score for ethical review, Score for data security;
[0019] Based on the allocation threshold, the institutions not selected in the list are assigned to the priority queue and the queue to be improved, and the remaining institutions are assigned to the normal queue, thus obtaining the complete queuing status.
[0020] Furthermore, based on the new performance evaluation results, the queuing score of the organization in the complete queuing state is calculated using a queuing score algorithm, including:
[0021] Based on the improved proof, the institutional dimension score vector is calculated; the improved proof is the proof file that needs to be updated for the institutional-reported dimension score vector used to characterize the institution.
[0022] For each institution in a fully queued state, increment the waiting round count to obtain a waiting round list;
[0023] By weighted summing of the score vectors for the organizational dimension and dynamic weight values, a list of total scores corresponding to the complete queuing state is obtained.
[0024] Based on the new performance evaluation results, the waiting round list, and the total score list, the queuing score is calculated using the following formula:
[0025]
[0026] in, To score points in the queue, The total score is represented by W, where W is the number of waiting rounds. The time decay coefficient, This adds points to the urgency level.
[0027] Furthermore, the urgency bonus is obtained through the following method:
[0028] Based on the performance evaluation results, the required data for the next round of model training is calculated to obtain a list of urgent needs.
[0029] Based on the urgency of improvement of demand data, an urgency score is assigned to each demand data to obtain an urgency mapping table;
[0030] The matching degree between the data characteristics of each institution and the list of urgent needs was analyzed to obtain the matching degree matrix of the corresponding institution; the data characteristics were extracted from the de-identified meta-dataset.
[0031] Based on the matching degree matrix and the urgency mapping table, a weighted summation is performed to obtain the urgency score.
[0032] Furthermore, based on the performance evaluation results of the federated learning model, the weight values of the score vectors for each dimension are determined, resulting in dynamic weight values, including:
[0033] Based on the performance evaluation results, specific areas where federated learning models have obvious defects are identified, and key performance shortcomings are obtained.
[0034] Analyze the correlation between each key performance weakness and each dimension of the corresponding dimensional score vector to obtain a correlation mapping table;
[0035] Based on the association mapping table, the severity scores of each key performance weakness associated with the same dimension are aggregated to obtain the initial weight values for each dimension.
[0036] Based on historical weight values, the initial weight values are adjusted to obtain new weight values; and the new weight values are then normalized to obtain dynamic weight values.
[0037] Furthermore, based on historical weight values, the initial weight values are adjusted to obtain new weight values, including:
[0038] The initial weight values are normalized to obtain the target weight vector;
[0039] Based on the target weight vector and historical weight values, the new weight values are calculated using the following formula:
[0040]
[0041] in, Let i be the new weight value for the i-th dimension. As a smoothing factor, Let be the historical weight value of the i-th dimension. Let be the target weight vector for the i-th dimension.
[0042] Secondly, this application also provides a cross-institutional medical data modeling apparatus based on federated learning, comprising:
[0043] The scoring module is used to obtain the de-identified meta-dataset and calculate the score vectors for each dimension based on the de-identified meta-dataset. The dimension score vectors include data security score, privacy protection score, ethical review score, number of cases score, data format score, specialty characteristics score, and data scarcity score.
[0044] The weighting module is used to determine the weight values of the score vectors of each dimension based on the performance evaluation results of the federated learning model, thus obtaining dynamic weight values.
[0045] The ranking module is used to calculate the comprehensive score of each institution based on the score vectors of each dimension and the dynamic weight values; and to rank the institutions based on the comprehensive scores to obtain a ranking list.
[0046] The selection module is used to select a preset number of institutions starting from the top of the ranking list to obtain the list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data participating in the training of the federated learning model in this round.
[0047] Thirdly, this application also provides a computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement any step of the method provided in the first aspect of this application.
[0048] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of the method provided in the first aspect of this application.
[0049] The aforementioned method, apparatus, equipment, and media for cross-institutional medical data modeling based on federated learning acquire anonymized metadata datasets. Based on these datasets, score vectors for each dimension are calculated. These vectors include data security scores, privacy protection scores, ethical review scores, case quantity scores, data format scores, specialty characteristics scores, and data scarcity scores. Based on the performance evaluation results of the federated learning model, the weights of each dimension's score vectors are determined, resulting in dynamic weight values. Based on the score vectors and dynamic weight values, a comprehensive score for each institution is calculated. Institutions are ranked based on their comprehensive scores, resulting in a ranking list. A predetermined number of institutions are selected from the top of the ranking list to obtain a list of selected institutions. This list represents the institutions corresponding to the data used in the current round of training for the federated learning model. By quantifying the evaluation of the training data reported by institutions, access resources can be rationally allocated even with limited network capacity. By comprehensively considering multiple dimensions and dynamically adjusting weights, the data scale advantages of large institutions and the specialty characteristics and data scarcity value of small institutions can be balanced, constructing a healthy and sustainable collaborative ecosystem and improving the upper limit of federated learning performance. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 A schematic diagram illustrating the process of a cross-institutional medical data modeling method based on federated learning, provided in an embodiment of the present invention;
[0052] Figure 2 This is a schematic diagram of the structure of a cross-institutional medical data modeling device based on federated learning, provided in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0054] In one embodiment, such as Figure 1As shown, a method for cross-institutional medical data modeling based on federated learning is provided. This embodiment illustrates the method by applying it to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0055] Step 101: Obtain the de-identified meta-dataset; and calculate the score vectors for each dimension based on the de-identified meta-dataset; the dimension score vectors include data security score, privacy protection score, ethical review score, number of cases score, data format score, specialty characteristics score, and data scarcity score.
[0056] The desensitized metadata dataset is a collection of data collected from various medical institutions that has undergone desensitization. Desensitization means that personal identification information has been removed from the data to protect privacy. Metadata refers to data that describes the characteristics of the data itself, such as the size, format, source, and creation time of the data, rather than the specific case content. The desensitized metadata dataset is used to evaluate the data quality of various institutions without involving sensitive information, thus ensuring compliance. The Dimensional Score Vector is a vector containing scores across multiple dimensions. Each score represents an institution's data quality assessment result on a specific dimension. The Dimensional Score Vector includes at least seven scores: Data Security Score, which assesses the level of security measures implemented by the medical institution during data storage, transmission, and processing, such as encryption strength and access control; a higher score indicates more secure data. Privacy Protection Score, which assesses the institution's compliance with privacy protection during data use, such as whether it complies with privacy regulations; a higher score indicates better privacy protection. Ethical Review Score, which assesses whether the institution has passed ethics committee review to ensure that data use complies with ethical standards, such as informed consent and research rationality; a higher score indicates stronger ethical compliance. Case Quantity Score, calculated based on the amount of case data provided by the institution; generally, more cases result in a higher score, representing a larger data scale. Data Format Score, which assesses the standardization and consistency of data formats, such as whether a unified data model is used and the degree of data cleaning; a higher score indicates a more standardized format. Specialty Specialty Score, which assesses the institution's data's distinctiveness or expertise in a specific medical specialty; a higher score indicates that the institution's data has representative or scarce value in a particular field. Data scarcity score assesses the rarity of an institution's data within the overall dataset. For example, data on certain rare diseases may score higher; a higher score indicates more unique and difficult-to-obtain data. The terminal collects pre-processed, anonymized metadata datasets from various medical institutions, containing only metadata and excluding patient privacy. Based on these datasets, scores are calculated for each institution across the seven dimensions mentioned above. For instance, the data security score might examine the institution's security certifications and audit reports, converting them into numerical scores. The case quantity score might be directly normalized or segmented based on the data volume. Other scores follow similar patterns. Each dimension has specific evaluation criteria. After calculation, each institution receives a dimensional score vector—a list containing the seven scores—to comprehensively reflect its data quality.
[0057] Step 102: Based on the performance evaluation results of the federated learning model, determine the weight values of the score vectors of each dimension to obtain dynamic weight values.
[0058] Specifically, the performance evaluation results of a federated learning model are the assessment results of the current federated learning model on the test or validation set, including metrics such as accuracy, recall, and F1 score. These results are used to identify areas where the model performs well or has weaknesses, helping to understand the model's needs. Dynamic weights are weight vectors corresponding to each dimension in the dimensional score vector, representing the relative importance of each dimension in the overall score calculation. The weights are dynamic, meaning they are adjusted based on each performance evaluation result, rather than being fixed, thus adapting to changing model requirements. The terminal analyzes the latest performance evaluation results of the federated learning model, identifies key performance weaknesses, establishes a mapping between weaknesses and each dimension of the dimensional score vector, assigns an initial weight value to each dimension based on the degree of correlation (stronger correlations result in higher weights), smoothly adjusts the weights by referencing historical weight values to avoid abrupt changes, and finally normalizes the weights to ensure the sum of all weights is 1. The resulting dynamic weight values reflect the aspects of the current model that most need improvement.
[0059] Step 103: Calculate the comprehensive score for each institution based on the score vectors and dynamic weight values of each dimension; and rank the institutions based on the comprehensive scores to obtain a ranking list.
[0060] Specifically, the overall score for each institution is a single numerical score, calculated by weighting and summing the dimensional score vectors with dynamic weight values. The overall score represents the institution's overall data quality level considering the current model requirements. The ranking list is a list in which all institutions are sorted from highest to lowest overall score. The ranking list visually displays the relative strengths and weaknesses of institutions, facilitating the selection of the best participants. The terminal calculates the overall score for each institution by multiplying the score of each dimension in the institution's dimensional score vector by its corresponding dynamic weight value, and then summing all products to obtain the overall score. After calculating the overall scores for all institutions, the institutions are sorted in descending order of score, generating a ranking list. The ranking list includes the institution's identifier and its overall score for use in subsequent steps.
[0061] Step 104: Select a preset number of institutions starting from the top of the ranking list to obtain the list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data participating in the training of the federated learning model in this round.
[0062] The preset number is a pre-defined integer representing the number of institutions to be selected in each training round, determined by the network capacity of the federated learning model. The list of selected institutions is a list containing the preset number of institutions selected starting from the top of the ranking list. Data reported by these institutions will be invited to participate in the training of the federated learning model in this round. The terminal selects the preset number of institutions downwards, starting from the institution with the highest overall score. For example, if the preset number is 5, the top 5 institutions are selected directly. This selection process continues from the first institution in the ranking list until the preset number is reached. The resulting list of selected institutions represents the institutions eligible to contribute data for training in this round.
[0063] This embodiment provides a federated learning-based cross-institutional medical data modeling method, which obtains an anonymized meta-dataset; calculates score vectors for each dimension based on the anonymized meta-dataset; the dimension score vectors include data security score, privacy protection score, ethical review score, case quantity score, data format score, specialty characteristic score, and data scarcity score; based on the performance evaluation results of the federated learning model, determines the weight values of each dimension score vector, obtaining dynamic weight values; calculates the comprehensive score of each institution based on the dimension score vectors and dynamic weight values; ranks the institutions based on the comprehensive scores, obtaining a ranking list; selects a preset number of institutions from the top of the ranking list, obtaining a list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data participating in the training of the federated learning model in this round. Through the above methods, the evaluation and quantification of the training data reported by institutions can be achieved, enabling reasonable allocation of access resources under limited network capacity, comprehensive multi-dimensional indicators and dynamic adjustment of weights, balancing the data scale advantage of large institutions with the specialty characteristics and data scarcity value of small institutions, building a healthy and sustainable collaborative ecosystem, and improving the upper limit of federated learning performance.
[0064] In one embodiment, after selecting a preset number of institutions from the top of the ranking list to obtain the list of selected institutions, the process further includes:
[0065] Step 201: Based on the ranking list and the list of selected institutions, assign the institutions not selected to queues of different priorities to obtain the complete queuing status.
[0066] The ranking list is a list of all institutions sorted from highest to lowest based on their overall score. The selected institutions list is a predetermined number of institutions selected from the top of the ranking list to participate in this round of training. The unselected institutions list is a list of institutions identified by comparing the ranking list and the selected institutions list. These institutions were not selected in this round of training but need to be included in the management system for consideration in subsequent rounds. Different priority queues: These are virtual queues established to manage unselected institutions, including priority queues, ordinary queues, and queues for improvement. Each queue represents a different priority, used to differentiate the treatment of these institutions in subsequent rounds. The complete queuing status is a data structure describing which priority queue each unselected institution is currently assigned to, and the possible queue order information, fully reflecting the current waiting status of all unselected institutions. The terminal performs a comparison operation, comparing the selected institutions list with the ranking list to identify institutions in the ranking list but not in the selected list, forming the unselected institutions list. According to preset rules, these unselected institutions are assigned to different priority queues, typically based on the institution's position in the ranking list and its scores in key dimensions.
[0067] Step 202: Based on the new performance evaluation results, calculate the queuing score of the organization in the complete queuing state using the queuing score algorithm; and update the complete queuing state based on the queuing score to obtain the updated queuing state.
[0068] Specifically, the new performance evaluation result refers to the latest result obtained after the federated learning model has completed one round of training and is re-evaluated, reflecting the model's current performance status and shortcomings. The queuing score algorithm is a specially designed mathematical formula used to calculate a queuing score for each institution in the queue. This algorithm comprehensively considers the institution's static data quality, waiting time, and the urgency of the current model requirements. The queuing score is a value calculated by the queuing score algorithm, used to re-evaluate the priority of institutions in the queue. The higher the score, the more likely the institution should be prioritized for the next round of training. Updating the queuing state is the new state obtained by adjusting the complete queuing state based on the newly calculated queuing score. Adjustments include changing the queue in which the institution belongs and adjusting the order of institutions within the queue. The terminal analyzes the latest performance of the model, identifies the model's key needs or shortcomings, uses the queuing score algorithm to calculate a queuing score for each institution in the complete queuing state, obtains the institution's current data quality indicators, waiting time, and model requirement matching degree, generates a queuing score through a formula, and updates the complete queuing state based on the queuing score.
[0069] Step 203: Based on the updated queuing status, select the institutions to participate in the new round of training of the federated learning model and obtain the list of participating institutions.
[0070] Specifically, the updated queuing status indicates the latest distribution and priority of institutions in the queue. The participating institution list is a list containing institutions selected to participate in a new round of training of the federated learning model, including institution identifiers for direct configuration of the training process. The terminal views the high-priority queue in the updated queuing status and selects institutions sequentially from high priority to low priority until a preset number is reached. It prioritizes selecting all institutions from the highest priority queue; if the number is insufficient, it continues selecting from the next highest priority queue until the requirement is met. After selection, a participating institution list is generated, listing all selected institutions.
[0071] This embodiment establishes a fair queuing mechanism to ensure that unselected institutions are not forgotten, but are classified into queues of different priorities based on data quality. This provides a basis for institution selection in subsequent rounds, enabling federated learning to dynamically adjust participating institutions, improving the overall model's inclusiveness and data diversity. Dynamically adjusting the queuing status ensures that institution selection better meets the model's latest needs. Through a queuing scoring algorithm, institutions that can compensate for model shortcomings are given priority, avoiding long waiting times for institutions, thereby improving training efficiency and quality.
[0072] In one embodiment, based on the ranking list and the list of selected institutions, institutions not included in the list are assigned to queues of different priorities to obtain a complete queuing status, including:
[0073] Step 301: Compare the ranking list and the list of selected institutions to identify the selected institutions and obtain the list of non-selected institutions.
[0074] The ranking list is an ordered list containing all participating institutions, sorted from highest to lowest by their overall score. Each institution in the list has its own overall score and unique identifier. The ranking list reflects the relative data quality level of the institutions and serves as the basis for institution selection. It includes institution identifiers and overall scores, allowing for comparison of institution priorities. The selected institution list is a list containing institutions chosen to participate in this round of federated learning model training. It is selected from the top of the ranking list based on a predetermined number. The selected institution list identifies the institutions actually participating in training in the current round and typically includes institution identifiers and other necessary metadata to ensure that the training process only includes high-quality institutions. The unselected institutions are those not included in the selected institution list from the ranking list. These unselected institutions represent potential but currently unused resources. The unselected institution list is a list containing all unselected institutions and is used for subsequent queue allocation. The terminal obtains the ranking list and the list of selected institutions, iterates through each institution in the ranking list, and checks whether the institution appears in the list of selected institutions. If an institution in the ranking list is not in the list of selected institutions, the institution is identified as an unselected institution. All unselected institutions are collected to form an unselected institution list, ensuring that each institution is processed only once.
[0075] Step 302: Based on the ranking list, calculate the allocation threshold for each queue using the following formula:
[0076]
[0077]
[0078] in, The priority queue score threshold, The threshold for the queue conditions to be improved is N, where N is the total number of institutions not selected. Preset the ratio for the priority queue. Ranked The overall score of the institution. To score for ethical review, Score for data security.
[0079] Specifically, the ranking list is a list of institutions sorted in descending order of their overall scores, providing information on their scores and rankings. The allocation threshold is a set of critical values used to determine which queue non-selected institutions are assigned to. The allocation threshold includes two specific thresholds: a priority queue score threshold, a numerical threshold calculated based on the institution's overall score; non-selected institutions with an overall score higher than or equal to the priority queue score threshold may be assigned to the priority queue to ensure that institutions with high data quality are given priority. The improvement queue condition threshold is a Boolean condition threshold based on the institution's ethics review score and data security score. It is defined as an ethics review score below 0.5 or a data security score below 0.5; institutions meeting this condition are assigned to the improvement queue to identify institutions with deficiencies in data compliance. The total number of non-selected institutions is the number of institutions obtained from the list of non-selected institutions, representing the total number of non-selected institutions and used as the position index for calculating the thresholds. The priority queue preset ratio is a preset percentage value representing the proportion of institutions that should be included in the priority queue, used to control the size of the priority queue and ensure that only top-tier non-selected institutions enter this queue. The ethics review score is an organization's score on the ethics review dimension, representing the assessment result of whether the organization's data use complies with ethical standards. The score is standardized between 0 and 1, with lower scores indicating greater ethical issues. The data security score is an organization's score on the data security dimension, representing the assessment result of the organization's data protection measures. Lower scores indicate greater security issues. The terminal obtains the total number of non-selected organizations from the list of non-selected organizations. Based on the ranking list, the priority queue score threshold is calculated using a formula. The conditional threshold for the queue to be improved is defined as a conditional expression for either an ethics review score below 0.5 or a data security score below 0.5. After calculation, the assigned threshold is output.
[0080] Step 303: Based on the allocation threshold, the institutions not selected in the list are allocated to the priority queue and the queue to be improved, and the remaining institutions are allocated to the normal queue to obtain the complete queuing status.
[0081] Specifically, the list of unselected institutions is a list containing all unselected institutions. The priority queue is a high-priority queue used to store unselected institutions with high overall scores. Institutions in the priority queue are given priority consideration for training in subsequent rounds, and their data quality is relatively excellent. The priority queue is a list of institutions sorted by overall score or other indicators. The queue for improvement is a low-priority queue used to store institutions with deficiencies in data security or ethics. Institutions in the queue for improvement need to improve their data quality before being considered, helping the system identify and manage high-risk institutions. The ordinary queue is a medium-priority queue used to store institutions that do not meet either the priority queue criteria or the queue for improvement criteria. Institutions in the ordinary queue have generally average data quality and do not require immediate improvement, but they are not given priority. As the default queue, it ensures that all unselected institutions are covered. The complete queuing state is a data structure representing the overall state of all unselected institutions after they have been assigned to the three queues. It includes the name of each queue, a list of institutions within it, and possible queue attributes. This is used by the system to track and manage the queuing status of institutions, providing a basis for subsequent updates and selections. The terminal iterates through each institution in the list of unselected institutions. For each institution, it checks whether the overall score is greater than or equal to the allocation threshold. If it is, the institution is allocated to the priority queue. If it is not, it checks whether the institution meets the threshold for the queue to be improved. If it does, the institution is allocated to the queue to be improved. For institutions that do not meet the above two conditions, they are allocated to the ordinary queue. During the allocation process, the institutions in the queue are sorted. After the allocation is completed, the three queues are integrated to form a complete queuing state.
[0082] This embodiment ensures the orderly management of institutions by scientifically classifying unselected institutions into queues of different priorities. Institutions in the priority queue are reused first, institutions in the improvement queue are required to improve, and institutions in the normal queue wait normally. This improves resource utilization efficiency and promotes the continuous improvement of the data quality used for model training.
[0083] In one embodiment, based on the new performance evaluation results, a queuing score for the organization in the complete queuing state is calculated using a queuing score algorithm, including:
[0084] Step 401: Calculate the organization dimension score vector based on the improved proof; the improved proof is the proof file that needs to be updated for the dimension score vector reported by the organization to represent the organization.
[0085] An improvement certificate is a document proactively submitted by an institution to demonstrate that its data quality has improved in one or more dimensions. Examples include a new data security certification to improve the data security score; a new approval from the ethics committee to improve the ethics review score; and a newly added list of case data to improve the case quantity score. The improvement certificate serves as the basis for triggering an update to the institution's dimension score vector. The institution's dimension score vector is a vector containing scores from multiple dimensions, used to quantify the institution's data quality. The terminal receives and reviews the improvement certificates submitted by institutions. For each certificate, it verifies its authenticity and validity, and updates the corresponding dimension score of the institution based on the specific improvements demonstrated in the certificate. This ensures that the latest data quality status of institutions in the queue is reflected in a timely and accurate manner, rather than relying on outdated historical scores.
[0086] Step 402: For each institution in a complete queue, increment the waiting round count to obtain a waiting round list.
[0087] Specifically, the complete queuing state includes the distribution of all unselected institutions in priority, normal, and waiting-to-improve queues. The waiting round count is a counter for each institution in the queue, recording the number of consecutive rounds in which that institution has not been selected for training. The waiting round list is a list that records the correspondence between the unique identifier of each queued institution and its current waiting round count. The terminal iterates through each institution in the complete queuing state. For each institution, the existing waiting round count is incremented by 1. If an institution is entering the queue for the first time, its waiting round count starts from 0 and becomes 1 after this operation. All institutions and their updated waiting round counts are then compiled into a clear list.
[0088] Step 403: Sum the weighted score vector of the organization dimension and the dynamic weight value to obtain the total score list corresponding to the complete queuing state.
[0089] The dynamic weights are the weights of each dimension's score, dynamically adjusted based on the latest performance limitations of the federated learning model, reflecting the data dimensions most important in the current training phase. The total score for each institution is a scalar value, obtained by multiplying each dimension's score by its corresponding dynamic weight and then summing the results. It represents the institution's overall data quality score, considering the specific needs of the current model. The total score list is a list containing the total score of each institution in the complete queuing state. The terminal performs a calculation once for each queuing institution, retrieving the institution's dimension score vector and a uniform dynamic weight value, multiplying each dimension's score by its corresponding weight, and then summing all products to obtain the institution's total score. The set of total scores for all institutions constitutes the total score list.
[0090] Step 404: Based on the new performance evaluation results, the waiting round list, and the total score list, calculate the queuing score using the following formula:
[0091]
[0092] in, To score points in the queue, The total score is represented by W, where W is the number of waiting rounds. The time decay coefficient, This adds points to the urgency level.
[0093] Specifically, the new performance evaluation results are the latest performance metrics for federated learning models, used to identify model weaknesses. The waiting rounds list contains the number of waiting rounds for each institution. The total score list contains the total score for each institution. The queuing score is the final calculated score used to determine the institution's priority in the queue, taking into account not only data quality but also waiting time and the urgency of the model. The time decay coefficient is a preset constant greater than 0 used to adjust the influence of waiting time on the queuing score; the larger the value, the faster the score increases for institutions with longer waiting times. The urgency bonus is a bonus calculated based on the new performance evaluation results; if an institution's data characteristics happen to compensate for the model's current major weaknesses, it will receive a higher bonus. The terminal uses a formula to calculate the queuing score for each institution.
[0094] This embodiment ensures that the agency selection strategy takes into account data quality, fairness, and efficiency by generating a comprehensive and balanced queuing score, enabling the federated learning process to both rapidly improve model performance and maintain the enthusiasm of participating agencies.
[0095] In one embodiment, the urgency score is obtained through the following method:
[0096] Step 501: Based on the performance evaluation results, calculate the required data for the next round of model training to obtain an urgent requirements list.
[0097] The performance evaluation results refer to the performance metrics exhibited by the federated learning model after the latest round of training or testing, used to diagnose the model's strengths and weaknesses. Demand data refers to the type or characteristics of data, derived from the analysis of model performance weaknesses, and represents categories of data urgently needed to improve model performance. For example, if the model has an extremely low recognition rate for rare disease A, then data related to rare disease A is a type of demand data; if the model performs poorly in processing image data format B, then image data in format B is another type of demand data. The urgent demand list is a list where each item represents a data category urgently needed by the model in the next round of training. Each item in the list is typically a clear label or description, indicating the priority direction for model improvement. The terminal deeply analyzes the performance evaluation results to identify areas or dimensions where the model's performance is significantly lower than expected, identifying these as key performance weaknesses. These weaknesses are then translated into specific data demands, and all identified demands of this kind are collected to form the urgent demand list.
[0098] Step 502: Based on the urgency of improving the demand data, assign an urgency score to each demand data to obtain an urgency mapping table.
[0099] Specifically, improvement urgency refers to the importance and urgency of each data requirement for improving the overall performance of the model. An urgency score is a numerical score assigned to each data requirement to quantify its urgency; a higher score indicates a more critical and urgent requirement. An urgency mapping table is a mapping structure where the keys are the labels of each data requirement in the urgent requirement list, and the values are the urgency scores assigned to that requirement, establishing a mapping relationship from data requirement to its importance weight. The terminal evaluates each data requirement in the urgent requirement list, assigning an urgency score based on factors such as the severity of the corresponding performance bottleneck and its impact on global metrics. The assignment rule can be automatically calculated. For example, the correlation coefficient between the bottleneck and core metrics is used as the urgency score. An urgency mapping table is generated based on these scores, clearly listing each requirement and its importance.
[0100] Step 503: Analyze the matching degree between the data characteristics of each institution and the list of urgent needs to obtain the matching degree matrix of the corresponding institution; the data characteristics are extracted from the de-identified metadata dataset.
[0101] Specifically, data features are characteristic information extracted from the de-identified metadata dataset to describe the attributes of an institution's data. They are more specific and granular metadata, and for example, may include a list of specialized disease types possessed by the institution, the main data format types, the models of imaging equipment, and the demographic distribution characteristics of cases. Matching degree refers to the degree of consistency between an institution's data features and a specific requirement in the emergency requirement list. It is a quantitative value indicating the extent to which the institution's data can meet that specific requirement. The matching degree matrix is a two-dimensional matrix where rows represent institutions, columns represent requirements in the emergency requirement list, and each element represents the degree of matching between the institution's data and the requirement data. The terminal extracts detailed data features from the de-identified metadata for each institution. For each institution, the data features are compared and analyzed one by one with each requirement in the emergency requirement list. The comparison process can be based on rules, similarity calculations, or machine learning models. Through calculation, a matching score is obtained for each requirement. The matching scores of all institutions with all requirements are organized to form the matching degree matrix.
[0102] Step 504: Based on the matching degree matrix and the urgency mapping table, perform a weighted summation calculation to obtain the urgency score.
[0103] The urgency score is a final calculated value assigned to each organization, comprehensively reflecting the organization's overall potential contribution to addressing all current urgent needs of the model. The terminal calculates a unique urgency score for each organization. For a single organization, it extracts the corresponding row of data in the matching matrix, uses the urgency scores corresponding to each need in the urgency mapping table as weights, and then performs a weighted sum of all matching scores for that row.
[0104] This embodiment calculates an urgency bonus, perfectly combining the breadth and depth of matching institutional data with the strategic importance of various needs. Institutions with data that can most accurately solve the most pressing problems of the model will receive the highest bonus, thus gaining a significant advantage in the queue and improving the relevance of model training.
[0105] In one embodiment, based on the performance evaluation results of the federated learning model, the weight values of each dimension's score vector are determined to obtain dynamic weight values, including:
[0106] Step 601: Based on the performance evaluation results, identify specific areas where the federated learning model has obvious defects and obtain key performance shortcomings.
[0107] The performance evaluation results are a set of performance metrics obtained by the federated learning model after the latest round of training or testing. These metrics include quantitative indicators such as accuracy, recall, F1 score, generalization error, and privacy risk. They comprehensively assess the model's strengths and weaknesses in its current state and serve as the basis for identifying model defects. Key performance bottlenecks (KBLs) are specific domains or tasks where the model exhibits significant deficiencies, identified from the performance evaluation results. KBLs are typically related to specific data types, specialties, or technical dimensions, indicating the areas where the model most needs improvement. The terminal analyzes the performance evaluation results, locating domains where the model performs poorly by comparing the performance metrics with target thresholds. For example, the system calculates the accuracy deviation of the model on different disease diagnosis tasks, marking the domains with the largest deviations as KBLs. The identification process involves clustering techniques to ensure the significance and representativeness of the KBLs. After completion, a list of KBLs is output, with each KBL including a domain description and severity information.
[0108] Step 602: Analyze the correlation between each key performance weakness and each dimension of the corresponding dimension score vector to obtain the correlation mapping table.
[0109] Specifically, the dimensional score vector comprises scores across seven dimensions: data security, privacy protection, ethical review, number of cases, data format, specialty characteristics, and data scarcity. Each dimension score represents an assessment of the institution's data quality in a specific area. The correlation strength indicates the degree of correlation between critical performance weaknesses and each dimension of the dimensional score vector. Correlation strength is a numerical value used to quantify the correlation. For example, if a critical performance weakness is a high risk of model privacy leakage, it will have a high correlation with the privacy protection score dimension; conversely, if a weakness is a low model recognition rate for rare diseases, it will have a high correlation with the data scarcity score dimension. The correlation mapping table is a data structure that records the correlation strength between each critical performance weakness and each dimension score. Rows in the correlation mapping table typically correspond to critical performance weaknesses, columns correspond to the dimensions of the dimensional score vector, and cell values represent the correlation strength. For each key performance weakness, the terminal analyzes the degree of correlation with each dimension of the dimensional score vector. The analysis process is based on predefined rules and statistical methods. Optionally, for weakness models with poor generalization ability, it is judged to be highly correlated with the number of cases score and specialty characteristic score, and a high correlation score is assigned. The correlation degree is calculated for all weakness dimension pairs, and a correlation mapping table is generated.
[0110] Step 603: Based on the association mapping table, aggregate the severity scores of each key performance weakness associated with the same dimension to obtain the initial weight values for each dimension.
[0111] Specifically, the severity score is a quantified value of the severity of each critical performance bottleneck, derived from the performance evaluation results. The severity score indicates the magnitude of the negative impact of the bottleneck on the overall model performance; a higher score indicates a more severe bottleneck. Aggregation is a mathematical operation that combines multiple values into a single composite value. Common aggregation methods include summation, weighted average, or maximum value selection. In this embodiment, it is used to merge the contributions of all bottlenecks associated with the same dimension. The initial weight value is the preliminary weight value for each dimension's score dimension, calculated based on the association mapping table and the severity score. The initial weight value reflects the relative importance of that dimension in addressing model bottlenecks, but it has not yet undergone smoothing. For each dimension, the terminal extracts all associated critical performance bottlenecks and their association degree values from the association mapping table. For each bottleneck, combining its severity score and association degree, the contribution value of that bottleneck to the dimension is calculated. The contribution values of all bottlenecks associated with that dimension are aggregated to obtain the initial weight value for that dimension. Aggregation is performed sequentially for all seven dimensions.
[0112] Step 604: Based on the historical weight values, adjust the initial weight values to obtain new weight values; and normalize the new weight values to obtain dynamic weight values.
[0113] Historical weights are the dimensional weights used in previous training rounds or the rounds before. These historical weights provide continuity and prevent drastic changes in weights due to model performance fluctuations. New weights are intermediate weights obtained by adjusting the initial and historical weights. These new weights balance current model requirements and historical stability, but are not yet normalized. Dynamic weights are the final output weight vector, containing normalized weights for each dimension's score. Dynamic weights are updated after each training round to adapt to model changes. For example, the terminal uses historical weights to smoothly adjust the initial weights using an exponential smoothing algorithm to control the influence of historical weights, avoid abrupt weight changes, and improve system stability. The new weights are then normalized by calculating the sum of all new weights for all dimensions and dividing each new weight by this sum until the sum of all weights equals 1. The resulting normalized weight vector is the dynamic weight vector.
[0114] This embodiment generates stable and adaptive dynamic weight values through historical vector smoothing and normalization, ensuring that the weights respond to the current needs of the model while maintaining smooth changes, accurately locating the weak links of the model, and ensuring that the weight allocation is targeted at the most urgent needs of the model, thereby improving training efficiency.
[0115] In one embodiment, the initial weight value is adjusted based on historical weight values to obtain a new weight value, including:
[0116] Step 701: Normalize the initial weight values to obtain the target weight vector.
[0117] The initial weight values are preliminary weights for each dimension's score, calculated based on the correlation between the model's key performance weaknesses and the dimensions. They reflect the relative importance of each dimension in addressing model deficiencies. However, since these values are obtained by aggregating severity scores, their sum may not be 1, requiring further processing. The target weight vector is a vector formed by normalizing the initial weight values. Each element in the target weight vector corresponds to a normalized weight value for one dimension, representing the ideal weight allocation based on the current model performance weakness analysis, without considering the influence of historical weights. The terminal obtains the initial weight values, which are a list containing weight values for the seven dimensions. The sum of the initial weight values is calculated, and each initial weight value is divided by this sum to obtain the normalized weight value for each dimension, resulting in the target weight vector. This vector directly reflects the relative importance of each dimension under the current model requirements.
[0118] Step 702: Based on the target weight vector and historical weight values, calculate the new weight values using the following formula:
[0119]
[0120] in, Let i be the new weight value for the i-th dimension. As a smoothing factor, Let be the historical weight value of the i-th dimension. Let be the target weight vector for the i-th dimension.
[0121] Specifically, the target weight vector is the normalized initial weight value vector, representing the ideal weights required by the current model. Historical weight values are the dimensional weight values used in previous training rounds or the rounds prior. For each dimension, the historical weight value is a numerical value, providing historical continuity in weight allocation, helping to avoid abrupt weight changes and ensuring system stability. The smoothing factor is a preset constant, ranging from 0 to 1, used to control the influence of historical weight values during the adjustment process: the larger the α value, the greater the influence of historical weight values, and the smoother the weight changes; the smaller the α value, the greater the influence of the target weight vector, and the more sensitive the weight changes. The new weight values are the adjusted weight values for each dimension calculated using the smoothing formula. The new weight values balance the target weight vector and historical weight values, responding to the current model requirements while maintaining historical continuity and avoiding drastic fluctuations. The terminal calculates the new weight values for each dimension sequentially, obtaining the historical and target weight values for that dimension, and then substitutes them into the formula for linear combination.
[0122] This embodiment uses a smoothing formula to make weight changes more stable, preventing large jumps in weight values due to short-term fluctuations in model performance, thus improving the robustness of the system and ensuring the continuity and predictability of the federated learning training process. The new weight values respond to the latest model requirements while inheriting the stability of historical weights.
[0123] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0124] Based on the same inventive concept, this application also provides a federated learning-based cross-institutional medical data modeling apparatus for implementing the federated learning-based cross-institutional medical data modeling method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more embodiments of the federated learning-based cross-institutional medical data modeling apparatus provided below can be found in the limitations of the federated learning-based cross-institutional medical data modeling method described above, and will not be repeated here.
[0125] In one exemplary embodiment, such as Figure 2 As shown, a cross-institutional medical data modeling device 800 based on federated learning is provided, comprising:
[0126] The scoring module 801 is used to obtain the de-identified meta-dataset and calculate the score vectors for each dimension based on the de-identified meta-dataset. The dimension score vectors include data security score, privacy protection score, ethical review score, number of cases score, data format score, specialty characteristics score, and data scarcity score.
[0127] The weight module 802 is used to determine the weight values of the score vectors of each dimension based on the performance evaluation results of the federated learning model, and obtain dynamic weight values.
[0128] The ranking module 803 is used to calculate the comprehensive score of each institution based on the score vectors of each dimension and the dynamic weight values; and to rank the institutions based on the comprehensive scores to obtain a ranking list.
[0129] The selection module 804 is used to select a preset number of institutions starting from the top of the ranking list to obtain the list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data participating in the training of the federated learning model in this round.
[0130] Furthermore, the device also includes a queuing module for:
[0131] Based on the ranking list and the list of selected institutions, institutions not selected are assigned to queues of different priorities to obtain the complete queuing status.
[0132] Based on the new performance evaluation results, the queuing score of the organization in the complete queuing state is calculated using the queuing score algorithm; and the complete queuing state is updated based on the queuing score to obtain the updated queuing state.
[0133] Based on the updated queuing status, institutions are selected to participate in the new round of training of the federated learning model, resulting in a list of participating institutions.
[0134] Furthermore, the queuing module is also used for:
[0135] By comparing the ranking list and the list of selected institutions, the selected institutions are identified, and the list of non-selected institutions is obtained.
[0136] Based on the ranking list, the allocation threshold for each queue is calculated using the following formula:
[0137]
[0138]
[0139] in, The priority queue score threshold, The threshold for the queue conditions to be improved is N, where N is the total number of institutions not selected. Preset the ratio for the priority queue. Ranked The overall score of the institution. To score for ethical review, Score for data security;
[0140] Based on the allocation threshold, the institutions not selected in the list are assigned to the priority queue and the queue to be improved, and the remaining institutions are assigned to the normal queue, thus obtaining the complete queuing status.
[0141] Furthermore, the queuing module is also used for:
[0142] Based on the improved proof, the institutional dimension score vector is calculated; the improved proof is the proof file that needs to be updated for the institutional-reported dimension score vector used to characterize the institution.
[0143] For each institution in a fully queued state, increment the waiting round count to obtain a waiting round list;
[0144] By weighted summing of the score vectors for the organizational dimension and dynamic weight values, a list of total scores corresponding to the complete queuing state is obtained.
[0145] Based on the new performance evaluation results, the waiting round list, and the total score list, the queuing score is calculated using the following formula:
[0146]
[0147] in, To score points in the queue, The total score is represented by W, where W is the number of waiting rounds. The time decay coefficient, This adds points to the urgency level.
[0148] Furthermore, the queuing module is also used for:
[0149] Based on the performance evaluation results, the required data for the next round of model training is calculated to obtain a list of urgent needs.
[0150] Based on the urgency of improvement of demand data, an urgency score is assigned to each demand data to obtain an urgency mapping table;
[0151] The matching degree between the data characteristics of each institution and the list of urgent needs was analyzed to obtain the matching degree matrix of the corresponding institution; the data characteristics were extracted from the de-identified meta-dataset.
[0152] Based on the matching degree matrix and the urgency mapping table, a weighted summation is performed to obtain the urgency score.
[0153] Furthermore, the weight module 802 is also used for:
[0154] Based on the performance evaluation results, specific areas where federated learning models have obvious defects are identified, and key performance shortcomings are obtained.
[0155] Analyze the correlation between each key performance weakness and each dimension of the corresponding dimensional score vector to obtain a correlation mapping table;
[0156] Based on the association mapping table, the severity scores of each key performance weakness associated with the same dimension are aggregated to obtain the initial weight values for each dimension.
[0157] Based on historical weight values, the initial weight values are adjusted to obtain new weight values; and the new weight values are then normalized to obtain dynamic weight values.
[0158] Furthermore, the weight module 802 is also used for:
[0159] The initial weight values are normalized to obtain the target weight vector;
[0160] Based on the target weight vector and historical weight values, the new weight values are calculated using the following formula:
[0161]
[0162] in, Let i be the new weight value for the i-th dimension. As a smoothing factor, Let be the historical weight value of the i-th dimension. Let be the target weight vector for the i-th dimension.
[0163] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of a federated learning-based cross-institutional medical data modeling method as described above.
[0164] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0165] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0166] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A method for modeling cross-institutional medical data based on federated learning, characterized in that, The method includes: Obtain the de-identified meta-dataset; and calculate the score vector for each dimension based on the de-identified meta-dataset; the dimension score vector includes data security score, privacy protection score, ethical review score, number of cases score, data format score, specialty characteristics score, and data scarcity score; Based on the performance evaluation results of the federated learning model, the weight values of the score vectors of each dimension are determined to obtain dynamic weight values; Based on the score vectors of each dimension and the dynamic weight values, calculate the comprehensive score of each institution; and rank the institutions based on the comprehensive scores to obtain a ranking list; A predetermined number of institutions are selected from the top of the ranking list to obtain a list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data participating in the training of the federated learning model in this round.
2. The method of claim 1, wherein, After selecting a preset number of institutions from the top of the ranking list to obtain the list of selected institutions, the process further includes: Based on the ranking list and the list of selected institutions, the institutions not selected from the list of institutions are assigned to queues of different priorities to obtain a complete queuing status; Based on the new performance evaluation results, the queuing score of the institution in the complete queuing state is calculated using a queuing score algorithm; and the complete queuing state is updated based on the queuing score to obtain the updated queuing state. Based on the updated queuing status, the institutions selected to participate in the new round of training of the federated learning model are obtained, resulting in a list of participating institutions.
3. The method according to claim 2, characterized in that, Based on the ranking list and the list of selected institutions, the institutions not included in the list are assigned to queues of different priorities to obtain a complete queuing status, including: By comparing the ranking list and the list of selected institutions, the selected institutions are identified, and a list of non-selected institutions is obtained. Based on the ranking list, the allocation threshold for each queue is calculated using the following formula: in, The priority queue score threshold, The threshold for the queue conditions to be improved is N, where N is the total number of institutions not selected. Preset the ratio for the priority queue. Ranked The overall score of the institution. To score for ethical review, Score for data security; Based on the allocation threshold, the institutions in the list of unselected institutions are allocated to the priority queue and the queue to be improved, and the remaining institutions are allocated to the normal queue, thus obtaining the complete queuing status.
4. The method according to claim 3, characterized in that, Based on the new performance evaluation results, the queuing score of the institution in the complete queuing state is calculated using a queuing score algorithm, including: Based on the improved proof, the institutional dimension score vector is calculated; the improved proof is the proof file reported by the institution that the institutional dimension score vector needs to be updated. For each of the institutions in the complete queuing state, increment the waiting round count to obtain a waiting round list; The weighted summation of the institutional dimension score vector and the dynamic weight value yields a total score list corresponding to the complete queuing state. Based on the new performance evaluation results, the waiting round list, and the total score list, the queuing score is calculated using the following formula: in, To score points in the queue, The total score is represented by W, where W is the number of waiting rounds. The time decay coefficient, This adds points to the urgency level.
5. The method according to claim 4, characterized in that, The urgency score is obtained through the following method: Based on the performance evaluation results, the required data for the next round of model training is calculated to obtain an urgent needs list. Based on the urgency of improving the demand data, an urgency score is assigned to each demand data to obtain an urgency mapping table; The matching degree between the data characteristics of each institution and the list of emergency needs is analyzed to obtain a matching degree matrix for each institution; the data characteristics are extracted from the de-identified metadata dataset. Based on the matching degree matrix and the urgency mapping table, a weighted summation calculation is performed to obtain the urgency score.
6. The method according to claim 1, characterized in that, The performance evaluation results based on the federated learning model determine the weight values of the score vectors for each dimension, resulting in dynamic weight values, including: Based on the performance evaluation results, specific domains where the federated learning model has obvious defects are identified, and key performance shortcomings are obtained. Analyze the correlation between each of the key performance weaknesses and each dimension of the corresponding dimensional score vector to obtain a correlation mapping table; Based on the association mapping table, the severity scores of each key performance weakness associated with the same dimension are aggregated to obtain the initial weight values for each dimension. Based on historical weight values, the initial weight values are adjusted to obtain new weight values; and the new weight values are normalized to obtain dynamic weight values.
7. The method according to claim 6, characterized in that, The process of adjusting the initial weight value based on historical weight values to obtain a new weight value includes: The initial weight values are normalized to obtain the target weight vector; Based on the target weight vector and the historical weight values, the new weight value is calculated using the following formula: in, Let i be the new weight value for the i-th dimension. As a smoothing factor, Let be the historical weight value of the i-th dimension. Let be the target weight vector for the i-th dimension.
8. A cross-institutional medical data modeling device based on federated learning, characterized in that, The device includes: The scoring module is used to obtain the de-identified meta-dataset and calculate the score vector for each dimension based on the de-identified meta-dataset. The score vector for each dimension includes data security score, privacy protection score, ethical review score, number of cases score, data format score, specialty characteristics score, and data scarcity score. The weighting module is used to determine the weight values of the score vectors of each dimension based on the performance evaluation results of the federated learning model, and obtain dynamic weight values. The ranking module is used to calculate the comprehensive score of each institution based on the score vectors of each dimension and the dynamic weight values; and to rank the institutions based on the comprehensive scores to obtain a ranking list. The selection module is used to select a preset number of institutions starting from the top of the ranking list to obtain a list of selected institutions; the list of selected institutions is used to represent the institutions corresponding to the data participating in the training of the federated learning model in this round.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.