Multi-modal federal learning method, system and equipment based on medical large model

By adopting multimodal feature fusion and cross-validation technology in the federated learning of medical big model, the problems of data privacy and model performance in traditional federated learning methods are solved, and global model training with more efficient and generalized capabilities is achieved.

CN119990260APending Publication Date: 2025-05-13PEKING UNION MEDICAL COLLEGE +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510081093.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The traditional medical data federated learning method faces the problems of data privacy leakage, incomplete information or loss, and fails to effectively consider the quality of local model features, resulting in limited model performance.

Method used

The multimodal federated learning method based on medical big models is adopted, and the initial global model is issued to the clients of each medical institution through the server side, local local training of the multimodal feature fusion data set is carried out, local models are obtained, and cross-verified on the server side, high-quality model features are retained, and global models are optimized.

Benefits of technology

It improves the training accuracy and efficiency of the global model, enhances the generalization ability of the model, provides a more comprehensive data perspective, can be applied in complex nursing scenarios of different diseases, and effectively solves user privacy protection problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990260A_ABST
    Figure CN119990260A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal federated learning method, system and device based on a medical large model, and relates to the field of federated learning, and the method comprises the steps: carrying out the local training of a received initial global model in a medical institution client through employing a multi-modal feature fusion data set, obtaining a local model, uploading the local model to a server, and carrying out the cross verification, through evaluation, high-quality model features are reserved to realize global training; stopping learning until the optimized global model reaches a first preset condition, performing medical institution client clustering to obtain a plurality of clusters, and inputting the optimal local model into the corresponding cluster to perform intra-cluster aggregation to obtain a cluster model; and sending the cluster model of each cluster to a medical institution client for local training, and after a second preset condition is satisfied, uploading the secondary optimization local model of each medical institution client to a server for aggregation to obtain an optimal federal global model. According to the method, the model generalization ability can be enhanced while the global model training accuracy and efficiency can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of federated learning, and in particular to a multimodal federated learning method, system and device based on a large medical model. Background Art

[0002] With the rapid development of artificial intelligence and machine learning technologies, data-driven applications of medical data are increasing. By analyzing large quantities of medical data to train models, it can be widely used in disease prediction, diagnosis support, and treatment decision optimization. The effect of model training is closely related to the collective quality of data. However, medical data is highly related to patient privacy and sensitivity. Traditional integrated machine learning methods need to store and process data sets in a centralized manner, facing data privacy leakage and security issues; the single modality data used in traditional federated learning can only provide limited information, resulting in incomplete or lost information, and it integrates the parameter features of each local model together without considering the quality of local model features, resulting in limited model performance. Summary of the invention

[0003] The purpose of this application is to provide a multimodal federated learning method, system and device based on a large medical model, which can ensure the accuracy and efficiency of global model training while enhancing the model generalization ability.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides a multimodal federated learning method based on a large medical model, comprising:

[0006] The initial global model is sent to the clients of each medical institution through the server; the initial global model is a large medical model;

[0007] In any of the medical institution clients, the received initial global model is locally trained using a multimodal feature fusion data set to obtain a local model and upload it to the server;

[0008] Cross-validating multiple local models on the server side, and implementing global training by judging and retaining high-quality model features to obtain an optimized global model;

[0009] When the optimized global model does not meet the first preset condition, the initial global model is updated to the optimized global model, and then the process returns to the step of sending the initial global model to each medical institution client through the server to perform the next federated learning, until the optimized global model meets the first preset condition, the learning is stopped, and the optimized global model is marked as the optimal global model, and the local models corresponding to each of the medical institution clients are marked as the optimal local models;

[0010] Clustering the medical institution clients based on the optimal global model to obtain a plurality of clusters, and then inputting the optimal local models of the medical institution clients into the corresponding clusters for intra-cluster aggregation to obtain a cluster model;

[0011] Sending the cluster model of each cluster to each of the medical institution clients for local training until a second preset condition is met to obtain a secondary optimized local model;

[0012] The secondary optimized local models of each of the medical institution clients are uploaded to the server for aggregation to obtain an optimal federated global model.

[0013] In the second aspect, the present application provides a multimodal federated learning system based on a large medical model, including:

[0014] The sending module is used to send the initial global model to the clients of each medical institution through the server; the initial global model is a large medical model;

[0015] A local multi-round training module is used to: in any of the medical institution clients, use a multimodal feature fusion data set to perform local local training on the received initial global model to obtain a local model and upload it to the server;

[0016] A cross-validation optimal module is used to: perform cross-validation on multiple local models on the server side, and implement global training by judging and retaining high-quality model features to obtain an optimized global model;

[0017] A cyclic learning module, used for: when the optimized global model does not meet the first preset condition, updating the initial global model to the optimized global model, and then returning to the step of sending the initial global model to each medical institution client through the server side to perform the next federated learning, until the optimized global model meets the first preset condition, stopping the learning, and marking the optimized global model as the optimal global model, and marking the local models of each medical institution client as the optimal local models;

[0018] A cluster processing module, used to: cluster the medical institution clients based on the optimal global model to obtain multiple clusters, and then input the optimal local model of each of the medical institution clients into the corresponding cluster for intra-cluster aggregation to obtain a cluster model;

[0019] The sending module is further used to: send the cluster model of each cluster to each of the medical institution clients, and perform local training in the local multi-round training module until a second preset condition is met to obtain a secondary optimized local model;

[0020] The final model determination module is used to aggregate the secondary optimized local models of each of the medical institution clients uploaded to the server to obtain the optimal federated global model.

[0021] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a multimodal federated learning method based on a large medical model.

[0022] According to the specific embodiments provided by the present application, the present application has the following technical effects: the present application provides a method, system and device for multimodal federated learning based on a large medical model, adopts a multimodal feature fusion data set for local model training, and cross-validates multiple local models on the server side, and realizes global training by judging and retaining high-quality model features, which can provide a more comprehensive data perspective while ensuring the accuracy and efficiency of global model training, enhance model accuracy and generalization ability, and can be applied in complex nursing scenarios of different diseases. After obtaining the optimal global model, clustering of medical institution clients is performed to obtain multiple clusters, and then intra-cluster aggregation is performed to obtain a cluster model. This setting takes into account the characteristics of endemic diseases, clustering of medical institution clients with high consistency, and fine-grained characterization of the image of different categories of medical institution clients can be achieved. Finally, the present application fuses multimodal data, breaks the medical institution data island through distributed database training, and solves the problem of user privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A flowchart of a multimodal federated learning method based on a large medical model provided in one embodiment of the present application;

[0025] Figure 2 Schematic diagram of multimodal data fusion of medical institution client provided for this application;

[0026] Figure 3 Schematic diagram of the graph clustering-based federated learning workflow provided for this application;

[0027] Figure 4 A schematic diagram of the clustering of the feature enhancement graph of the medical institution client provided for this application;

[0028] Figure 5A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0030] In order to make the purpose, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0031] In an exemplary embodiment, Figure 1 As shown, a multimodal federated learning method based on a large medical model is provided. The method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In an embodiment of the present application, the method includes the following steps 101 to 107.

[0032] Step 101: The server sends the initial global model to the clients of each medical institution. Specifically, the server transfers the medical global model to the preprocessing module, and randomly initializes all model parameters (or features) in the preprocessing module to obtain the initialized global model f 1 , and then the server will f 1 Pass in the medical institution client as the local initialization model.

[0033] Step 102: In any of the medical institution clients, the received initial global model is locally trained using a multimodal feature fusion data set to obtain a local model and upload it to the server.

[0034] In an application example, Figure 2 As shown, the process of acquiring the multimodal feature fusion data set includes: collecting the multimodal data sets of the institution collected by N medical institution clients The data is passed to the multimodal feature fusion module, and the following steps (21) to (23) are executed to realize feature fusion and obtain local data enhancement features. The local data of the medical institution after feature enhancement is the multimodal feature fusion data set {D1, D2, ···, D i ,···,D N}.

[0035] (21) Obtain text data, voice data and image data from the medical institution service end.

[0036] (22) Feature extraction is performed on the text data, voice data and image data of the medical institution service end to obtain text features, voice features and image features.

[0037] (23) Based on the attention mechanism, multimodal feature fusion is performed on the text features, the speech features and the image features to obtain a multimodal feature fusion data set. Specifically, the attention mechanism is used to explore the correlation and difference between the features of each modality, and attention weights are added to different modalities; the input of the attention mechanism consists of a query (Q), a key (K) and a value (V), and the calculation formula of the attention weight is as follows:

[0038]

[0039] Among them, s(x i ,q) is the probability output of the attention mechanism, and the formula is: s(x i ,q)=V T tanh(Wx i +U q ), V, W, and U are used as parameter matrices to be learned. After obtaining the attention weights, they are applied to the input features.

[0040] In another application example, in the process of using the multimodal feature fusion data set to perform local training on the received initial global model, the following loss function is used:

[0041]

[0042] Among them, L i (f i p ) is f i p The value of the loss function, f i p is the model feature of medical institution client i in round p, D i is the multimodal feature fusion dataset of medical institution client i, k is D i Any multimodal feature fusion data in k (f i p ) is to find the optimal model parameter f * , so that the loss function is minimized; f * The formal expression is:

[0043] f * =argminL i (f i p ).

[0044] Use the gradient descent method to iteratively update F(f i p ) to minimize the local loss, the formula used is as follows:

[0045]

[0046] Where λ is the step size, and λ>0, f i p-1 is the model feature of the medical institution client i in the p-1th round, is the gradient operator. After t iterations, each medical institution client obtains the updated local model f i p , and then input into the cross-validation optimal module on the server side.

[0047] In another application example, the initial global model is a medical big model, that is, the present application uses a medical big model as a federated learning distributed database training method. The medical big model is used to realize image segmentation, disease knowledge reasoning or clinical auxiliary diagnosis, and based on the above-mentioned different functional needs, the model output corresponding to the model input is set. For example, when faced with the need for image segmentation, the model input of the medical big model is multimodal feature fusion data, and the model output of the medical big model is the segmentation result of the image data in the multimodal feature fusion data. For example, when faced with the need for disease knowledge reasoning, the model input of the medical big model is multimodal feature fusion data, and the model output of the medical big model is the inferred disease type or the inferred disease stage (such as negative, positive, etc.).

[0048] Step 103, cross-validate multiple local models on the server side, and implement global training by judging and retaining high-quality model features to obtain an optimized global model. The above-mentioned global training is to aggregate the local models of each medical institution client. Traditional aggregation methods include hard voting and soft voting. In the embodiment of the present application, by judging the model feature quality of the local model of each medical institution client, high-quality features are screened and retained, and low-quality features are rejected. Judging the quality of model features includes filtering, packaging and embedding methods. This embodiment takes the filtering method as an example for detailed description, and includes the following steps:

[0049] First, the server uses the feature modeling of the local model of each medical institution client to perform cross-validation to obtain the score of the corresponding model feature; select q model features with the highest scores, q≤N, to form a high-quality feature set, and the rest constitute a low-quality set. The global model of the p+1th round is as follows:

[0050]

[0051] Among them, f p+1is the optimized global model of the p+1th round, N is the number of medical institution clients; ω is the model feature weight, when the feature quality is excellent, ω is 1, otherwise it is 0; f i p is the model feature of medical institution client i in round p, D i It is a multimodal feature fusion dataset of medical institution client i.

[0052] Step 104, when the optimized global model does not meet the first preset condition, the initial global model is updated to the optimized global model, and then the process returns to the step of sending the initial global model to each medical institution client through the server side to perform the next federated learning, until the optimized global model meets the first preset condition, then the learning is stopped, and the optimized global model is marked as the optimal global model, and the corresponding local models of each medical institution client are marked as the optimal local models.

[0053] In an application example, the optimized global model f of the p+1th round obtained in step 103 is p+1 , is sent to each medical institution client through step 101, and is trained locally in the local multi-round training module to obtain the updated local model f i p+2 , and then the model f i p+2 Upload the cross-validation optimal module to the server to obtain the p+2th round global model f p+2 The above steps are repeated until the optimized global model reaches the optimal threshold and the optimal global model f is obtained. P and the optimal local model f of each medical institution client i P .

[0054] Step 104: based on the optimal global model f P Medical institution clients are clustered, and medical institution clients with consistent characteristics are clustered into one cluster to obtain multiple clusters. Then, the optimal local model of each of the medical institution clients is input into the corresponding cluster for intra-cluster aggregation to obtain a cluster model, that is, an intra-cluster adaptation model.

[0055] In another application example, Figure 3 and Figure 4 As shown, medical institution clients are clustered based on the optimal global model to obtain multiple clusters, including:

[0056] Step (41), sending the optimal global model to each of the medical institution clients through the server, and performing local training to obtain an optimized local model and upload it to the server.

[0057] Step (42), using a multi-layer neural network to perform feature mapping on the multiple optimized local models to obtain a feature mapping graph G i ; Wherein, the model features of each of the optimized local models are converted into the weights of each edge in the corresponding feature map.

[0058] Step (43), performing personalized enhanced feature extraction on the feature map to obtain a feature enhanced map G i ', the function used is:

[0059]

[0060] in, When the value is 1, it means that the feature weights before and after feature map enhancement are valid, and the edge is retained. When the value is 0, it means that the feature weights before and after feature map enhancement are invalid, and the edges are deleted. Side Weights before and after feature map enhancement; δ is the threshold; are the sets of edges connecting the s-1th layer and the sth layer before and after feature map enhancement, 1≤s≤S.

[0061] Step (44), converting the feature enhancement graph into a similarity distance vector, includes executing the following steps on the server side:

[0062] 1) Based on the feature enhancement graph, generate a weighted adjacency matrix W i .

[0063] 2) According to the weighted adjacency matrix W i Adjoint diagonal matrix O i Get the feature enhancement graph Laplacian operator

[0064]

[0065] 3) Calculate the eigenvalue sequence λ of the feature enhancement graph Laplacian operator i And the feature vector sequence. It is calculated using the following formula: ν i for The feature vector sequence of i = {ν i1 ,ν i2 ,···,ν in},λ i ={λ i1 ,λ i2 ,···,λ in}.

[0066] 4) According to the eigenvalue sequence, the eigenvalue similarity is calculated using the following formula:

[0067]

[0068] 5) According to the feature vector sequence, the feature vector similarity is calculated using the following formula, specifically the F-norm between the feature vector sequences of the medical institution client:

[0069]

[0070] 6) Use the following formula for weighting to calculate the global similarity:

[0071] Dis comb (i,j)=αDis eig (i,j)+(1-α)Dis vec (i,j).

[0072] 7) The global similarity of each feature enhancement graph is stored as a similarity distance vector Dis corresponding to the medical institution client i =[Dis i1 ,Dis i2 ,···,Dis ij ,···,Dis iN ]. Among them, Dis ij is the feature enhancement graph G i 'And feature enhancement graph G' j α∈[0,1] is the weight parameter, which is determined by the actual training environment; Dis comb (i,j) is the feature enhancement graph G i 'And feature enhancement graph G' j The global similarity of Dis eig (i,j) is the feature enhancement graph G i 'And feature enhancement graph G' j The similarity of the eigenvalues, Dis vec (i,j) is the feature enhancement graph G i 'And feature enhancement graph G' j The feature vector similarity, λ ik is the feature enhancement graph G i 'The kth eigenvalue, λ jk is the feature enhancement graph G' j The kth eigenvalue in i,kI is the feature enhancement graph G i 'The kIth eigenvector, v j,kI is the feature enhancement graph G' j The kIth eigenvector in .

[0073] Step (45) clusters the plurality of medical institution clients based on the similarity distance vector using a clustering algorithm to obtain a plurality of clusters. Specifically, the clustering algorithm includes K-means clustering, hierarchical clustering, density clustering, and the like.

[0074] The above steps (41) to (45) can be implemented in the client clustering module of the server to obtain the first cluster of medical institution clients, the second cluster of medical institution clients, ..., the jth cluster of medical institution clients, ..., the Jth cluster of medical institution clients. In this application, a total of N medical institution clients are preset, and after training, they are aggregated into J clusters. The number of terminal devices in the jth cluster is N. j ,j∈{1,2,···,i,···,J},

[0075] After the clustering of the medical institution clients is completed, the server or the core medical institution clients in the cluster execute the cluster aggregation module to perform cluster aggregation on the updated local models of each medical institution in the cluster to obtain the P+1 round cluster model, which is:

[0076]

[0077] in, is the cluster model of the P+1th round, is the medical institution client n corresponding to cluster j j Multimodal feature fusion dataset, is the medical institution client n corresponding to the jth cluster in the P+1th round j The optimal local model, N j is the number of medical institution clients in the jth cluster, N is the number of medical institution clients.

[0078] Step 106, the cluster model of each cluster is sent to each of the medical institution clients for local local training until the second preset condition is met to obtain a secondary optimized local model; during this local local training process, the above steps of "inputting the optimal local model of each of the medical institution clients into the corresponding cluster for intra-cluster aggregation to obtain a cluster model; sending the cluster model of each cluster to each of the medical institution clients for local local training" are repeated until the threshold is reached or the preset training time is met, and the training is stopped.

[0079] Step 107, uploading the secondary optimized local models of the clients of each medical institution to the server for aggregation to obtain the optimal federated global model.

[0080] In summary, this application takes endemic diseases as a consideration and uses graph clustering-based federated learning as an example. The graph clustering-based federated learning recorded in this application is suitable for grouping medical institution clients into different clusters based on the characteristics of endemic disease data, and further realizing distributed database training. Its application scenario is that the number of patients with primary hypertension in northern China is higher than that in southern China. Dividing the federated training clients into two clusters, northern and southern, can make the final results more realistic and accurate.

[0081] Compared with the prior art, this application also has the following advantages:

[0082] (1) This application provides a more comprehensive information perspective by integrating multi-modal medical data such as text, voice, and images, improves the accuracy and robustness of the federated training model of the medical large model, and can avoid the loss of information in a single modality data.

[0083] (2) This application can significantly improve the performance and generalization ability of the global model by judging the quality of model features uploaded by the medical institution client and integrating only high-quality features to complete the global model iteration, so that the model of this method can achieve better results in complex nursing scenarios of different diseases, and set off a new wave of artificial intelligence ecology for the nursing industry.

[0084] (3) This application takes into account the characteristics of endemic diseases and clusters medical institution clients with high consistency through a graph clustering multimodal federated learning method to achieve a fine-grained characterization of the images of different categories of medical institutions.

[0085] Based on the same inventive concept, the embodiment of the present application also provides a multimodal federated learning system based on a large medical model. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more system embodiments provided below can refer to the limitations of the method above, and will not be repeated here.

[0086] The present application discloses a multimodal federated learning system based on a large medical model, including a sending module, a local multi-round training module, a cross-validation optimal module, a cyclic learning module, a cluster processing module and a final model determination module. Among them, the sending module, the cross-validation optimal module, the cyclic learning module, the cluster processing module and the final model determination module are set on the server side, and the local multi-round training module is set on the client side of the medical institution.

[0087] The sending module is used to send the initial global model to the clients of various medical institutions through the server.

[0088] The local multi-round training module is used to: in any of the medical institution clients, use a multimodal feature fusion data set to perform local local training on the received initial global model to obtain a local model and upload it to the server.

[0089] The cross-validation optimal module is used to: cross-validate multiple local models on the server side, and implement global training by judging and retaining high-quality model features to obtain an optimized global model.

[0090] The cyclic learning module is used to: when the optimized global model does not meet the first preset condition, update the initial global model to the optimized global model, and then return to the step of sending the initial global model to each medical institution client through the server to carry out the next federated learning, until the optimized global model meets the first preset condition, stop learning, and mark the optimized global model as the optimal global model, and mark the corresponding local model of each medical institution client as the optimal local model.

[0091] The cluster processing module is used to: cluster the medical institution clients based on the optimal global model to obtain multiple clusters, and then input the optimal local model of each of the medical institution clients into the corresponding cluster for intra-cluster aggregation to obtain a cluster model.

[0092] The sending module is also used to: send the cluster model of each cluster to each of the medical institution clients, perform local training in the local multi-round training module, and obtain a secondary optimized local model when the second preset condition is met.

[0093] The final model determination module is used to aggregate the secondary optimized local models of each of the medical institution clients uploaded to the server to obtain the optimal federated global model.

[0094] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multimodal federated learning method based on a large medical model is implemented.

[0095] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0096] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0097] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0099] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0100] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0101] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0102] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A multimodal federated learning method based on a large medical model, characterized in that: The multimodal federated learning method based on a large medical model includes: The initial global model is sent to the clients of each medical institution through the server; the initial global model is a large medical model; In any of the medical institution clients, the received initial global model is locally trained using a multimodal feature fusion data set to obtain a local model and upload it to the server; Cross-validating multiple local models on the server side, and implementing global training by judging and retaining high-quality model features to obtain an optimized global model; When the optimized global model does not meet the first preset condition, the initial global model is updated to the optimized global model, and then the process returns to the step of sending the initial global model to each medical institution client through the server to perform the next federated learning, until the optimized global model meets the first preset condition, the learning is stopped, and the optimized global model is marked as the optimal global model, and the local models corresponding to each of the medical institution clients are marked as the optimal local models; Clustering the medical institution clients based on the optimal global model to obtain a plurality of clusters, and then inputting the optimal local models of the medical institution clients into the corresponding clusters for intra-cluster aggregation to obtain a cluster model; Sending the cluster model of each cluster to each of the medical institution clients for local training until a second preset condition is met to obtain a secondary optimized local model; The secondary optimized local models of each of the medical institution clients are uploaded to the server for aggregation to obtain an optimal federated global model.

2. The multimodal federated learning method based on a large medical model according to claim 1 is characterized in that: The process of acquiring the multimodal feature fusion data set includes: Acquiring text data, voice data and image data from the medical institution service end; Extracting features of the text data, voice data and image data of the medical institution service end to obtain text features, voice features and image features; Based on the attention mechanism, multimodal feature fusion is performed on the text features, the speech features and the image features to obtain a multimodal feature fusion data set.

3. The multimodal federated learning method based on a large medical model according to claim 1, characterized in that: In the process of locally training the received initial global model using the multimodal feature fusion dataset, the following loss function is used: Among them, L i (f i p ) is f i p The value of the loss function, f i p is the model feature of the medical institution client i in round p, D i is the multimodal feature fusion dataset of medical institution client i, k is D i Any multimodal feature fusion data in k (f i p ) is to find the optimal model parameter f * , so that the loss function is minimized; f * The formal expression is: f * =argminL i (f i p ); Use the gradient descent method to iteratively update F(f i p ) to minimize the local loss, the formula used is as follows: f i p =f i p-1 -λ▽F(f i p ); Among them, λ is the step size, f i p-1 is the model feature of the medical institution client i in the p-1th round, ▽F(f i p ) is the gradient operator.

4. The multimodal federated learning method based on a large medical model according to claim 1, characterized in that: The function formula of the optimized global model is: Among them, f p+1 is the optimized global model of the p+1th round, N is the number of medical institution clients; ω is the model feature weight; f i p is the model feature of the medical institution client i in round p, D i It is a multimodal feature fusion dataset of medical institution client i.

5. The multimodal federated learning method based on a large medical model according to claim 1, characterized in that: Medical institution clients are clustered based on the optimal global model to obtain multiple clusters, including: The optimal global model is sent to each of the medical institution clients through the server, and local training is performed to obtain an optimized local model and upload it to the server; A multi-layer neural network is used to perform feature mapping on the plurality of the optimized local models to obtain a feature map; wherein the model features of each of the optimized local models are converted into the weights of each edge in the corresponding feature map; Performing personalized enhanced feature extraction on the feature map to obtain a feature enhanced map; Converting the feature enhancement map into a similarity distance vector; Based on the similarity distance vector, a clustering algorithm is used to cluster the plurality of medical institution clients to obtain a plurality of clusters.

6. The multimodal federated learning method based on a large medical model according to claim 5 is characterized in that: Converting the feature enhancement map into a similarity distance vector includes: Based on the feature enhancement graph, generating a weighted adjacency matrix; Obtaining a feature enhanced graph Laplacian operator according to the weighted adjacency matrix adjoint diagonal matrix; Calculating the eigenvalue sequence and eigenvector sequence of the feature enhancement graph Laplacian operator; According to the eigenvalue sequence, the eigenvalue similarity is calculated using the following formula: According to the feature vector sequence, the feature vector similarity is calculated using the following formula: The following formula is used for weighting to calculate the global similarity: Dis comb (i,j)=αDis eig (i,j)+(1-α)Dis vec (i,j); Storing the global similarity of each of the feature enhancement graphs as a similarity distance vector of the corresponding medical institution client; Among them, α∈[0,1] is the weight parameter, Dis comb (i,j) is the feature enhancement graph G i 'And feature enhancement graph G' j The global similarity of Dis eig (i,j) is the feature enhancement graph G i 'And feature enhancement graph G' j The similarity of the eigenvalues, Dis vec (i,j) is the feature enhancement graph G i 'And feature enhancement graph G' j The feature vector similarity, λ ik is the feature enhancement graph G i 'The kth eigenvalue, λ jk is the feature enhancement graph G' j The kth eigenvalue in i,kI is the feature enhancement graph G i 'The kIth eigenvector, v j,kI is the feature enhancement graph G' j The kIth eigenvector in .

7. The multimodal federated learning method based on a large medical model according to claim 5, characterized in that: In the step of performing personalized enhanced feature extraction on the feature map, the function used is: in, When the value is 1, it means that the feature weights before and after feature map enhancement are valid, and the edge is retained. When the value is 0, it means that the feature weights before and after feature map enhancement are invalid, and the edges are deleted. Side Weights before and after feature map enhancement; δ is the threshold; are the sets of edges connecting the s-1th layer and the sth layer before and after feature map enhancement, 1≤s≤S.

8. The multimodal federated learning method based on a large medical model according to claim 1, characterized in that: The cluster model is: in, is the cluster model of the P+1th round, D nj is the medical institution client n corresponding to cluster j j Multimodal feature fusion dataset, is the medical institution client n corresponding to the jth cluster in the P+1th round j The optimal local model, N j is the number of medical institution clients in the jth cluster, N is the number of medical institution clients.

9. A multimodal federated learning system based on a large medical model, characterized in that: The multimodal federated learning system based on the medical big model includes: The sending module is used to send the initial global model to the clients of each medical institution through the server; the initial global model is a large medical model; A local multi-round training module is used to: in any of the medical institution clients, use a multimodal feature fusion data set to perform local local training on the received initial global model to obtain a local model and upload it to the server; A cross-validation optimal module is used to: perform cross-validation on multiple local models on the server side, and implement global training by judging and retaining high-quality model features to obtain an optimized global model; A cyclic learning module, used for: when the optimized global model does not meet the first preset condition, updating the initial global model to the optimized global model, and then returning to the step of sending the initial global model to each medical institution client through the server side to perform the next federated learning, until the optimized global model meets the first preset condition, stopping the learning, and marking the optimized global model as the optimal global model, and marking the local models of each medical institution client as the optimal local models; A cluster processing module, used to: cluster the medical institution clients based on the optimal global model to obtain multiple clusters, and then input the optimal local model of each of the medical institution clients into the corresponding cluster for intra-cluster aggregation to obtain a cluster model; The sending module is further used to: send the cluster model of each cluster to each of the medical institution clients, and perform local training in the local multi-round training module until a second preset condition is met to obtain a secondary optimized local model; The final model determination module is used to aggregate the secondary optimized local models of each of the medical institution clients uploaded to the server to obtain the optimal federated global model.

10. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal federated learning method based on a medical big model as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Federal learning optimization method, device and system, and storage medium

    CN120705612A

  • Federal learning system based on medical care health data lake

    CN121096686A