Medical model training method and device and computer readable storage medium

By using the low-rank adaptation method to perform low-rank decomposition and training on the local model in medical model training, the problem of low-rank training efficiency of medical model is solved, and the effect of improving training efficiency and ensuring privacy is achieved.

CN120106244APending Publication Date: 2025-06-06HANGZHOU YIKANG HUILIAN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510186460.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Medical models are less efficient in training, mainly due to the complexity of the model, the sensitivity and difficulty in obtaining data.

Method used

The local model is trained by low-rank adaptation method. By performing low-rank decomposition of the model gradient, the decomposed local first low-rank matrix and local second low-rank matrix are obtained, and trained on the client to reduce the weights of local data training and improve training efficiency.

Benefits of technology

By reducing the size of the model that needs to be delivered and reducing communication consumption, the overall efficiency of medical model training is improved and the security of patient privacy is guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106244A_ABST
    Figure CN120106244A_ABST
Patent Text Reader

Abstract

The invention provides a medical model training method and device and a computer readable storage medium, and the method comprises the steps: carrying out the low-rank decomposition of the gradient of a local model, and obtaining a local first low-rank matrix and a local second low-rank matrix; training the local first low-rank matrix and the local second low-rank matrix through the received updated global first low-rank matrix, the updated global second low-rank matrix and the local data until the training of the local first low-rank matrix and the training of the local second low-rank matrix reach a set iteration condition; and generating a local medical model according to the finally obtained global first low-rank matrix, the finally obtained global second low-rank matrix, the finally obtained local first low-rank matrix, the finally obtained local second low-rank matrix and the local model. According to the method, the training efficiency is improved and the communication consumption caused by model transmission is reduced by reducing the number of the training weights and the size of the transmitted model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical big data, and specifically, to a medical model training method, device and computer-readable storage medium. Background Art

[0002] As a knowledge-intensive field, the medical field mainly contains a large amount of medical texts, including but not limited to patient medical records, medical literature, clinical reports, etc. These texts contain a wealth of medical knowledge. The large language model provides unprecedented opportunities for the automated analysis and understanding of medical texts through its powerful semantic understanding and context perception capabilities. It can efficiently extract potential knowledge, accelerate the progress of medical research, and provide more comprehensive support for medical decision-making.

[0003] However, due to the complexity of medical models, the current medical model training efficiency is low. Summary of the invention

[0004] In view of this, the purpose of the embodiments of the present application is to provide a medical model training method, device and computer-readable storage medium, which can improve the efficiency of medical model training.

[0005] In a first aspect, an embodiment of the present application provides a medical model training method, which is applied to a client, and the method comprises: performing low-rank decomposition on the gradient of the local model to obtain a decomposed local first low-rank matrix and a local second low-rank matrix; training the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client; sending the trained local first low-rank matrix and the local second low-rank matrix to a server; wherein the global first low-rank matrix and the global second low-rank matrix are configured to be transmitted through the client The local first low-rank matrix and the local second low-rank matrix after training are updated; the local first low-rank matrix and the local second low-rank matrix are trained by the received updated global first low-rank matrix, the updated global second low-rank matrix and the local data until the training of the local first low-rank matrix and the local second low-rank matrix reaches the set iteration condition; according to the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above training process, a local medical model is generated.

[0006] In the above implementation process, by training the local first low-rank matrix and the local second low-rank matrix after the low-rank decomposition of the local model, since only the update and calculation of the parameters of the local first low-rank matrix and the local second low-rank matrix are set during the training process, the number of weights for local data training can be reduced and the training efficiency can be improved. Moreover, after the local model is updated, only the updated local first low-rank matrix and the local second low-rank matrix need to be transmitted, which can effectively reduce the size of the model that needs to be transmitted in federated learning, reduce the communication consumption caused by model transmission, and improve the overall efficiency of medical model training.

[0007] In one embodiment, the local first low rank matrix and the local second low rank matrix are trained based on the local model, the received global first low rank matrix, the received global second low rank matrix and the local data of the client, including: performing forward propagation calculation based on the local model, the local data, the local first low rank matrix and the local second low rank matrix to obtain a forward propagation output result; calculating the gradient of the local first low rank matrix and the gradient of the local second low rank matrix based on the forward propagation output result, the global first low rank matrix and the global second low rank matrix; updating the local first low rank matrix and the local second low rank matrix according to the gradient of the local first low rank matrix and the gradient of the local second low rank matrix.

[0008] In the above implementation process, during the training of the local first low-rank matrix and the local second low-rank matrix, the parameters involved in the update only include the local first low-rank matrix and the local second low-rank matrix. Since the local first low-rank matrix and the local second low-rank matrix only occupy a very small part of the local model, the amount of parameter calculation for local training is greatly reduced, the computing resources occupied by model training are reduced, and the training efficiency is improved.

[0009] In one embodiment, there are multiple clients; sending the trained local first low-rank matrix and local second low-rank matrix to the server includes: each client generates a key random number with one or more other clients through a DH key exchange algorithm; encrypting the trained local first low-rank matrix and local second low-rank matrix of the client through the key random number; and sending the encrypted local first low-rank matrix and local second low-rank matrix to the server.

[0010] In the above implementation process, the local first low-rank matrix and the local second low-rank matrix are encrypted by adopting the secure aggregation method of the DH key exchange algorithm. Since the encryption complexity of the secure aggregation method mainly depends on the number of clients rather than the number of model parameters, and this method does not affect the final effect of the model. Therefore, the security of the local first low-rank matrix and the local second low-rank matrix can be improved while reducing the impact on the final effect of the medical model training.

[0011] In one embodiment, before the local first low rank matrix and the local second low rank matrix are trained based on the local model, the received global first low rank matrix, the received global second low rank matrix and the local data of the client, the method also includes: desensitizing the local data according to multiple desensitization requirements to obtain multiple desensitized local data; the training of the local first low rank matrix and the local second low rank matrix based on the local model, the received global first low rank matrix, the received global second low rank matrix and the local data of the client includes: training the local first low rank matrix and the local second low rank matrix based on the local model, the received global first low rank matrix, the received global second low rank matrix and the desensitized local data.

[0012] In the above implementation process, the local first low-rank matrix and the local second low-rank matrix are trained by using the desensitized local data after desensitization processing. Since the local data involved in the training of the local first low-rank matrix and the local second low-rank matrix are desensitized data, the privacy and security of the local data can be effectively protected, thereby realizing data sharing and cooperation among multiple clients, and solving the problem of medical data islands.

[0013] In one embodiment, the multiple desensitized local data of the client include first desensitized local data; after generating the local medical model according to the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above-mentioned training process, the method also includes: training the local medical model through the first desensitized local data to obtain the corresponding personalized medical model of the client; wherein, the first desensitized local data is the local data processed in a desensitization method with lower desensitization requirements among the multiple desensitized local data.

[0014] In the above implementation process, the local medical model is trained by the first desensitized local data. Since the desensitization requirements of the first desensitized local data are relatively low, the information corresponding to the first desensitized local data is relatively more comprehensive. Therefore, by training the local medical model with the first desensitized local data, a personalized medical model that is more in line with the actual situation of the client can be obtained, thereby improving the accuracy and flexibility of the personalized medical model.

[0015] In a second aspect, an embodiment of the present application further provides a medical model training method, which is applied to a server side, and the method comprises: sending a global first low-rank matrix and a global second low-rank matrix to a client; wherein the global first low-rank matrix and the global second low-rank matrix are configured to train a local first low-rank matrix and a local second low-rank matrix with a local model of the client and local data of the client; the local first low-rank matrix and the local second low-rank matrix are obtained by performing low-rank decomposition on the gradient of the local model; updating the global first low-rank matrix and the global second low-rank matrix according to the received trained local first low-rank matrix and local second low-rank matrix; and sending The client sends the updated global first low-rank matrix and the global second low-rank matrix; updates the global first low-rank matrix and the global second low-rank matrix according to the trained local first low-rank matrix and the local second low-rank matrix until the updates of the global first low-rank matrix and the global second low-rank matrix reach the set iteration condition; sends the final global first low-rank matrix and the final global second low-rank matrix to the client; wherein the final global first low-rank matrix and the final global second low-rank matrix are configured to generate a local medical model with the final local first low-rank matrix, the final local second low-rank matrix and the local model.

[0016] In the above implementation process, when the server updates the global first low-rank matrix and the global second low-rank matrix, it only needs to use the trained local first low-rank matrix and the local second low-rank matrix for update, which can effectively reduce the model size that needs to be transmitted in federated learning, reduce the communication consumption caused by model transmission, and improve the overall efficiency of medical model training.

[0017] In one embodiment, before sending the global first low-rank matrix and the global second low-rank matrix to the client, the method also includes: setting a general model through public data set training to obtain the global medical model; wherein the data in the public data set is data in the public data source of the corresponding medical field; sending the global medical model to the client; wherein the global medical model is configured to perform an initial training on the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client.

[0018] In the above implementation process, by using the data in the public data set to train the set general model, the set general model can be trained into a global medical model commonly used in the medical field, thereby improving the training effect of the global medical model.

[0019] In a third aspect, an embodiment of the present application further provides a medical model training device, which is applied to a client, and the device includes: a decomposition module, which is used to perform low-rank decomposition on the gradient of the local model to obtain a decomposed local first low-rank matrix and a local second low-rank matrix; a first training module, which is used to train the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client; a first sending module, which is used to send the trained local first low-rank matrix and the local second low-rank matrix to the server; wherein the global first low-rank matrix and the global second low-rank matrix are configured as The local first low-rank matrix and the local second low-rank matrix trained by the client are updated; the first iteration module is used to train the local first low-rank matrix and the local second low-rank matrix through the received updated global first low-rank matrix, the updated global second low-rank matrix and the local data until the training of the local first low-rank matrix and the local second low-rank matrix reaches the set iteration condition; the generation module is used to generate a local medical model according to the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above training process.

[0020] In a fourth aspect, an embodiment of the present application further provides a medical model training device, which is applied to a server side, and the device comprises: a second sending module, which is used to send a global first low-rank matrix and a global second low-rank matrix to a client; wherein the global first low-rank matrix and the global second low-rank matrix are configured to train the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client; the local first low-rank matrix and the local second low-rank matrix are obtained by performing low-rank decomposition on the gradient of the local model; an updating module, which is used to update the global first low-rank matrix and the global second low-rank matrix according to the received trained local first low-rank matrix and the local second low-rank matrix; the second sending module, It is also used to send the updated global first low-rank matrix and the global second low-rank matrix to the client; the second iteration module is used to update the global first low-rank matrix and the global second low-rank matrix according to the trained local first low-rank matrix and the local second low-rank matrix until the updates of the global first low-rank matrix and the global second low-rank matrix reach the set iteration conditions; the second sending module is also used to send the final global first low-rank matrix and the final global second low-rank matrix to the client; wherein the final global first low-rank matrix and the final global second low-rank matrix are configured to generate a local medical model with the final local first low-rank matrix, the final local second low-rank matrix and the local model.

[0021] In a fifth aspect, an embodiment of the present application further provides an electronic device, comprising: a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method in the above-mentioned first aspect, or any possible implementation of the first aspect.

[0022] In a sixth aspect, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the medical model training method in the above-mentioned first aspect, or any possible implementation of the first aspect are executed.

[0023] In order to make the above-mentioned objects, features and advantages of the present application more obvious and understandable, embodiments are given below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0025] Figure 1 A flowchart of a medical model training method applied to a client provided in an embodiment of the present application;

[0026] Figure 2 A flowchart of a medical model training method applied to a server provided in an embodiment of the present application;

[0027] Figure 3 A schematic diagram of functional modules of a medical model training device applied to a client provided in an embodiment of the present application;

[0028] Figure 4 A schematic diagram of the functional modules of a medical model training device applied to a server provided in an embodiment of the present application;

[0029] Figure 5 A schematic diagram of the interaction between the server and the client provided in the embodiment of the present application;

[0030] Figure 6 A block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application.

[0032] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0033] In today's digital age, the explosive development and widespread application of large language models have become a highlight in the field of science and technology. These large models have achieved remarkable results in tasks such as natural language processing and image recognition, bringing huge opportunities to all walks of life. If large language models are applied to the medical field and introduced into the medical field, it can provide a new perspective for medical text analysis and knowledge mining.

[0034] However, just like other technologies, the widespread application of large language models also faces some challenges. In the medical field, it is necessary to consider the high sensitivity of patient privacy and the professionalism and complexity of medical texts. In order to overcome these challenges, the medical field not only needs the technical support of large language models, but also needs to combine these models with the special needs of medical scenarios to form large language models in the medical field.

[0035] After long-term research, the inventors of this application found that due to the particularity of data in the medical field and the high restrictions on privacy requirements, the training of medical models faces unique difficulties. The difference between the medical field and other industries is that its data is sensitive and difficult to obtain. Patients' personal health information is extremely private and needs to be subject to strict regulations and ethical protection. At the same time, medical data is scattered in different medical institutions, forming data islands, and these institutions are often unwilling or restricted to share patient data. In addition, due to the high complexity of the medical model, the training efficiency of the medical model is low when it is trained.

[0036] In view of this, the present application proposes a medical model training method based on federated learning technology, which enables various medical institutions to jointly build and fine-tune large language models without directly sharing patient sensitive data, which not only meets the demand for large language models in the medical field, but also ensures the security of patient privacy. In addition, by adopting a low-rank adaptation method to train the local model, since the low-rank adaptation method only involves the update of part of the weights, the number of weights for local data training can be reduced, and the training efficiency can be improved. Moreover, after the local model is updated, only part of the updated weights need to be transferred, which can effectively reduce the size of the model that needs to be transferred in federated learning, reduce the communication consumption caused by model transfer, and improve the overall efficiency of medical model training.

[0037] See also Figure 1, is a flow chart of a medical model training method applied to a client provided in an embodiment of the present application. Figure 1 The specific process shown is described in detail.

[0038] Step S201, performing low-rank decomposition on the gradient of the local model to obtain a decomposed local first low-rank matrix and a local second low-rank matrix.

[0039] Among them, the gradient of the local model is low-rank decomposed by adopting the LoRA low-rank adaptation method. This low-rank adaptation method is the LoRA low-rank adaptation method in the fine-tuning method of transfer learning.

[0040] It should be understood that large language models usually have a large number of model parameters, and the transmission volume of the model can reach tens or hundreds of GB. The transmission efficiency of such a large number of parameters for model interaction in federated learning is low and it is difficult to implement. However, neural networks contain many dense layers that perform matrix multiplication, and the weight matrices in these layers usually have full rank. When adapted to specific tasks, language models have a lower "intrinsic dimension" and can still be effectively learned despite random projections to smaller subspaces, which provides a theoretical basis for LoRA's fine-tuning training method. By applying the fine-tuning method of transfer learning, the LoRA low-rank adaptation method can complete the fine-tuning training of local models in federated learning.

[0041] In one embodiment, the gradient of the local model is subjected to low-rank decomposition to obtain the decomposed first low-rank matrix and the second low-rank matrix by the following formula:

[0042] W 0 +ΔW=W 0 +BA;

[0043] Among them, W 0 is the local model, ΔW is the gradient of the local model during training, A is the first low-rank matrix, and B is the second low-rank matrix.

[0044] Here r 1 <<min(d,k),r 2 <<min(d,k). Among them, is the field of real numbers, r 1 is the number of rows of the second lowest rank matrix, r 2 is the number of columns of the first low-rank matrix, d is the number of rows of the local model, and k is the number of columns of the local model. 1 =r 2 .

[0045] The above-mentioned local first low-rank matrix and local second low-rank matrix have smaller ranks, and the weights contained in the local first low-rank matrix B and the local second low-rank matrix A are much smaller than ΔW.

[0046] In one embodiment, BA multiplied by the input local data and ΔW multiplied by the input local data have outputs of the same shape.

[0047] Step S202: training the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client.

[0048] The local model is a model pre-trained by the local data of the client before step S201. During the training of the local first low-rank matrix and the local second low-rank matrix, the local model is frozen and does not receive gradient updates.

[0049] The above-mentioned local data is the data in the client's data set.

[0050] Optionally, the client may be one or more. For example, the client may be a hospital, a medical institution, a medical research institution, etc., and the local data may be patient data in a hospital database, patient data in a medical institution database, research subject data in a medical research institution database, etc. The local data may be selected according to actual conditions.

[0051] Here, the global first low-rank matrix and the global second low-rank matrix are determined by the server according to the local first low-rank matrix and the local second low-rank matrix of each client. The global first low-rank matrix and the global second low-rank matrix are configured to be updated by the local first low-rank matrix and the local second low-rank matrix after client training.

[0052] In one embodiment, when the local first low rank matrix and the local second low rank matrix are trained, the global first low rank matrix and the global second low rank matrix are in a frozen state (ie, the global first low rank matrix and the global second low rank matrix only participate in the calculation and are not updated).

[0053] In one embodiment, the training of the local first low-rank matrix and the local second low-rank matrix can be expressed by the following formula:

[0054]

[0055] Among them, t is the number of iterations, L(W (t) ,W public ) is the loss function of the local first low-rank matrix and the local second low-rank matrix for the tth iteration, λ is the learning rate, are the local first low-rank matrix and the local second low-rank matrix of the t-th iteration, is the local first low rank matrix and the local second low rank matrix of the t+1th iteration, D iis the i-th type of data set.

[0056] Step S203: Send the trained local first low-rank matrix and the local second low-rank matrix to the server.

[0057] It should be understood that after the server obtains the local first low-rank matrix and the local second low-rank matrix sent by each client, it aggregates the local first low-rank matrix and the local second low-rank matrix of each client to obtain the global first low-rank matrix and the global second low-rank matrix. Step S204, train the local first low-rank matrix and the local second low-rank matrix through the received updated global first low-rank matrix, the updated global second low-rank matrix and the local data until the training of the local first low-rank matrix and the local second low-rank matrix reaches the set iteration condition.

[0058] The set iteration condition here may be a set number of iterations, a set iteration time, or reaching a set accuracy, etc. The set iteration condition may be selected according to actual conditions.

[0059] Step S205, generates a local medical model based on the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above training process.

[0060] The local first low-rank matrix and the local second low-rank matrix finally obtained are the local first low-rank matrix and the local second low-rank matrix when the set iteration condition is reached. The local first low-rank matrix and the local second low-rank matrix finally obtained are also used to train the global first low-rank matrix and the global second low-rank matrix to determine the global first low-rank matrix and the global second low-rank matrix finally obtained.

[0061] It should be understood that after the client receives the final global first low-rank matrix and the global second low-rank matrix, the local first low-rank matrix and the local second low-rank matrix are trained again based on the final global first low-rank matrix and the global second low-rank matrix to obtain the final local first low-rank matrix and the local second low-rank matrix. And a local medical model is generated based on the final local first low-rank matrix and the final local second low-rank matrix.

[0062] In the above implementation process, by training the local first low-rank matrix and the local second low-rank matrix after the low-rank decomposition of the local model, since only the update and calculation of the parameters of the local first low-rank matrix and the local second low-rank matrix are set during the training process, the number of weights for local data training can be reduced and the training efficiency can be improved. Moreover, after the local model is updated, only the updated local first low-rank matrix and the local second low-rank matrix need to be transmitted, which can effectively reduce the size of the model that needs to be transmitted in federated learning, reduce the communication consumption caused by model transmission, and improve the overall efficiency of medical model training.

[0063] In one possible implementation, step S202 includes: performing forward propagation calculations based on the local model, local data, a local first low-rank matrix, and a local second low-rank matrix to obtain a forward propagation output result; calculating the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix based on the forward propagation output result, the global first low-rank matrix, and the global second low-rank matrix; and updating the local first low-rank matrix and the local second low-rank matrix according to the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix.

[0064] Among them, the local model, the global first low-rank matrix and the global second low-rank matrix only participate in the parameter calculation of training, and the local first low-rank matrix and the local second low-rank matrix are updated through training.

[0065] In one embodiment, the calculation process of the forward propagation during the training process may be:

[0066] h=W 0 x+ΔWx=W 0 x+BAx;

[0067] Among them, W 0 is the local model, ΔW is the gradient of the local model during training, A is the first low-rank matrix, B is the second low-rank matrix, and x is the input local data.

[0068] Here, the first low-rank matrix can be initialized using random Gaussian initialization, and the second low-rank matrix can be initialized using zero initialization, thereby ensuring that ΔW=BA is 0 at the beginning of training.

[0069] Understandably, in the Transformer architecture, there are four weight matrices in the self-attention module, which can be W q , W k , W v , W o Among them, W q is the weight matrix used to generate the Query vector, W k is the weight matrix used to generate the Key vector, W vis the weight matrix used to generate the Value vector, W o For local models.

[0070] Among them, the Query vector is obtained by inputting the embedding vector and the weight matrix W q Multiply them together to find the key vector associated with it. The key vector is obtained by inputting the embedding vector and the weight matrix W k The Value vector is obtained by multiplying the input embedding vector and the weight matrix W. v Multiply together to obtain the specific information of the word.

[0071] It should be understood that when training the local first low-rank matrix and the local second low-rank matrix, only W q and W v The corresponding local first low-rank matrix and local second low-rank matrix are added after the weight matrix. In the backward propagation calculation, the output gradient of the forward propagation output result is calculated according to the loss function, and the output gradient is back-propagated to each layer in the network to calculate the partial derivatives of the loss function to the local first low-rank matrix and the local second low-rank matrix (that is, the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix), and then the local first low-rank matrix and the local second low-rank matrix are updated based on the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix.

[0072] In one embodiment, during the entire training process of the local first low-rank matrix and the local second low-rank matrix, the original local model is frozen to reduce the number of weights for local data training.

[0073] In the above implementation process, during the training of the local first low-rank matrix and the local second low-rank matrix, the parameters involved in the update only include the local first low-rank matrix and the local second low-rank matrix. Since the local first low-rank matrix and the local second low-rank matrix only occupy a very small part of the local model, the amount of parameter calculation for local training is greatly reduced, the computing resources occupied by model training are reduced, and the training efficiency is improved.

[0074] In a possible implementation, step S203 includes: each client generates a key random number with one or more other clients through a DH key exchange algorithm; encrypts the local first low-rank matrix and the local second low-rank matrix trained by the client through the key random number; and sends the encrypted local first low-rank matrix and the local second low-rank matrix to the server.

[0075] Understandably, in each round of federated learning updates, the server can infer some attributes of the client data through iterative updates of the local model. Therefore, in the process of federated learning, the local model also needs to be protected to prevent the server from viewing the local model. Since the training of the medical vertical federated large model has a large number of parameters, and the clients are mostly large institutions such as hospitals, and the number is not too large, a secure aggregation method can be used to protect the privacy of the models of all parties.

[0076] The encryption complexity of the secure aggregation method here is only related to the number of clients and has little to do with the number of model parameters, and will not have a negative impact on the results.

[0077] In one embodiment, it is assumed that client C u Have a local low-rank matrix x u For any other client C in the system v (u≠v), client C u and client C v Generate a secret random number s through the Diffie-Hellman key exchange algorithm (DH) uv After that, client C u The local low-rank matrix x is calculated according to the following formula u Encrypt as follows:

[0078] x u =x u +∑ v∈n,u<v s uv -∑ v∈n,u>v s uv ;

[0079] Among them, x u is the local low-rank matrix, v is the client number of any other client, and u is the local low-rank matrix x u The client number, n is the total number of clients in the system, s uv For client C u and client C v Generate a secret random number.

[0080] It should be understood that the above encryption process is the process of adding random number disturbance to the client's local low-rank matrix. When the server receives the local low-rank matrix of each client, it will not see the plaintext of the local low-rank matrix. After submission to the server, the aggregation method on the server does not require special processing. The encrypted local low-rank matrix will also eliminate the secret key after aggregation, which will not affect the aggregation result.

[0081] In the above implementation process, the local first low-rank matrix and the local second low-rank matrix are encrypted by adopting the secure aggregation method of the DH key exchange algorithm. Since the encryption complexity of the secure aggregation method mainly depends on the number of clients rather than the number of model parameters, and this method does not affect the final effect of the model. Therefore, the security of the local first low-rank matrix and the local second low-rank matrix can be improved while reducing the impact on the final effect of the medical model training.

[0082] In a possible implementation, before step S202, the method further includes: desensitizing the local data according to a plurality of desensitizing requirements to obtain a plurality of desensitized local data.

[0083] Among them, the multiple desensitized local data of the client include first desensitized local data and second desensitized local data. The first desensitized local data is local data processed according to a desensitization method with lower desensitization requirements among the multiple desensitized local data. The second desensitized local data is local data processed according to a desensitization method with higher desensitization requirements among the multiple desensitized local data.

[0084] It should be understood that since the private data of the client involves the patient's personal health information, such as medical records, diagnosis, treatment plans, etc. This information is extremely sensitive personal privacy, and directly using the original data for model training may lead to the risk of patient privacy leakage. Through desensitization processing, the threat to patient privacy can be minimized.

[0085] The desensitized data scope here can be desensitized according to the sensitive data classification standards stipulated in the relevant standards of medical data information security technology, and the data of the set level can be desensitized. Among them, the desensitization means may include removal, generalization, etc.

[0086] Exemplarily, personal attribute data may be removed by removing information that can uniquely identify an individual or that would have a significant impact on the individual if disclosed (such as name; ID card / driver's license number, telephone number, fax, email, medical insurance number, medical record number, etc.).

[0087] Personal attribute data that can be indirectly related to personal information (such as date of birth, consultation time, examination time, treatment time, hospitalization and discharge time, work unit, etc.) can be processed in a generalized manner. Among them, for data that still has medical significance after fuzzification, the fuzzified results can be retained.

[0088] In one embodiment, step S202 includes: training the local first low rank matrix and the local second low rank matrix based on the local model, the received global first low rank matrix, the received global second low rank matrix and the desensitized local data.

[0089] The first desensitized local data and the second desensitized local data are used for local model training in different scenarios. For example, the first desensitized local data can be used to train a local medical model, and the second desensitized local data is used to train a local first low-rank matrix and a local second low-rank matrix.

[0090] In one embodiment, before step S202, the method further includes: preprocessing a plurality of desensitized local data.

[0091] The preprocessing here may include data cleaning, word segmentation, removal of stop words and other data preprocessing methods.

[0092] Data cleaning is used to clean the original local data to remove HTML tags, special characters, extra spaces, etc. to ensure that the text is clean. Word segmentation is used to divide the continuous original local data into independent words or phrases. Stop word removal is used to remove common words that do not contribute much in the local data.

[0093] In the above implementation process, the local first low-rank matrix and the local second low-rank matrix are trained by using the desensitized local data after desensitization processing. Since the local data involved in the training of the local first low-rank matrix and the local second low-rank matrix are desensitized data, the privacy and security of the local data can be effectively protected, thereby realizing data sharing and cooperation among multiple clients, and solving the problem of medical data islands.

[0094] In a possible implementation, after step S205, the method further includes: training a local medical model using the first desensitized local data to obtain a personalized medical model corresponding to the client.

[0095] It can be understood that after the client generates the local medical model based on the final global first low-rank matrix, the final global second low-rank matrix, the final local first low-rank matrix, the final local second low-rank matrix and the local model, it can also continue to train the generated local medical model based on the client's local data to obtain the client's own personalized medical model.

[0096] In one embodiment, the personalized medical model can be trained by the following process:

[0097]

[0098] in, is the personalized medical model of the i-th client at the t+1th iteration, is the personalized medical model of the i-th client at the t-th iteration, is the loss function of the tth iteration of the personalized medicine model.

[0099] In the above implementation process, the local medical model is trained by the first desensitized local data. Since the desensitization requirements of the first desensitized local data are relatively low, the information corresponding to the first desensitized local data is relatively more comprehensive. Therefore, by training the local medical model with the first desensitized local data, a personalized medical model that is more in line with the actual situation of the client can be obtained, thereby improving the accuracy and flexibility of the personalized medical model.

[0100] See also Figure 2 , is a flow chart of a medical model training method applied to a server provided in an embodiment of the present application. Figure 2 The specific process shown is described in detail.

[0101] Step S301, sending a global first low-rank matrix and a global second low-rank matrix to a client.

[0102] The global first low-rank matrix and the global second low-rank matrix are configured to train the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client. The local first low-rank matrix and the local second low-rank matrix are obtained by performing low-rank decomposition on the gradient of the local model.

[0103] The global medical model here is trained by the server based on the local model of the client, or trained based on data in the public data source in the corresponding medical field.

[0104] Step S302: Update the global first low rank matrix and the global second low rank matrix according to the received trained local first low rank matrix and local second low rank matrix.

[0105] The update of the global first low rank matrix and the global second low rank matrix here can be expressed by the following formula:

[0106]

[0107] Among them, t is the number of iterations, L(W (t) ,D public ) is the loss function of the tth iteration, λ is the learning rate, W (t) is the global first low rank matrix and the global second low rank matrix of the t-th iteration, W (t+1) are the global first low-rank matrix and the global second low-rank matrix of the t+1th iteration.

[0108] In one embodiment, the server further aggregates the received trained local first low-rank matrix and local second low-rank matrix. The aggregation formula may be as follows:

[0109]

[0110] in, are the local first low-rank matrix and the local second low-rank matrix of the t+1th iteration, n is the number of clients, i is the i-th client, are the local first low-rank matrix and the local second low-rank matrix of the i-th client at the t+1-th iteration.

[0111] Step S303: Send the updated global first low-rank matrix and the global second low-rank matrix to the client.

[0112] Step S304, updating the global first low rank matrix and the global second low rank matrix according to the trained local first low rank matrix and the local second low rank matrix until the update of the global first low rank matrix and the global second low rank matrix reaches the set iteration condition.

[0113] The set iteration condition here may be the set number of iterations, the set iteration time, or the reaching of the set accuracy, etc. The set iteration condition may be selected according to the actual situation. Step S305, sending the finally obtained global first low-rank matrix and the finally obtained global second low-rank matrix to the client.

[0114] Among them, the final global first low-rank matrix and the final global second low-rank matrix are configured to generate a local medical model with the final local first low-rank matrix, the final local second low-rank matrix and the local model.

[0115] It should be understood that after the client receives the final global first low-rank matrix and the global second low-rank matrix, the local first low-rank matrix and the local second low-rank matrix are trained again based on the final global first low-rank matrix and the global second low-rank matrix to obtain the final local first low-rank matrix and the local second low-rank matrix. And a local medical model is generated based on the final local first low-rank matrix and the final local second low-rank matrix.

[0116] In the above implementation process, when the server updates the global first low-rank matrix and the global second low-rank matrix, it only needs to use the trained local first low-rank matrix and the local second low-rank matrix for update, which can effectively reduce the size of the model that needs to be transmitted in federated learning, reduce the communication consumption caused by model transmission, and improve the overall efficiency of medical model training.

[0117] In a possible implementation, before step S301, the method further includes: setting a general model through public data set training to obtain a global medical model; and sending the global medical model to the client.

[0118] The data in the public datasets are the data in the public data sources in the corresponding medical field, such as medical records, disease descriptions, treatment plans, and other data in medical-related papers, medical-related public patents, medical news, academic journals, and other texts.

[0119] The set general model here can be GPT2, ChatGLM-6B, LLaMA-7B, DeBERTa, etc. The set general model can be selected according to actual conditions.

[0120] The above-mentioned global medical model is configured to perform an initial training on the local first low-rank matrix and the local second low-rank matrix together with the local model of the client and the local data of the client.

[0121] It should be understood that when the server obtains the global medical model through training, the server does not yet have the local first low-rank matrix and the local second low-rank matrix sent by the client. At this time, the server sends the trained global medical model to the client so that the client can perform the first training of the local first low-rank matrix and the local second low-rank matrix according to the acquired global medical model. In the above implementation process, by using the data in the public data set to train the set general model, the set general model can be trained as a global medical model commonly used in the medical field, thereby improving the training effect of the global medical model.

[0122] Based on the same application concept, the embodiment of the present application also provides a medical model training device applied to the client corresponding to the medical model training method applied to the client. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the aforementioned embodiment of the medical model training method applied to the client, the implementation of the device in this embodiment can refer to the description in the embodiment of the above-mentioned method, and the repeated parts will not be repeated.

[0123] See also Figure 3 , is a functional module diagram of a medical model training device applied to a client provided in an embodiment of the present application. Each module in the medical model training device applied to a client in this embodiment is used to execute each step in the above method embodiment. The medical model training device applied to a client includes a decomposition module 301, a first training module 302, a first sending module 303, a first iteration module 304, and a generation module 305; wherein,

[0124] The decomposition module 301 is used to perform low-rank decomposition on the gradient of the local model to obtain a decomposed local first low-rank matrix and a local second low-rank matrix.

[0125] The first training module 302 is used to train the local first low rank matrix and the local second low rank matrix based on the local model, the received global first low rank matrix, the received global second low rank matrix and the local data of the client.

[0126] The first sending module 303 is used to send the trained local first low-rank matrix and the local second low-rank matrix to the server; wherein the global first low-rank matrix and the global second low-rank matrix are configured to be updated by the local first low-rank matrix and the local second low-rank matrix trained by the client.

[0127] The first iteration module 304 is used to train the local first low rank matrix and the local second low rank matrix through the received updated global first low rank matrix, the updated global second low rank matrix and the local data until the training of the local first low rank matrix and the local second low rank matrix reaches the set iteration condition.

[0128] The generation module 305 is used to generate a local medical model based on the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above training process.

[0129] In one possible implementation, the first training module 302 is specifically used to: perform forward propagation calculations based on the local model, the local data, the local first low-rank matrix, and the local second low-rank matrix to obtain a forward propagation output result; calculate the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix based on the forward propagation output result, the global first low-rank matrix, and the global second low-rank matrix; and update the local first low-rank matrix and the local second low-rank matrix according to the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix.

[0130] In one possible implementation, the first sending module 303 is specifically used to: each client generates a key random number with one or more other clients through a DH key exchange algorithm; encrypts the local first low-rank matrix and the local second low-rank matrix trained by the client through the key random number; and sends the encrypted local first low-rank matrix and the local second low-rank matrix to the server.

[0131] In a possible implementation, the medical model training device applied to the client further includes a desensitization module for desensitizing the local data according to a variety of desensitization requirements to obtain a variety of desensitized local data.

[0132] In one possible implementation, the first training module 302 is specifically used to train the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the desensitized local data.

[0133] In a possible implementation, the first training module 302 is also used to: train the local medical model through the first desensitized local data to obtain a corresponding personalized medical model for the client; wherein the first desensitized local data is local data processed in a desensitization method with lower desensitization requirements among multiple desensitized local data.

[0134] Based on the same application concept, the embodiment of the present application also provides a medical model training device applied to the server side corresponding to the medical model training method applied to the server side. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the aforementioned embodiment of the medical model training method applied to the server side, the implementation of the device in this embodiment can refer to the description in the embodiment of the above-mentioned method, and the repeated parts will not be repeated.

[0135] See also Figure 4 , is a functional module diagram of a medical model training device applied to a server provided in an embodiment of the present application. Each module in the medical model training device applied to a server in this embodiment is used to execute each step in the above method embodiment. The medical model training device applied to a server includes a second sending module 401, an updating module 402, and a second iteration module 403; wherein,

[0136] The second sending module 401 is used to send a global first low-rank matrix and a global second low-rank matrix to the client; wherein the global first low-rank matrix and the global second low-rank matrix are configured to train the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client; the local first low-rank matrix and the local second low-rank matrix are obtained by performing low-rank decomposition on the gradient of the local model.

[0137] The updating module 402 is used to update the global first low rank matrix and the global second low rank matrix according to the received trained local first low rank matrix and local second low rank matrix.

[0138] The second sending module 401 is further used to send the updated global first low-rank matrix and the global second low-rank matrix to the client.

[0139] The second iteration module 403 is used to update the global first low rank matrix and the global second low rank matrix according to the trained local first low rank matrix and the local second low rank matrix until the updates of the global first low rank matrix and the global second low rank matrix meet the set iteration conditions.

[0140] The second sending module 401 is also used to send the final global first low-rank matrix and the final global second low-rank matrix to the client; wherein the final global first low-rank matrix and the final global second low-rank matrix are configured to generate a local medical model with the final local first low-rank matrix, the final local second low-rank matrix and the local model.

[0141] In a possible implementation, the medical model training device applied to the server side also includes a second training module, which is used to set a general model through public data set training to obtain the global medical model; wherein the data in the public data set is data in the public data source of the corresponding medical field.

[0142] In one possible implementation, the global medical model is sent to the client; wherein the global medical model is configured to perform an initial training on the local first low-rank matrix and the local second low-rank matrix with the client's local model and the client's local data.

[0143] In a possible implementation, the second sending module 401 is also used to send the global medical model to the client; wherein the global medical model is configured to perform an initial training on the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client.

[0144] To facilitate understanding of this embodiment, the operating environment for executing a medical model training method disclosed in the embodiment of the present application is introduced in detail below.

[0145] like Figure 5 , which is a schematic diagram of the interaction between the server and the client provided in the embodiment of the present application. The server is connected to one or more clients through a network for data communication or interaction. The server may be a network server, a database server, etc. The local terminal may be a personal computer (PC), a tablet computer, a smart phone, a personal digital assistant (PDA), etc.

[0146] To facilitate understanding of this embodiment, the electronic device for executing the medical model training method disclosed in the embodiment of the present application is described in detail below. The electronic device in the embodiment of the present application can be a server-side device or a client-side device. The specific setting method of the electronic device can be selected according to actual conditions.

[0147] like Figure 6 , which is a block diagram of an electronic device. The electronic device 100 may include a memory 111 and a processor 113. A person skilled in the art may understand that Figure 6 The structure shown is only for illustration and does not limit the structure of the electronic device 100. For example, the electronic device 100 may further include Figure 6 More or fewer components as shown, or with Figure 6 Different configurations are shown.

[0148] The memory 111 and the processor 113 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these elements can be electrically connected to each other via one or more communication buses or signal lines. The processor 113 is used to execute the executable module stored in the memory.

[0149] The memory 111 may be, but not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory 111 is used to store programs, and the processor 113 executes the program after receiving the execution instruction. The method executed by the electronic device 100 defined by the process disclosed in any embodiment of the present application can be applied to the processor 113, or implemented by the processor 113.

[0150] The processor 113 may be an integrated circuit chip with signal processing capability. The processor 113 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0151] The electronic device 100 in this embodiment can be used to execute each step in each method provided in the embodiments of the present application.

[0152] In addition, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the medical model training method described in the above method embodiment are executed.

[0153] The computer program product of the medical model training method provided in the embodiment of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the medical model training method described in the above method embodiment. Please refer to the above method embodiment for details, which will not be repeated here.

[0154] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0155] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0156] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk. It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or apparatus. In the absence of more restrictions, the elements defined by the sentence "includes..." do not exclude the presence of other identical elements in the process, method, article or apparatus that includes the elements.

[0157] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application. It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0158] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A medical model training method, characterized in that: Applied to a client, the method comprises: Perform low-rank decomposition on the gradient of the local model to obtain a decomposed local first low-rank matrix and a local second low-rank matrix; Based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client, the local first low-rank matrix and the local second low-rank matrix are trained; wherein the global first low-rank matrix and the global second low-rank matrix are configured to be updated by the local first low-rank matrix and the local second low-rank matrix trained by the client; Sending the trained local first low-rank matrix and the local second low-rank matrix to the server; Training the local first low rank matrix and the local second low rank matrix by using the received updated global first low rank matrix, the updated global second low rank matrix and the local data until the training of the local first low rank matrix and the local second low rank matrix reaches a set iteration condition; A local medical model is generated based on the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above training process.

2. The method according to claim 1, characterized in that The training of the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client includes: Performing forward propagation calculation based on the local model, the local data, the local first low-rank matrix, and the local second low-rank matrix to obtain a forward propagation output result; Calculate the gradient of the local first low-rank matrix and the gradient of the local second low-rank matrix based on the forward propagation output result, the global first low-rank matrix, and the global second low-rank matrix; The local first low rank matrix and the local second low rank matrix are updated according to the gradient of the local first low rank matrix and the gradient of the local second low rank matrix.

3. The method according to claim 1, characterized in that in, There are multiple clients; sending the trained local first low-rank matrix and the local second low-rank matrix to the server includes: Each client generates a key random number with one or more other clients through the DH key exchange algorithm; Encrypting the local first low-rank matrix and the local second low-rank matrix trained by the client by using the key random number; The encrypted local first low-rank matrix and the local second low-rank matrix are sent to the server.

4. The method according to claim 1, characterized in that: Before training the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client, the method further includes: Desensitizing the local data according to multiple desensitization requirements to obtain multiple desensitized local data; The training of the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the local data of the client includes: Based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and the desensitized local data, the local first low-rank matrix and the local second low-rank matrix are trained.

5. The method according to claim 4, characterized in that in, The multiple desensitized local data of the client include first desensitized local data; after generating a local medical model according to the global first low-rank matrix finally obtained in the training process, the global second low-rank matrix finally obtained, the local first low-rank matrix finally obtained, the local second low-rank matrix finally obtained and the local model, the method further includes: Training the local medical model through the first desensitized local data to obtain a corresponding personalized medical model for the client; Among them, the first desensitized local data is local data processed in a desensitization method with lower desensitization requirements among multiple desensitized local data.

6. A medical model training method, characterized in that: Applied to the server side, the method includes: Sending a global first low-rank matrix and a global second low-rank matrix to the client; wherein the global first low-rank matrix and the global second low-rank matrix are configured to train the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client; the local first low-rank matrix and the local second low-rank matrix are obtained by performing low-rank decomposition on the gradient of the local model; Update the global first low rank matrix and the global second low rank matrix according to the received trained local first low rank matrix and the local second low rank matrix; Sending the updated global first low-rank matrix and the global second low-rank matrix to the client; Update the global first low rank matrix and the global second low rank matrix according to the trained local first low rank matrix and the local second low rank matrix until the updates of the global first low rank matrix and the global second low rank matrix reach a set iteration condition; The final global first low-rank matrix and the final global second low-rank matrix are sent to the client; wherein the final global first low-rank matrix and the final global second low-rank matrix are configured to generate a local medical model with the final local first low-rank matrix, the final local second low-rank matrix and the local model.

7. The method according to claim 6, characterized in that Before sending the global first low-rank matrix and the global second low-rank matrix to the client, the method further includes: The global medical model is obtained by setting a universal model through training of a public data set; wherein the data in the public data set is data in a public data source in the corresponding medical field; The global medical model is sent to the client; wherein the global medical model is configured to perform an initial training on the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client.

8. A medical model training device, characterized in that: Applied to a client, the device comprises: A decomposition module, used for performing low-rank decomposition on the gradient of the local model to obtain a decomposed local first low-rank matrix and a local second low-rank matrix; A first training module, configured to train the local first low-rank matrix and the local second low-rank matrix based on the local model, the received global first low-rank matrix, the received global second low-rank matrix and local data of the client; A first sending module is used to send the trained local first low-rank matrix and the local second low-rank matrix to the server; wherein the global first low-rank matrix and the global second low-rank matrix are configured to be updated by the local first low-rank matrix and the local second low-rank matrix trained by the client; A first iteration module, used for training the local first low rank matrix and the local second low rank matrix by using the received updated global first low rank matrix, the updated global second low rank matrix and the local data until the training of the local first low rank matrix and the local second low rank matrix reaches a set iteration condition; A generation module is used to generate a local medical model based on the global first low-rank matrix, the global second low-rank matrix, the local first low-rank matrix, the local second low-rank matrix and the local model finally obtained in the above training process.

9. A medical model training device, characterized in that: Applied to the server side, the device comprises: A second sending module is used to send a global first low-rank matrix and a global second low-rank matrix to the client; wherein the global first low-rank matrix and the global second low-rank matrix are configured to train the local first low-rank matrix and the local second low-rank matrix with the local model of the client and the local data of the client; the local first low-rank matrix and the local second low-rank matrix are obtained by performing low-rank decomposition on the gradient of the local model; An updating module, configured to update the global first low-rank matrix and the global second low-rank matrix according to the received trained local first low-rank matrix and the local second low-rank matrix; The second sending module is further used to send the updated global first low-rank matrix and the global second low-rank matrix to the client; A second iteration module is used to update the global first low rank matrix and the global second low rank matrix according to the trained local first low rank matrix and the local second low rank matrix until the global first low rank matrix and the global second low rank matrix are updated to meet the set iteration condition; The second sending module is also used to send the final global first low-rank matrix and the final global second low-rank matrix to the client; wherein the final global first low-rank matrix and the final global second low-rank matrix are configured to generate a local medical model with the final local first low-rank matrix, the final local second low-rank matrix and the local model.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are executed.

Citation Information

Cited By

  • Distributed power model updating method and device based on model increment training and electronic equipment

    CN120408010A