Modal heterogeneous federated learning privacy protection method based on cross-modal prototype
By adopting a cross-modal prototype method in multimodal federated learning scenarios with modal heterogeneity and task heterogeneity, the problem of relying on strong assumptions and computational communication overhead is solved, efficient multimodal federated learning is achieved, application scenarios are expanded and communication overhead is reduced.
Patent Information
- Application Number
- CN202510119564.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-20
AI Technical Summary
The multimodal federated learning techniques in existing modal heterogeneous and task heterogeneous scenarios rely on strong assumptions and have high overhead for communication computing, making it difficult to achieve effective multimodal federated learning without relying on public data sets or with labels on each client.
Modal heterogeneous federated learning privacy protection method based on cross-modal prototypes is adopted to pass prototype information between the server and the client, aggregation and update of the model is realized, dependence on public data sets and labels is reduced, and prototypes are calculated through clustering methods, reducing calculation and communication overhead.
This method does not require strong assumptions, and can effectively expand application scenarios, reduce communication overhead, and retain multimodal paired information, helping single-modal clients obtain information about missing modal prototypes.
Smart Images

Figure CN120180463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a privacy protection method for modality heterogeneous federated learning based on cross-modal prototypes, belonging to the technical fields of artificial intelligence and information security. Background Art
[0002] With the development of information technology, the value of data as a key production factor has become increasingly prominent. However, due to the distribution characteristics of data and privacy protection requirements, traditional centralized data processing faces many challenges, including the risk of privacy leakage, regulatory compliance problems, and the data island effect. Federated learning technology provides a distributed training solution that allows multiple participants to collaboratively train a model without exposing local data, realizing the trusted sharing of data between different participants. Federated learning usually includes two types of participants: servers and clients. Among them, the client is responsible for local model training and sends the model parameters or gradients to the server. The server is responsible for aggregating the models or gradients of the clients and returning the aggregated results to the clients to help the clients update their local models.
[0003] Existing federated learning methods usually assume that clients have data of the same modality, and the data of each client is unimodal. However, with the progress of sensing devices, participants can collect data of multiple modalities (such as text, images, audio, etc.), and the unimodal federated learning training architecture can no longer meet the needs of collaborative training of multimodal clients. Multimodal federated learning technology focuses on how to use multimodal data to cooperate in training a model between different clients, further expanding the application scenarios of federated learning.
[0004] However, due to differences between devices, the modalities of data collected by different participants may be heterogeneous. For example, one device can collect text and images, while another device can only collect images. In addition, the inconsistency of modalities will also cause differences in model tasks. For example, clients with image-text data usually process multimodal tasks such as image-text retrieval and visual question answering, while clients with only one modality usually process unimodal tasks such as image classification and text classification. The differences in modalities and tasks lead to a decline in the performance of the model in multimodal federated learning.
[0005] Existing multimodal federated learning technologies in scenarios of modality heterogeneity and task heterogeneity mainly include the following three categories:
[0006] The first category is multimodal federated learning methods based on public datasets. This type of method either uses public datasets as prior knowledge to enhance missing modal information, thereby transforming modal heterogeneous multimodal federated learning into modal homogeneous multimodal federated learning; or uses public datasets containing all modal information as a medium for knowledge transfer to achieve local knowledge sharing between multimodal clients and unimodal clients. The model performance of this type of method depends on the quality of the public dataset.
[0007] The second category is prototype-based multimodal federated learning methods. In this type of method, the prototype is used to represent local modal information. The client first performs local training to obtain a local prototype and sends it to the server. The server aggregates the prototypes and sends the aggregated prototypes to the client. Each client obtains knowledge of other modalities from the prototypes aggregated by the server, thereby solving the problem of modal heterogeneity. This type of method usually assumes that each client has a label, while some multimodal tasks in actual scenarios (such as multimodal retrieval) do not have label information.
[0008] The third category is the block-based multimodal federated learning method. This type of method takes into account the heterogeneity of modalities and tasks, divides the model of each client into different modules, and realizes knowledge sharing among different clients by aggregating similar modules. However, in this type of method, all blocks of the client model participate in the aggregation process, which has high computational and communication overhead.
[0009] Therefore, in scenarios with heterogeneous modalities and tasks, how to achieve multimodal federated learning that does not rely on strong assumptions (such as the existence of a public dataset or that each client has a label) is a key technical problem that needs to be solved urgently. Summary of the invention
[0010] The purpose of the present invention is to solve the technical problems that the existing modality-heterogeneous multimodal federated learning technology relies on strong assumptions and has high communication and computing overhead, and creatively proposes a modality-heterogeneous federated learning privacy protection method based on cross-modal prototypes.
[0011] First, the relevant concepts and contents involved in the present invention are explained.
[0012] Federated learning: A distributed machine learning method in which clients participating in federated learning can collaboratively train models without exposing their privacy.
[0013] Single-modal client: refers to an entity that only has data of one modality and participates in the federated learning training process.
[0014] Multimodal client: refers to an entity that has two or more modal data and participates in the federated learning training process.
[0015] Server: refers to the entity that performs local prototype and model gathering and distribution in the federated learning process.
[0016] Prototype: refers to a set of representative samples extracted from the training data, which are used to characterize the core features of each category or cluster. The prototype can represent the data distribution characteristics.
[0017] Contrastive loss: Contrastive loss is used to learn data representation in unsupervised or semi-supervised situations, optimizing the data distribution in the embedding space by bringing similar pairs of samples closer and pushing dissimilar pairs of samples further apart.
[0018] Cosine similarity: It is an indicator that measures the degree of similarity between two vectors in direction, expressed by calculating the cosine value of the angle between the two vectors.
[0019] K-means clustering: A distance-based unsupervised clustering method that assigns data points to the cluster with the closest centroid and continuously updates the centroid position, ultimately dividing the data into K clusters to maximize the similarity of data within a cluster and minimize the difference between clusters.
[0020] The present invention is implemented by adopting the following technical solutions.
[0021] A privacy protection method for modality heterogeneous federated learning based on cross-modal prototypes includes the following steps:
[0022] Step 1: Model initialization.
[0023] The server first initializes the global model for each client. The global model includes the mapping module of the client's corresponding modality. The server sends the initialized model to the corresponding client. The multimodal client initializes a private prototype learning model.
[0024] Step 2: Using the global model, the single-modal client and the multi-modal client perform local model training respectively. This includes the following steps:
[0025] Step 2.1: The single-modal client completes local training using supervised learning methods and calculates local prototype values.
[0026] Specifically, it includes:
[0027] Step 2.1.1: The single-modal client first updates the mapping module of the local model to the mapping module of the corresponding modality in the global model.
[0028] Step 2.1.2: Under the guidance of cross entropy loss and global modal knowledge transfer loss, the single-modal client completes local training after multiple rounds of iterations to obtain a local model.
[0029] Among them, the local model includes an encoder, a mapping module and a classification module.
[0030] Step 2.1.3: After the training is completed, the unimodal client inputs the local data into the encoder and the mapping module to obtain the final sample embeddings. The unimodal client calculates the mean of the sample embeddings with the same label to obtain the local prototype.
[0031] Step 2.1.4: After completing the operations in Step 2.1.2 and Step 2.1.3, the unimodal client sends the local prototype and the mapping module in the local model to the server;
[0032] Step 2.2: The multimodal client uses the unsupervised learning method to complete the local training, including the prototype learning model training and the task model training, to obtain the local cross-modal prototype pairs and the local task model.
[0033] Specifically, it includes the following steps:
[0034] Step 2.2.1: The multimodal client uses the local data to complete the training of the prototype learning model, and the losses are the task loss, the intra-modal contrast loss, and the inter-modal contrast loss.
[0035] Step 2.2.2: After completing the training of the prototype learning model, the multimodal client fuses the information of multiple modalities to obtain the fused embeddings, and uses the K-means method to cluster the fused embeddings to obtain the pseudo-labels of the multimodal embedding pairs.
[0036] Step 2.2.3: The multimodal client calculates the mean of the embedding pairs with the same pseudo-label as the local image-text embedding pairs.
[0037] Step 2.2.4: The multimodal client trains the task model under the guidance of the local prototype learning model and the global prototype, and the losses are the task loss, the global modality knowledge transfer loss, and the local mapping module regularization loss.
[0038] Step 2.2.5: After completing Step 2.2.2 to Step 2.2.4, the multimodal client sends the local prototype pairs and the mapping module in the local task model to the server.
[0039] Step 3: After receiving the local prototype and the local model sent by the client, the server performs prototype aggregation and model aggregation respectively.
[0040] Step 3.1: After the server receives the local prototypes sent by the clients participating in the training, it first uses the multimodal prototype knowledge to perform modality completion on the prototypes of the unimodal clients so that the unimodal clients have the prototypes of the complete modalities. After the prototype completion, the server aggregates the local prototypes by the clustering method to obtain K global image-text prototype pairs.
[0041] Step 3.2: The server aggregates the client's local mapping modules using a model adaptive aggregation method.
[0042] Specifically, the steps include:
[0043] Step 3.2.1: For a mapping module of a certain modality of a certain client, the server calculates the similarity between the mapping module of the modality and the mapping modules of the same modality of other clients, and converts the similarity into a weight.
[0044] Step 3.2.2: The modality mapping module is multiplied by the weight to obtain the global mapping module of the client. The mapping modules of other modalities are aggregated in the same way.
[0045] Step 4: After receiving the global prototype pair and the global mapping module, the client updates the local mapping module and starts a new round of training.
[0046] Step 5: Repeat steps 2 to 4 until a specific number of training rounds is reached.
[0047] Beneficial Effects
[0048] Compared with the prior art, the method of the present invention has the following advantages:
[0049] 1. In the method of the present invention, information is transmitted between the server and the client through prototypes, without the need for prior knowledge of public data sets. Unlabeled multimodal clients can calculate prototypes by clustering, further expanding the application scenarios. Compared with existing methods, the training process of the present invention does not rely on strong assumptions (such as the existence of a public data set or each client has a label).
[0050] 2. In the method of the present invention, the model transmitted between the server and the client is only a part of the whole model. Compared with the third method, it can reduce the communication overhead and does not require that the local models have the same structure.
[0051] 3. In the method of the present invention, the local prototype calculated by the multimodal client and the prototype aggregation process of the server both retain the multimodal paired information, which is conducive to the single-modal client to obtain information from the missing modal prototype in a targeted manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0053] The present invention is further described below in conjunction with the accompanying drawings and embodiments. It should be noted that the implementation of the present invention is not limited to the following embodiments, and any form of modification or change made to the present invention will fall within the protection scope of the present invention.
[0054] Example
[0055] As shown Figure 1 in the figure, a privacy protection method for modality heterogeneous federated learning based on cross-modal prototypes.
[0056] This embodiment elaborates in detail the model training process in a modality heterogeneous federated learning scenario with an image unimodal client, a text unimodal client, and an image-text multimodal client.
[0057] This scenario has M I image clients, M T text clients, and M M image-text multimodal clients. Among them, the unimodal clients all have labels, and the multimodal clients have no label information. The tasks of the unimodal clients and the multimodal clients are classification tasks and multimodal retrieval tasks respectively. Each client has an encoder for the corresponding modality and a mapping module f i * , * ∈ {I, T}, and the unimodal clients also have a classification module I represents the image modality, and T represents the text modality.
[0058] It includes the following steps:
[0059] Step 1: Model initialization.
[0060] The server first initializes the global model. The server distributes the global model for the corresponding modality to the unimodal clients and the multimodal clients. The multimodal clients initialize the private prototype learning model.
[0061] In this embodiment, the server S initializes the model for each client, including (M I + M M ) image mapping modules and (M T + M M ) text mapping modules. The multimodal clients initialize the local private prototype learning model, including the image mapping module and the text mapping module.
[0062] Step 2: Using the global model, the image client, the text client, and the image-text client respectively perform local model training.
[0063] Step 2.1: The unimodal clients complete local training using the supervised learning method.
[0064] The specific implementation steps are as follows:
[0065] Step 2.1.1: The image client first updates the mapping module f i I of the local model to the mapping module initialized by the server where \(i\in[M I +M T +M M \).
[0066] Step 2.1.2: The image client completes local training through multiple rounds of iteration, and the loss is the cross-entropy loss and the global modality knowledge transfer loss \(L GKT . For the \(j\)-th sample of the image client, the embedding output by the mapping module is denoted as
[0067] Specifically, the calculation method of the global modality knowledge transfer loss \(L GKT is as follows:
[0068]
[0069] where, represents the probability that the embedding belongs to \(K\) global image prototypes; represents the probability that the embedding belongs to \(K\) global text prototypes; \(D JS represents the Jensen-Shannon divergence (JS divergence), and \(D KL represents the Kullback-Leibler divergence (KL divergence).
[0070] Step 2.1.3: After completing the training, the image client inputs local data into the local model to obtain the set of sample embeddings
[0071] The unimodal client calculates the mean of the sample embeddings with the same label to obtain the local prototype set where represents the number of categories, represents the local image prototype corresponding to the \(y\)-th category.
[0072] Step 2.1.4: The unimodal client sends the local prototype and the mapping module in the local model to the server \(S\).
[0073] Step 2.2: The multimodal client completes local training using unsupervised learning methods, including prototype learning model training and task model training, to obtain local cross-modal prototype pairs and local prototype learning models and task models.
[0074] Specifically, it includes the following steps:
[0075] Step 2.2.1: The multimodal client uses local data Complete the training of the prototype learning model. Then, the multimodal client inputs the local image-text pairs into the image-text encoders respectively, and the output results are passed into the image-text mapping module to obtain the image-text pair embeddings. The image-text pair embeddings are output through the fusion module to obtain the fused features. After that, the multimodal client uses the K-means method to cluster the fused features to obtain the pseudo-labels of the image-text pairs.
[0076] The loss of the above process is the multimodal retrieval task loss and the intra-modal contrast loss and the inter-modal contrast loss The intra-modal contrast loss encourages the embeddings of the same modality within the same cluster to be closer and closer, and the inter-modal contrast loss reduces the differences between different modalities, realizing the alignment of the image-text embeddings belonging to the same cluster.
[0077] During the training process of the prototype learning model, the total loss function L M is as follows:
[0078]
[0079] where L task represents the task loss, and N M represents the number of samples of the multimodal client.
[0080] Step 2.2.2: After completing the training of the prototype learning model, the multimodal client inputs the local data into the prototype learning model to obtain the pseudo-labels of the final image-text embeddings;
[0081] Step 2.2.3: The multimodal client calculates the means of the image embeddings and text embeddings with the same pseudo-labels respectively as the local image-text prototype pairs where K represents the number of local prototype pairs of the multimodal client.
[0082] Step 2.2.4: The multimodal client trains the task model under the guidance of the local prototype learning model and the global prototype. The training process is similar to the supervised learning process shown in Step 2.2.1.
[0083] The loss is the task loss L task 、the global modality knowledge transfer loss L GKT and the local mapping module regularization loss L LMR . The task loss is the loss function of multimodal retrieval. The purpose of global modality knowledge transfer is to encourage the multimodal client to obtain knowledge from the global prototype through self-supervised learning to enhance the local learning performance, and the calculation method is the same as that in Step 2.1.2. The purpose of the local mapping module regularization loss is to reduce the differences between the mapping module of the task model and the mapping module of the private prototype learning model. The calculation process is as follows:
[0084]
[0085] Among them, L LMR represents the regularization loss of the local mapping module; λ represents the balance factor used to balance local knowledge and global knowledge; θ * represents the mapping module of the task module, and represents the private mapping module obtained during the training process of the clustering model.
[0086] Step 2.2.5: After the multimodal client completes Steps 2.2.2 to 2.2.4, it sends the local prototype pair and the multimodal mapping module {θ *} *∈{I,T} in the local task model to the server.
[0087] Step 3: After receiving the local prototype and the local model sent by the client, the server performs prototype aggregation and model aggregation respectively.
[0088] Step 3.1: After the server receives the local prototypes sent by the unimodal clients and multimodal clients participating in the training, it first uses the multimodal prototypes to complement the prototypes of the unimodal clients, so that the unimodal clients have prototypes with complete modalities.
[0089] Specifically, it includes the following steps:
[0090] Step 3.1.1: The server first calculates the similarity between the image prototype of and the image prototype of the multimodal client, intercepts the top k most similar multimodal image prototypes, and records the corresponding text prototypes as
[0091] Step 3.1.2: The server converts the similarity into a weight value through softmax.
[0092] Step 3.1.3: The server multiplies the elements in by the corresponding weights to obtain the text prototypes paired with the image prototypes of , and forms new image-text prototype pairs.
[0093] Step 3.2: After the prototype complementation, the server aggregates all the local prototype pairs of the clients through the K-means clustering method to obtain K global image-text prototype pairs.
[0094] Specifically, it includes the following steps:
[0095] Step 3.2.1: The server first fuses the image-text prototype pairs after complementation, and the fused prototype is denoted as p iRepresents the image prototype, p t Represents the text prototype.
[0096] Step 3.2.2: The server clusters the fused prototypes, and the number of clusters is K.
[0097] Step 3.2.3: After clustering is completed, the server calculates the global image prototype and text prototype pairs. The global image prototype is the mean of the image local prototypes corresponding to the fused prototypes belonging to the same cluster. The global text prototype corresponding to the global image prototype is calculated in the same way, and finally K global image-text prototype pairs are obtained.
[0098] Step 3.3: The server aggregates the local mapping modules of the clients using the model adaptive aggregation method.
[0099] Taking the image client as an example, the process for the server to calculate the aggregated model is as follows:
[0100] Step 3.3.1: The server calculates the similarity between the image mapping module of and the image mapping modules of other clients, and converts the similarity into weights.
[0101] Step 3.3.2: The image mapping modules of all clients are multiplied by the weights of to obtain the global mapping module of . The mapping modules of other clients are aggregated according to the above method, and finally personalized aggregated models for each client are obtained.
[0102] Step 3.4: The server sends the aggregated cross-modal prototype pairs and the global mapping module to the corresponding clients.
[0103] Step 4: After receiving the global prototype pairs and the global mapping module, the client updates the local mapping module and starts a new round of training.
[0104] Step 5: Repeat Steps 2 to 4 until a specific number of training rounds is reached. At this time, the federated learning training is completed.
[0105] The above specific description further details the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A privacy protection method for modal heterogeneous federated learning based on cross-modal prototypes, characterized in that: The following steps are involved: Step 1: Model initialization; The server first initializes the global model for each client; the global model includes the mapping module of the client's corresponding modality; the server sends the initialized model to the corresponding client; the multimodal client initializes the private prototype learning model; Step 2: Using the global model, the single-modal client and the multi-modal client perform local model training respectively; Step 2.1: The single-modal client completes local training using a supervised learning method and calculates the local prototype value; Step 2.1.1: The single-modal client first updates the mapping module of the local model to the mapping module of the corresponding modality in the global model; Step 2.1.2: Under the guidance of cross entropy loss and global modal knowledge transfer loss, the single-modal client completes local training through multiple rounds of iterations to obtain a local model; Among them, the local model includes an encoder, a mapping module, and a classification module; Step 2.1.3: After completing the training, the unimodal client inputs the local data into the encoder and mapping module to obtain the final sample embedding; the unimodal client calculates the mean of the sample embeddings with the same label to obtain the local prototype; Step 2.1.4: After completing the operations of step 2.1.2 and step 2.1.3, the single-modal client sends the mapping module in the local prototype and the local model to the server; Step 2.2: The multimodal client completes local training using an unsupervised learning method, including prototype learning model training and task model training, to obtain local cross-modal prototype pairs and local task models; Step 2.2.1: The multimodal client completes the training of the prototype learning model using local data. The losses are task loss, intra-modality contrast loss, and inter-modality contrast loss. Step 2.2.2: After completing the training of the prototype learning model, the multimodal client fuses the information of multiple modalities to obtain the fused embedding, and clusters the fused embedding using the K-means method to obtain the pseudo-label of the multimodal embedding pair; Step 2.2.3: The multimodal client calculates the mean of the embedding pairs with the same pseudo-label as the local image-text embedding pair; Step 2.2.4: The multimodal client trains the task model under the guidance of the local prototype learning model and the global prototype, and the loss is the task loss, the global modality knowledge transfer loss and the local mapping module regularization loss; Step 2.2.5: After completing steps 2.2.2 to 2.2.4, the multimodal client sends the local prototype pair and the mapping module in the local task model to the server; Step 3: After receiving the local prototype and local model sent by the client, the server performs prototype aggregation and model aggregation respectively; Step 4: After receiving the global prototype pair and the global mapping module, the client updates the local mapping module and starts a new round of training; Step 5: Repeat steps 2 to 4 until a specific number of training rounds is reached.
2. The method for privacy protection of modal heterogeneous federated learning based on cross-modal prototypes as claimed in claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: After receiving the local prototype sent by the client participating in the training, the server first uses the multimodal prototype knowledge to complete the modality of the prototype of the single-modal client, so that the single-modal client has a prototype with a complete modality; after the prototype is completed, the server aggregates the local prototypes by clustering to obtain K global image-text prototype pairs; Step 3.2: The server aggregates the client's local mapping module using a model adaptive aggregation method; Step 3.2.1: For a mapping module of a certain modality of a certain client, the server calculates the similarity between the mapping module of the modality and the mapping modules of the same modality of other clients, and converts the similarity into a weight; Step 3.2.2: The modality mapping module is multiplied by the weight to obtain the global mapping module of the client; the mapping modules of other modalities are aggregated in the same way.
3. The method for privacy protection of modal heterogeneous federated learning based on cross-modal prototypes as claimed in claim 1, characterized in that: In step 2.1.2, the global modal knowledge transfer loss L GKT The calculation method is as follows: in, Representation Embedding The probability of belonging to K global image prototypes; Represents the embedding of the output of the mapping module The probability of belonging to K global text prototypes; D JS Denotes Jensen-Shannon divergence, D KL represents the Kullback-Leibler divergence; I represents the image modality, and T represents the text modality.
4. The method for privacy protection of modal heterogeneous federated learning based on cross-modal prototypes as claimed in claim 1, characterized in that: In step 2.2.1, the total loss function L M As shown below: Among them, L task represents the task loss, N M Indicates the number of samples of multimodal clients; is the intra-modality contrast loss, is the inter-modality contrast loss; I represents the image modality and T represents the text modality.
5. The method for privacy protection of modal heterogeneous federated learning based on cross-modal prototypes as claimed in claim 1, characterized in that: In step 2.2.4, the loss is the task loss L task , Global modal knowledge transfer The purpose of the local mapping module regularization loss is to narrow the difference between the task model mapping module and the private prototype learning model mapping module. The calculation process is as follows: Among them, L LMR represents the regularization loss of the local mapping module; λ represents the balance factor, which is used to balance local knowledge and global knowledge; θ * A mapping module representing a task module, Represents the private mapping module obtained during clustering model training.
6. The method for privacy protection of modal heterogeneous federated learning based on cross-modal prototypes as described in claim 2, characterized in that: Step 3.1 includes the following steps: Step 3.1.1: The server first calculates the image client The similarity between the image prototype of and the image prototype of the multimodal client is intercepted, and the first k most similar multimodal image prototypes are recorded as Step 3.1.2: The server converts the similarity into a weight value through softmax; Step 3.1.3: The server will Multiply the elements in by the corresponding weights to get The image prototype is paired with the text prototype to form a new image-text prototype pair.
7. The method for privacy protection of modal heterogeneous federated learning based on cross-modal prototypes as described in claim 2, characterized in that: Step 3.2 includes the following steps: Step 3.2.1: The server first fuses the completed image-text prototype pair, and the fused prototype is recorded as p i Represents the image prototype, p t Represents a text prototype; Step 3.2.2: The server clusters the fused prototypes, and the number of clusters is K; Step 3.2.3: After clustering is completed, the server calculates the global image prototype and text prototype pairs. The global image prototype is the mean of the local image prototypes corresponding to the fusion prototype belonging to the same cluster; the global text prototype corresponding to the global image prototype is calculated in the same way, and finally K global image-text prototype pairs are obtained.
Citation Information
Cited By
Personalized federal learning method and system based on domain invariant text representation and intra-domain global prior
CN120508883A
Multi-modal federal cross-domain fault diagnosis method based on prototype comparative learning and application
CN121980502A