Non-equilibrium mode-oriented federal size model collaborative task enhancement method

By generating balanced multimodal data on the server side and adopting a bidirectional knowledge distillation mechanism and a dynamic weight aggregation strategy, the problem of uneven distribution of multimodal data in federated learning is solved, the accuracy and generalization ability of the model are improved, and efficient collaboration between resource-constrained devices and large cloud models is achieved.

CN120597995APending Publication Date: 2025-09-05HENAN POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510738486.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In federated learning, the uneven distribution of multimodal data leads to a decline in model performance. Existing methods find it difficult to balance privacy protection and multimodal collaborative optimization, and traditional knowledge distillation technology lacks a two-way interaction mechanism, resulting in insufficient model generalization capabilities.

Method used

By generating balanced multimodal data on the server side and adopting a two-way knowledge distillation mechanism, combined with a dynamic weight aggregation strategy, knowledge transfer between small and large models and updating of global model parameters are achieved.

Benefits of technology

It significantly improves the accuracy and generalization ability of the model in cross-modal tasks, realizes efficient collaboration between resource-constrained devices and large cloud models, and provides a safe and reliable multimodal federated learning solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597995A_ABST
    Figure CN120597995A_ABST
Patent Text Reader

Abstract

The invention relates to an unbalanced modal-oriented federal size model collaborative task enhancement method, which comprises the following steps of: obtaining balanced multi-modal data generated at a server side, respectively inputting the multi-modal data into different small models, and obtaining output characterization; the small model is obtained by training a training set; the training set comprises an equipment fault data set; and in the process of training the small models by using the training set, jointly training a cross-modal anomaly detection model as a large model by using the small models, performing knowledge distillation between the small models and the large model and between different small models by using a bidirectional knowledge distillation mechanism, and introducing a dynamic weight aggregation strategy to update global model parameters. According to the method, the problems of unbalanced modal distribution and privacy disclosure in federated learning are effectively solved, the model performance and generalization ability of a cross-modal task are improved, and the method is suitable for demand scenes such as intelligent manufacturing and industrial detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning and multimodal data processing, and in particular to a method for enhancing federated large and small model collaborative tasks for unbalanced modalities. Background Art

[0002] In the current federated learning framework, the problem of uneven distribution of multimodal data is prevalent. For example, in the field of intelligent manufacturing, the data modalities of various manufacturing enterprises vary significantly, resulting in the performance of the global model jointly trained by various manufacturing enterprises being degraded due to modal skew. Existing methods mostly use one-way knowledge distillation or simple data aggregation, which makes it difficult to balance privacy protection and multimodal collaborative optimization requirements. At the same time, traditional knowledge distillation technology focuses on transferring knowledge from small models to large models, lacks a two-way interaction mechanism, and limits the complementarity between models. In addition, the synthetic data generated under uneven data often finds it difficult to cover the diversity of multimodal distributions, exacerbating the problem of insufficient model generalization ability. Therefore, there is an urgent need for a collaborative optimization method that takes into account privacy security, modal balance, and two-way knowledge transfer to improve the effectiveness of federated learning in uneven multimodal scenarios. Summary of the Invention

[0003] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a collaborative task enhancement method for federated large and small models of unbalanced modalities, which solves the challenge of unbalanced multimodal data distribution through the deep integration of server-side global data enhancement and bidirectional knowledge distillation.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A method for enhancing collaborative tasks of federated large and small models with non-balanced modes, including:

[0006] Obtaining balanced multimodal data generated by the server, the multimodal data being used to represent information about devices with different faults, inputting the multimodal data into different small models to obtain output representations, the small models being used to represent the local model; the small models being trained using a training set; the training set comprising: a data set of device faults;

[0007] In the process of using the training set to train the small model, the small model is used to jointly train the cross-modal anomaly detection model as the large model. The large model is used to characterize the server model. A bidirectional knowledge distillation mechanism is used to perform knowledge distillation between the small model and the large model and between different small models, and a dynamic weight aggregation strategy is introduced to update the global model parameters.

[0008] Optionally, obtaining the balanced multimodal data generated by the server side includes:

[0009] Obtaining local privacy-protected domain data generated by the server corresponding to different small models; the domain data includes: device images of different fault types, fault reports and maintenance records, and image-text pairs of images and quality inspection text;

[0010] The domain data is screened and enhanced, and the original data distribution is statistically analyzed. Based on the original data distribution, the domain data is extracted according to a target ratio to obtain the multimodal data.

[0011] Optionally, performing knowledge distillation between the small model and the large model and between different small models includes:

[0012] Inputting the training set into the large model and different small models respectively, constructing different cross entropy loss functions output by the large model and the small models;

[0013] Combining the different cross entropy loss functions, generating loss functions for different modalities, adding the loss functions, and minimizing the loss functions, thereby performing knowledge distillation between the small model and the large model and between different small models.

[0014] Optionally, the different small models include:

[0015] A first small model is used to process the device images in the training set and output a local image representation;

[0016] The second small model is used to process the fault reports and maintenance records in the training set and output local text representations;

[0017] The third small model is used to process the image-text pairs in the training set and output a joint representation.

[0018] Optionally, the different cross entropy loss functions include:

[0019] A first cross entropy loss function is used to enhance the small model using the knowledge of the large model on the same data modality;

[0020] A second cross entropy loss function is used to supplement the small model knowledge with the large model knowledge on different data modalities;

[0021] The third cross entropy loss function is used to distill the local historical task knowledge in the small model to the current task.

[0022] Optionally, generating the loss function includes:

[0023] The loss function for the image modality is:

[0024]

[0025] in, is the local image representation, The text representation output by the large model, Image representation of the output of the large model, For joint representation, L inter , L intra , L acc Both are cross entropy loss functions;

[0026] The loss function for text modality is:

[0027]

[0028] in, is the local text representation, is the output representation of the k-th text in the local model;

[0029] The loss function for the image-text modality is:

[0030]

[0031] in, represents the output representation of the k-th image on the local last round model.

[0032] Optionally, introducing the dynamic weight aggregation strategy to update the global model parameters of the large model includes:

[0033] Obtaining the representation weights output by the small model, normalizing them to obtain target weights, and aggregating the small model based on the target weights;

[0034] The difference between the image representation and text representation output by the large model and the aggregation result is calculated, and the global model parameters are updated by minimizing the difference.

[0035] Optionally, aggregating the small models based on the target weight includes:

[0036]

[0037] Among them, i (k) , t (k) is the aggregation result, represents the output of the k-th image sample on the small model of user i, α (k,i) , β (k,i) is the target weight, represents the output of the k-th text sample on the small model of user i.

[0038] Optionally, calculating the difference between the image representation and the text representation output by the large model and the aggregation result includes:

[0039]

[0040] Among them, θ pub is the global model parameter.

[0041] Optionally, minimizing the difference and updating the global model parameters includes:

[0042]

[0043] Among them, [i (k) ,t (k) ] is the aggregation result after vector splicing, η is the learning rate, D pub is public multimodal data, is the global model parameter value at t iterations, is the gradient operator.

[0044] The beneficial effects of the present invention are:

[0045] This invention effectively solves the problems of uneven modal distribution and privacy protection in federated learning by constructing a balanced multimodal public dataset and a two-way knowledge distillation mechanism on the server side, significantly improving the accuracy and generalization ability of the model in cross-modal tasks (such as image and text retrieval and multimodal classification). At the same time, through lightweight small models and dynamic weight aggregation strategies, efficient collaboration between resource-constrained devices and large cloud models is achieved, taking into account both computing efficiency and task performance, providing a safe and reliable multimodal federated learning solution for demand scenarios such as intelligent manufacturing and industrial inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 This is a flow chart of a method for enhancing collaborative tasks of federated large and small models in an unbalanced mode according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] In the field of smart manufacturing, multiple manufacturers need to jointly train a cross-modal anomaly detection model to identify equipment failures and abnormal behavior in real time on the production line. However, the data modalities of each manufacturer vary significantly: Manufacturer A provides images of equipment in operation (image modality) to capture visual and thermal imaging anomalies; Manufacturer B provides production and maintenance log text (text modality) to analyze fault descriptions and repair records; and Manufacturer C provides paired images and corresponding quality inspection reports (image-text pairs) to compare image and text information to identify abnormal patterns.

[0051] like Figure 1 As shown, this embodiment discloses a method for enhancing the collaborative task of federated large and small models for unbalanced modalities, including: obtaining balanced multimodal data generated on the server side, the multimodal data is used to characterize equipment information of different faults, inputting the multimodal data into different small models respectively, obtaining output representations, and the small models are used to characterize local models; the small models are trained using training sets; the training sets include: equipment fault data sets; in the process of training the small models using the training sets, the small models are used to jointly train a cross-modal anomaly detection model as a large model, the large model is used to characterize the server model, and a bidirectional knowledge distillation mechanism is used to perform knowledge distillation between the small model and the large model and between different small models, and a dynamic weight aggregation strategy is introduced to update the global model parameters.

[0052] Furthermore, obtaining balanced multimodal data generated on the server side includes: obtaining local privacy-protected domain data generated on the server side corresponding to different small models; the domain data includes: device images of different fault types, fault reports and maintenance records, and image-text pairs of images and quality inspection text;

[0053] The domain data is filtered and enhanced, and the original data distribution is statistically analyzed. Based on the original data distribution, the domain data is extracted according to the target ratio to obtain multimodal data.

[0054] Specifically, balanced multimodal data is generated on the server side. To alleviate the problem of reduced model utility caused by model imbalance, each user first generates privacy-protected domain data locally, and then aggregates it on the server side according to the principles of diversity and balance. This step includes the following two steps:

[0055] Manufacturing company A uses SGD training generators locally to synthesize equipment images of different fault types, and uploads them after desensitization using the differential privacy mechanism; manufacturing company B fine-tunes text generation models locally (such as the GPT series) to generate synthetic fault reports and maintenance records, and filters sensitive information; manufacturing company C uses a multimodal generation framework to locally generate matching pairs of images and quality inspection texts, and ensures the diversity and privacy of image-text alignment. Each manufacturing company only uploads synthetic data to the server.

[0056] The server filters and enhances the user's data according to the principles of diversity and balance, ensuring that the server has balanced multimodal data for both the server and the user to access. The server statistics show that the original data distribution is: 60% images, 30% text, and 10% image-text pairs. By sampling the image-text pair data of manufacturing company C and screening the diversity samples of manufacturing companies A and B, a balanced public dataset D is generated. pub , containing 33% images, 33% text and 33% image-text pairs.

[0057] Furthermore, knowledge distillation between the small model and the large model, as well as between different small models, includes: inputting the training set into the large model and different small models respectively, constructing different cross-entropy loss functions for the outputs of the large model and the small model; combining different cross-entropy loss functions to generate loss functions for different modalities, adding the loss functions, and minimizing the loss functions, thereby performing knowledge distillation between the small model and the large model, as well as between different small models.

[0058] Furthermore, different small models include: a first small model, used to process device images in the training set and output local image representations; a second small model, used to process fault reports and maintenance records in the training set and output local text representations; a third small model, used to process image-text pairs in the training set and output joint representations.

[0059] Furthermore, different cross-entropy loss functions include: a first cross-entropy loss function, which is used to enhance the small model with the knowledge of the large model on the same data modality; a second cross-entropy loss function, which is used to supplement the small model knowledge with the knowledge of the large model on different data modalities; and a third cross-entropy loss function, which is used to distill the local historical task knowledge in the small model to the current task.

[0060] Specifically, knowledge is distilled from the large model to the small models and between the small models. Using a public dataset, the knowledge of the large model and other small models is distilled into the current user's small model. This step includes the following four steps:

[0061] 1. Input public datasets into large models (such as CLIP) to extract image representations and text representation

[0062] 2. Input the public data into the small model of each manufacturing enterprise to obtain the local small model output representation. A manufacturing enterprise’s MobileNet processes the device image i in the public dataset. (k) , output local image representation TinyBERT at manufacturing company B processes maintenance report text and outputs local text representation The multimodal small model of the manufacturing company C processes the image and text pairs simultaneously and outputs a joint representation

[0063] 3. Construct a cross entropy loss function between the output of the large model, the output of other small models (if the current small model is the model trained by manufacturing company A, the other small models refer to the models trained by manufacturing companies B and C) and the output of the current small model, and transfer the knowledge of the large model and other small models to the current small model. The idea is to use the knowledge of the large model to enhance the user's small model (the loss function is L intra ), for example: A manufacturing company’s MobileNet uses the loss function L intra ,make Approximating CLIP Using large model knowledge to supplement user small model knowledge on different data modalities (loss function is L inter ). Through the image representation of A manufacturing company Text Representation with CLIP Align and distill the local historical task knowledge of the small model to the current task (the loss function is L acc ). Combined with the loss function L acc , ensuring that after training for the new task, MobileNet can still accurately identify the features in the historical equipment fault identification task. Represents the output representation of the k-th image and text on the local last round model.

[0064] For image modality users, the loss function is:

[0065]

[0066] For text modality users, the loss function is:

[0067]

[0068] For users in the image and text mode, the loss function is:

[0069]

[0070] 4. Local model update. Add formulas (1), (2), and (3) to minimize the loss function, and achieve distillation of the small model and the large model between data modalities, within modalities, and the historical knowledge of the small model to the current small model.

[0071] Furthermore, a dynamic weight aggregation strategy is introduced to update the global model parameters of the large model, including: obtaining the representation weights output by the small models, normalizing them, obtaining the target weights, and aggregating the small models based on the target weights; calculating the differences between the image representation and text representation output by the large model and the aggregation results, minimizing the differences, and updating the global model parameters.

[0072] Specifically, knowledge distillation from the small model to the large model is performed. The performance differences between the large and small models on public datasets are leveraged to migrate the knowledge from the small model to the large model. For the current large model, different weights are assigned to the user knowledge that matches it, and the knowledge from the small model is distilled to the large model. This step includes the following two steps:

[0073] Calculate the small model representation weights. Assume represents the local image representation output of the kth sample on the small model of user i. Then, during inter-modal knowledge distillation, Relative to The weight is:

[0074]

[0075] Relative to The weight is:

[0076]

[0077] For s (k,i) ,t (k,i) Normalize to get the weight α (k,i) ,β (k,i) .

[0078] Small model knowledge distillation large model. Aggregate local small models to obtain Computational large models Representation and small model aggregation results on i (k) ,t (k) The difference Minimize the difference and update the global model parameters:

[0079]

[0080] Among them, [i (k) ,t (k) ] represents the concatenation of vectors.

[0081] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for enhancing collaborative tasks of federated large and small models in non-equilibrium mode, characterized by: include: Obtaining balanced multimodal data generated by the server, wherein the multimodal data is used to represent information about devices with different faults, inputting the multimodal data into different small models, and obtaining output representations, wherein the small models are used to represent the local model; The small model is obtained by training using a training set; the training set includes: an equipment failure data set; In the process of using the training set to train the small model, the small model is used to jointly train the cross-modal anomaly detection model as the large model. The large model is used to characterize the server model. A bidirectional knowledge distillation mechanism is used to perform knowledge distillation between the small model and the large model and between different small models, and a dynamic weight aggregation strategy is introduced to update the global model parameters.

2. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 1 is characterized in that: Acquiring the balanced multimodal data generated by the server side includes: Obtaining local privacy-protected domain data generated by the server corresponding to different small models; the domain data includes: device images of different fault types, fault reports and maintenance records, and image-text pairs of images and quality inspection text; The domain data is screened and enhanced, and the original data distribution is statistically analyzed. Based on the original data distribution, the domain data is extracted according to a target ratio to obtain the multimodal data.

3. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 1 is characterized in that: Performing knowledge distillation between the small model and the large model and between different small models includes: Inputting the training set into the large model and different small models respectively, constructing different cross entropy loss functions output by the large model and the small models; Combining the different cross entropy loss functions, generating loss functions for different modalities, adding the loss functions, and minimizing the loss functions, thereby performing knowledge distillation between the small model and the large model and between different small models.

4. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 3 is characterized in that: The different small models include: A first small model is used to process the device images in the training set and output a local image representation; The second small model is used to process the fault reports and maintenance records in the training set and output local text representations; The third small model is used to process the image-text pairs in the training set and output a joint representation.

5. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 3 is characterized in that: The different cross entropy loss functions include: A first cross entropy loss function is used to enhance the small model using the knowledge of the large model on the same data modality; A second cross entropy loss function is used to supplement the small model knowledge with the large model knowledge on different data modalities; The third cross entropy loss function is used to distill the local historical task knowledge in the small model to the current task.

6. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 3 is characterized in that: Generating the loss function includes: The loss function for the image modality is: in, is the local image representation, The text representation output by the large model, Image representation of the output of the large model, For joint representation, L inter , L intra , L acc Both are cross entropy loss functions; The loss function for text modality is: in, is the local text representation, is the output representation of the k-th text in the local model; The loss function for the image-text modality is: in, represents the output representation of the k-th image on the local last round model.

7. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 1 is characterized in that: Introducing the dynamic weight aggregation strategy to update the global model parameters of the large model includes: Obtaining the representation weights output by the small model, normalizing them to obtain target weights, and aggregating the small model based on the target weights; The difference between the image representation and text representation output by the large model and the aggregation result is calculated, and the global model parameters are updated by minimizing the difference.

8. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 7 is characterized in that: Aggregating the small models based on the target weights includes: Among them, i (k) , t (k) is the aggregation result, represents the output of the k-th image sample on the small model of user i, α (k,i) , β (k,i) is the target weight, represents the output of the k-th text sample on the small model of user i.

9. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 7, characterized in that: Calculating the difference between the image representation and text representation output by the large model and the aggregation result includes: Among them, θ pub is the global model parameter.

10. The method for enhancing the collaborative task of federated large and small models in an unbalanced mode according to claim 7, characterized in that: Minimizing the difference to update the global model parameters includes: Among them, i (k) ,t (k) is the aggregation result after vector splicing, η is the learning rate, D pub is public multimodal data, is the global model parameter value at t iterations, is the gradient operator.