Federal learning big and small model collaborative optimization method based on cross-modal knowledge fusion

By adopting the method of cross-modal knowledge fusion in federated learning, the synthetic data set is constructed and optimized and semantic alignment, the resource and efficiency challenges of text-generated image models in federated learning in professional fields are solved, and the model performance improvement and resource conservation are achieved.

CN120163261APending Publication Date: 2025-06-17TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510048401.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing federated learning is difficult to carry out professional field data training and enhancement of text-generated image models while ensuring the security of private data. Especially in the federated fine-tuning of large-scale models, resource and efficiency challenges are great.

Method used

The federated learning size model collaborative optimization method based on cross-modal knowledge fusion is adopted. The preset text generation image model on the server and the image generation text model acquired by calling the API interface is used to generate images and their corresponding text descriptions. The synthetic data set is constructed, and data optimization and semantic alignment are performed through tag voting and multimodal representation fusion. Finally, the big model on the server and small models on the client are fine-tuned and trained.

Benefits of technology

On the premise of ensuring the security of private data, the performance of text-generated image models is effectively enhanced, the performance of the model on specific tasks is improved, and the consumption of resource and communication bandwidth is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163261A_ABST
    Figure CN120163261A_ABST
Patent Text Reader

Abstract

The invention provides a federal learning size model collaborative optimization method based on cross-modal knowledge fusion, which comprises the following steps: generating an image and corresponding text description through a text generation image model preset by a server and an image generation text model obtained by calling an API interface, and constructing a synthetic data set through the image and the text description; performing data optimization through label voting based on the synthetic data set, constructing unified representation through multi-modal representation fusion based on the optimized synthetic data set, and completing cross-client semantic alignment; and performing fine tuning training on the large model of the server and the small model of the client based on the unified representation to complete collaborative optimization. The method solves the problem that the existing federal learning is difficult to carry out professional field data training enhancement on the text generation image model on the premise of ensuring the security of private data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-modal federated learning, and particularly to a collaborative optimization method for large and small models of federated learning based on cross-modal knowledge fusion. Background Art

[0002] Related models for generating images from text have been rapidly developed in various application fields. To improve the performance of a pre-trained text-to-image generation model on a specific task, targeted fine-tuning is required, but this is usually restricted by privacy constraints of domain-specific data. Therefore, there is an urgent need to develop innovative methods to fine-tune the model while ensuring data privacy and security.

[0003] Federated Learning (FL) is a paradigm that allows multiple clients to collaboratively train a model without directly exchanging data. In traditional methods, FL usually trains a local model on each client and then aggregates the updated parameters in a central server. Directly applying this method to train a large-scale text-to-image generation model requires a large amount of computing resources and communication bandwidth on distributed client devices, which is somewhat difficult to implement. To address the resource and efficiency challenges in federated fine-tuning of large generative models, federated parameter-efficient fine-tuning methods have been proposed in the prior art, which freeze most of the pre-trained model and only update selected layers. However, the client still needs to locally host a large-scale model. The split training strategy distributes the computation to the server at the cost of communication cost and privacy risk. Another study established cross-silo knowledge edge transfer between a client small model (SM) and a server large language model (LM). However, existing methods are limited to unimodal clients.

[0004] With the development of mobile intelligence and the Internet of Things (IoT), a large amount of multimodal data has been generated. Integrating this multimodal data shows great promise. However, effectively utilizing multimodal data in FL remains a huge challenge. In such an environment, the heterogeneity of data modalities, distributions, and model architectures makes direct knowledge alignment and fusion extremely difficult. Existing multimodal methods rely on public datasets to unify multimodal data from different sources into a consistent and aligned representation.

[0005] To overcome these challenges, a new federated multimodal knowledge transfer framework is needed to achieve collaborative training and mutual enhancement of heterogeneous multimodal client models (SMS) and a server text-to-image generation model without exposing private data, while promoting efficient and effective cross-island knowledge transfer. Summary of the Invention

[0006] The present invention provides a collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion, which is used to solve the problem that it is difficult for existing federated learning to enhance the training of text-to-image models with professional domain data while ensuring the security of private data.

[0007] The present invention provides a collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion, including: Generating an image and its corresponding text description through a text-to-image model preset by the server and an image-to-text model obtained by calling an API interface, and constructing a synthetic dataset through the image and the text description; Optimizing the data through label voting based on the synthetic dataset, and constructing a unified representation through multi-modal representation fusion based on the optimized synthetic dataset to complete cross-client semantic alignment; Fine-tuning and training the large model on the server and the small model on the client respectively based on the unified representation to complete collaborative optimization.

[0008] According to the collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention, the step of generating an image and its corresponding text description through a text-to-image model preset by the server and an image-to-text model obtained by calling an API interface, and constructing a synthetic dataset through the image and the text description specifically includes: Randomly selecting a class label from the existing historical database on the server and converting the class label into a label description; Inputting the label description into the text-to-image model preset by the server to generate an image; Inputting the image into the image-to-text model obtained by calling the API interface to generate a corresponding text description; Forming a synthetic dataset through the image and the text description and sending it to the client.

[0009] According to the collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention, the step of optimizing the data through label voting based on the synthetic dataset specifically includes: Extracting local prediction labels on the client based on the synthetic dataset; Counting the number of votes for each category of each piece of synthetic data in the synthetic dataset based on the local prediction labels, and determining the consensus label through majority voting; In the case where the number of votes for the consensus label is less than the set value of the total number of clients, the corresponding data is discarded to obtain an optimized synthetic dataset.

[0010] A collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention. The unified representation is constructed through multi-modal representation fusion based on the optimized synthetic dataset to complete cross-client semantic alignment, specifically including: Extract image representations through clients based on the optimized synthetic dataset; Perform multi-modal representation alignment through the cross-attention mechanism between modalities based on the image representations; Construct a unified representation through contrastive aggregation based on the image representations after representation alignment to complete cross-client semantic alignment.

[0011] A collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention. The unified representation is constructed through contrastive aggregation based on the image representations after representation alignment to complete cross-client semantic alignment, specifically including: Compare and weight the image representations after representation alignment with the average representation; Calculate the representation weight values of the clients according to the set formula, perform aggregation and normalization operations on the representations with high weight values according to the representation weight values, construct a unified representation, and complete cross-client semantic alignment.

[0012] A collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention. The large model on the server side and the small model on the client side are fine-tuned and trained respectively based on the unified representation to complete collaborative optimization, specifically including: Optimize the decoder of the text-to-image model on the server side based on the unified representation, and introduce Bernoulli random variables for fine-tuning of the decoder to complete the fine-tuning training of the large model on the server side; Re-train the local model by combining the unified representation with the global representation on the client side, align the unified representation with the global representation using the federated learning loss, and align the local prediction labels with the consensus labels using the cross-entropy loss to complete the fine-tuning training of the small model on the client side.

[0013] The present invention also provides a collaborative optimization system for large and small models in federated learning based on cross-modal knowledge fusion. The system includes: A cross-modal data synthesis module for generating images and their corresponding text descriptions through a text-to-image model preset on the server side and an image-to-text model obtained by calling an API interface, and constructing a synthetic dataset through the images and text descriptions; A cross-client semantic alignment module for optimizing data through label voting based on the synthetic dataset, and constructing a unified representation through multi-modal representation fusion based on the optimized synthetic dataset to complete cross-client semantic alignment; A knowledge transfer module for fine-tuning and training the large model on the server side and the small model on the client side respectively based on the unified representation to complete collaborative optimization.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion as described in any one of the above.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion as described in any one of the above.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion as described in any one of the above.

[0017] A collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention synergistically enhances the text-to-image model on the server side and the large model on the client side through multi-modal knowledge transfer, and without sharing customer data or model parameters, ensuring privacy; using synthetic data as an intermediary for knowledge transfer, combining label voting to effectively improve data quality, and integrating multi-modal representations from heterogeneous clients to promote two-way knowledge transfer between the text-to-image model on the server side and the client model. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of the collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention.

[0020] Figure 2 It is an architecture diagram of the collaborative optimization system for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention.

[0021] Figure 3 It is a schematic diagram of the module connection of the collaborative optimization system for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention.

[0022] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention.

[0023] Reference numerals: 110: Cross-modal data synthesis module; 120: Cross-client semantic alignment module; 130: Knowledge transfer module; 410: Processor; 420: Communication interface; 430: Memory; 440: Communication bus. Detailed implementation manners

[0024] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0025] The following Figure 1 describes a collaborative optimization method for large and small models of federated learning based on cross-modal knowledge fusion according to the present invention, including: Step 100, generating an image and its corresponding text description through a text-to-image model preset by the server and an image-to-text model obtained by calling an API interface, and constructing a synthetic dataset through the image and the text description.

[0026] Specifically, randomly select a class label from the existing historical database of the server, and convert the class label into a label description; Input the label description into the text-to-image model preset by the server to generate an image; Input the image into the image-to-text model obtained by calling the API interface to generate a corresponding text description; Form a synthetic dataset through the image and the text description and send it to the client.

[0027] In the present invention, since private data cannot be directly obtained, first use the generation ability of the text-to-image model (T2I model) of the server to synthesize a dataset in a specific field , for transmitting knowledge between the client and the server. First, randomly select a class label , and convert it into a label description prompt , such as "a picture of ". Then, input the prompt to to generate the corresponding picture .

[0028] To make full use of the data in the fields of images and texts, further use the API of the pre-trained image-to-text model (I2T model) to obtain the text descriptions corresponding to each generated image, denoted as . Note that the generated by the I2T model provides a richer description of the image content, capturing specific features and contexts. The constructed synthetic dataset will be distributed to each client according to the modality, that is, distributed to the text client, distributed to the image client.

[0029] Step 200: Optimize the data through label voting based on the synthetic dataset, and construct a unified representation through multi-modal representation fusion based on the optimized synthetic dataset to complete cross-client semantic alignment.

[0030] In the present invention, the client receives the synthetic dataset transmitted by the server, and each client uses the local model to perform local inference. Specifically, is split into a feature extractor and a classifier . is used to map the input space into a representation, is used to map the representation to the category space. The picture representation k extracted from the synthetic data by the picture client and the local prediction label are obtained according to the following formula: For the text client, and are obtained in the same way. Since the client uses local private data to train the local model, the representations generated by the local client are significantly different, which poses a challenge to effectively aligning and integrating them into a unified global representation. To efficiently integrate the multi-modal knowledge of different clients in the professional field, a two-step fusion strategy is performed: label voting (LabVote) and multi-modal representation fusion (MultiRepFusion). The goal of the LabVote label voting design is to correct the synthetic data labels and eliminate unreliable samples to ensure the quality of the synthetic data. The MultiRepFusion multi-modal representation fusion focuses on integrating the multi-modal domain knowledge from different clients to create a comprehensive and rich representation for promoting knowledge transfer.

[0031] Among them, optimizing the data through label voting based on the synthetic dataset includes: Extract local prediction labels at the client based on the synthetic dataset; Based on the local prediction labels, count the number of votes for each category in each synthetic data in the synthetic dataset, and determine the consensus label through majority voting; If the number of votes for the consensus label is less than the set value of the total number of clients, then discard the corresponding data to obtain an optimized synthetic dataset.

[0032] In the present invention, to solve the problem of low quality of synthetic data, client knowledge is used to correct the labels of synthetic data. For each synthetic data, count the number of votes of the client for each category, and use majority voting to determine the consensus label: where is an indicator function. To ensure the quality and reliability of synthetic data, through a screening mechanism. If the consensus label has fewer votes than of the total number of clients, then the i th data is considered unreliable and discarded.

[0033] where represents the number of votes obtained. Based on experimental evaluation, is set in this project. This threshold can ensure label accuracy to a high degree while retaining sufficient data for subsequent knowledge fusion. The corrected high-quality data is then input into the next module for multimodal representation fusion.

[0034] Based on the optimized synthetic dataset, construct a unified representation through multimodal representation fusion to complete cross-client semantic alignment, specifically including: Extract image representations through the client based on the optimized synthetic dataset; Perform multimodal representation alignment through an inter-modal cross-attention mechanism based on the image representations; Based on the image representations after representation alignment, construct a unified representation through contrast aggregation to complete cross-client semantic alignment.

[0035] In the present invention, in order to integrate the representations and modalities from different clients to achieve knowledge transfer between the client and the server. Since each client processes different data modalities and tasks, the uploaded representations will inevitably have modality biases and distribution biases. The main reason for the modality bias is that the client does not have access to data information of another modality during training. This modality separation will lead to differences in the representations generated for the same synthetic data subsequently. The distribution bias stems from the non-independent and identically distributed nature of the client data, which results in disagreements in the representations of the same synthetic data. To address the representation differences from different clients, the present invention designs two mechanisms.

[0036] The cross-modal cross-attention mechanism. To address the modality bias existing in single-modal clients, first, the cross-modal cross-attention mechanism is used for multi-modal representation alignment. Specifically, the cross-modal cross-attention uses the complementary information of another modality to integrate the local single-modal representations: Among them, represents the cross-attention mechanism. Aggregating all clients of the same modality can reduce the influence of client-specific differences and emphasize the common knowledge of all clients, thus providing a more balanced and consistent perspective for another modality. This step reduces the bias of specific patterns and achieves more coherent multi-modal knowledge transfer.

[0037] Based on the image representations after representation alignment, a unified representation is constructed through contrastive aggregation to complete cross-client semantic alignment, specifically including: Comparing and weighting the image representations after representation alignment with the average representation; Calculating the representation weight value of the client according to the set formula, and performing aggregation and normalization operations on the representations with high weight values according to the representation weight value to construct a unified representation and complete cross-client semantic alignment.

[0038] In the present invention, through the cross-client contrastive aggregation mechanism, after representation alignment, the multi-modal representations from all clients are fused into a unified representation for subsequent knowledge transfer. However, due to the distribution bias between clients, it is necessary to accurately evaluate the quality and contribution of each representation. A cross-client and cross-modal contrast score is designed for weighting. Specifically, a representation that is "closer" to the average representation of the paired samples and "farther" from other irrelevant samples is considered to be able to better capture the potential cross-modal semantic information. Therefore, the i weight of the n representation of the th client for the Among them, Label cosine similarity. Next, the server performs normalization and aggregation operations according to the following formula: Step 300: Based on the unified representation, fine-tune and train the large model on the server side and the small model on the client side respectively to complete collaborative optimization.

[0039] In the present invention, the fused unified representation is combined into to improve the performance of the server and client models.

[0040] Fine-tune the text-to-image model on the server side. The original T2I model consists of an encoder and a decoder . The encoder projects the text condition into the semantic latent space, and the decoder uses the semantic latent space to generate coherent images. In this project, the aggregated global representation is integrated with the text embedding of the encoder, expressed as: where the label is converted into the label description . The decoder takes z as input and generates a picture that reflects the expected meaning of the text prompt and the fused representation.

[0041] The main training objective here is to optimize the decoder , while keeping the encoder and other components frozen. This can ensure that the generated pictures are consistent with the conditions set by the provided multimodal input, thereby improving the context accuracy of the generated pictures.

[0042] Finally, since the fused representation from the client is not available during the inference phase, the model needs to remain reliable in the absence of the fused representation. To solve this problem, a random component is introduced during the fine-tuning phase. Specifically, a Bernoulli random variable B ( p ) is adopted, where p represents the probability of omitting the image representation, as shown below: Set p = 0.2. This random method ensures that the model can maintain its generation ability even when the fused representation is not available. After training is completed, a fine-tuned decoder will be obtained. During the inference process, the fine-tuned T2I model only uses the text prompt input.

[0043] When fine-tuning the large model for the client, due to non-independent and identically distributed data and missing modalities, the client model trained with private data will also show biases. To address these biases, the corrected synthetic data is combined with the global representation to retrain the local model to obtain an enhanced model. Specifically, the local model is encouraged to align its representation with the global representation using the KL loss and align its local predicted labels with the consensus labels using the cross-entropy loss.

[0044] Based on a collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention, through multi-modal knowledge transfer, the text-to-image model on the server side and the large model on the client side are collaboratively enhanced, and there is no need to share customer data or model parameters, ensuring privacy; using synthetic data as an intermediary for knowledge transfer, combined with label voting to effectively improve data quality, and integrating multi-modal representations from heterogeneous clients to promote two-way knowledge transfer between the text-to-image model on the server side and the client model.

[0045] Reference Figure 2 and Figure 3 According to, the present invention also discloses a collaborative optimization system for large and small models in federated learning based on cross-modal knowledge fusion, and the system includes: A cross-modal data synthesis module 110, configured to generate an image and its corresponding text description through a text-to-image model preset on the server side and an image-to-text model obtained by calling an API interface, and construct a synthetic data set through the image and the text description; A cross-client semantic alignment module 120, configured to optimize the data through label voting based on the synthetic data set, and construct a unified representation through multi-modal representation fusion based on the optimized synthetic data set to complete cross-client semantic alignment; A knowledge transfer module 130, configured to perform fine-tuning training on the large model on the server side and the small model on the client side respectively based on the unified representation to complete collaborative optimization.

[0046] Among them, generating an image and its corresponding text description through a text-to-image model preset on the server side and an image-to-text model obtained by calling an API interface, and constructing a synthetic data set through the image and the text description specifically includes: Randomly select a class label from the existing historical database on the server side, and convert the class label into a label description; Input the label description into the text-to-image model preset on the server side to generate an image; Input the image into the image - to - text model obtained by calling the API interface to generate the corresponding text description; Construct a synthetic dataset with the image and the text description and send it to the client.

[0047] Optimize the data through label voting based on the synthetic dataset, specifically including: Extract local prediction labels on the client based on the synthetic dataset; Based on the local prediction labels, count the number of votes for each category in each synthetic data in the synthetic dataset, and determine the consensus label through majority voting; If the number of votes for the consensus label is less than the set value of the total number of clients, discard the corresponding data to obtain the optimized synthetic dataset.

[0048] Construct a unified representation through multi - modal representation fusion based on the optimized synthetic dataset to complete cross - client semantic alignment, specifically including: Extract image representations on the client based on the optimized synthetic dataset; Perform multi - modal representation alignment through the cross - attention mechanism between modalities based on the image representations; Based on the image representations after representation alignment, construct a unified representation through contrastive aggregation to complete cross - client semantic alignment.

[0049] Among them, based on the image representations after representation alignment, construct a unified representation through contrastive aggregation to complete cross - client semantic alignment, specifically including: Compare and weight the image representations after representation alignment with the average representation; Calculate the representation weight value of the client according to the set formula, perform aggregation and normalization operations on the representations with high weight values according to the representation weight value, construct a unified representation, and complete cross - client semantic alignment.

[0050] Fine - tune and train the large model on the server and the small model on the client respectively based on the unified representation to complete collaborative optimization, specifically including: Optimize the decoder of the text - to - image model on the server based on the unified representation, and introduce Bernoulli random variables for fine - tuning the decoder to complete the fine - tuning training of the large model on the server; Combine the unified representation with the global representation of the client to retrain the local model, align the unified representation with the global representation using the federated learning loss, and align the local prediction labels with the consensus labels using the cross - entropy loss to complete the fine - tuning training of the small model on the client.

[0051] A collaborative optimization system for large and small models in federated learning based on cross-modal knowledge fusion provided by the present invention collaboratively enhances the text-to-image model on the server side and the large model on the client side through multi-modal knowledge transfer, and without sharing customer data or model parameters, ensuring privacy; uses synthetic data as an intermediary for knowledge transfer, combines label voting to effectively improve data quality, and integrates multi-modal representations from heterogeneous clients to promote two-way knowledge transfer between the text-to-image model on the server side and the client model.

[0052] Figure 4 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute a collaborative optimization method for large and small models in federated learning based on cross-modal knowledge fusion. The method includes: generating an image and its corresponding text description through a text-to-image model preset on the server side and an image-to-text model obtained by calling an API interface, and constructing a synthetic data set through the image and the text description; performing data optimization through label voting based on the synthetic data set, and constructing a unified representation through multi-modal representation fusion based on the optimized synthetic data set to complete cross-client semantic alignment; performing fine-tuning training on the large model on the server side and the small model on the client side respectively based on the unified representation to complete collaborative optimization.

[0053] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0054] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a collaborative optimization method for large and small models of federated learning based on cross-modal knowledge fusion provided by the above-mentioned various methods. The method includes: generating an image and its corresponding text description through a text-to-image model preset by a server and an image-to-text model obtained by calling an API interface, and constructing a synthetic dataset through the image and the text description; performing data optimization through label voting based on the synthetic dataset, constructing a unified representation through multi-modal representation fusion based on the optimized synthetic dataset, and completing cross-client semantic alignment; respectively performing fine-tuning training on the large model of the server and the small model of the client based on the unified representation to complete collaborative optimization.

[0055] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes a collaborative optimization method for large and small models of federated learning based on cross-modal knowledge fusion provided by the above-mentioned various methods. The method includes: generating an image and its corresponding text description through a text-to-image model preset by a server and an image-to-text model obtained by calling an API interface, and constructing a synthetic dataset through the image and the text description; performing data optimization through label voting based on the synthetic dataset, constructing a unified representation through multi-modal representation fusion based on the optimized synthetic dataset, and completing cross-client semantic alignment; respectively performing fine-tuning training on the large model of the server and the small model of the client based on the unified representation to complete collaborative optimization.

[0056] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0057] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for collaborative optimization of large and small models of federated learning based on cross-modal knowledge fusion, characterized in that: include: Generate images and their corresponding text descriptions through the text-to-image model preset on the server and the image-to-text model obtained by calling the API interface, and build a synthetic dataset through images and text descriptions; Based on the synthetic data set, data optimization is performed through label voting, and based on the optimized synthetic data set, a unified representation is constructed through multimodal representation fusion to complete cross-client semantic alignment; Based on the unified representation, the large model on the server and the small model on the client are fine-tuned and trained respectively to complete collaborative optimization.

2. The method for collaborative optimization of large and small models based on cross-modal knowledge fusion in federated learning according to claim 1 is characterized in that: The method generates an image and its corresponding text description by using a text-to-image model preset by the server and an image-to-text model obtained by calling an API interface, and constructs a synthetic data set by using the image and text description, specifically including: Randomly select a class label based on the existing historical database on the server, and convert the class label into a label description; Input the label description into a text-to-image model preset by the server to generate an image; Input the image into the image generation text model obtained by calling the API interface to generate a corresponding text description; A synthetic dataset is constructed by using the images and text descriptions and sent to the client.

3. The method for collaborative optimization of large and small models based on cross-modal knowledge fusion in federated learning according to claim 1, characterized in that: The data optimization based on the synthetic data set through label voting specifically includes: Extracting local prediction labels on the client based on the synthetic dataset; Based on the local predicted label, counting the number of votes for each category for each synthetic data in the synthetic data set, and determining the consensus label by majority voting; When the number of votes for the consensus tag is less than the set value of the total number of clients, the corresponding data is discarded to obtain an optimized synthetic data set.

4. The method for collaborative optimization of large and small models based on cross-modal knowledge fusion in federated learning according to claim 1, characterized in that: The method of constructing a unified representation based on the optimized synthetic dataset by fusion of multimodal representations to complete cross-client semantic alignment specifically includes: Extract image representations through the client based on the optimized synthetic dataset; Multimodal representation alignment based on image representation through inter-modal cross-attention mechanism; Based on the aligned image representations, a unified representation is constructed through comparison and aggregation to complete cross-client semantic alignment.

5. The method for collaborative optimization of large and small models based on cross-modal knowledge fusion in federated learning according to claim 4 is characterized in that: The image representations after representation alignment are compared and aggregated to construct a unified representation, thereby completing cross-client semantic alignment, specifically including: Compare and weight the image representation after representation alignment with the average representation; The client's representation weight value is calculated according to the set formula, and the representations with high weight values ​​are aggregated and normalized according to the representation weight value to build a unified representation and complete cross-client semantic alignment.

6. The method for collaborative optimization of large and small models based on cross-modal knowledge fusion in federated learning according to claim 1, characterized in that: The fine-tuning training of the large model on the server and the small model on the client are respectively performed based on the unified representation to complete the collaborative optimization, specifically including: Based on the unified representation, the decoder of the text-to-image model on the server is optimized, and Bernoulli random variables are introduced to fine-tune the decoder to complete the fine-tuning training of the large model on the server; The local model is retrained based on the combination of the unified representation and the global representation of the client, the unified representation is aligned with the global representation using the federated learning loss, and the local prediction label is aligned with the consensus label using the cross entropy loss to complete the fine-tuning training of the small client model.

7. A federated learning large and small model collaborative optimization system based on cross-modal knowledge fusion, characterized in that: The system comprises: The cross-modal data synthesis module is used to generate images and their corresponding text descriptions through the text-to-image model preset on the server and the image-to-text model obtained by calling the API interface, and to construct a synthetic dataset through images and text descriptions; A cross-client semantic alignment module, used to perform data optimization through label voting based on the synthetic data set, and to construct a unified representation through multimodal representation fusion based on the optimized synthetic data set to complete cross-client semantic alignment; The knowledge transfer module is used to fine-tune the large model on the server and the small model on the client based on the unified representation to complete collaborative optimization.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for collaborative optimization of large and small models of federated learning based on cross-modal knowledge fusion as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for collaborative optimization of large and small models of federated learning based on cross-modal knowledge fusion as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for collaborative optimization of large and small models of federated learning based on cross-modal knowledge fusion as described in any one of claims 1 to 6 is implemented.