Cross-modal retrieval model training method, unsupervised federated learning system and program product
By performing unsupervised training locally and generating local enhanced data, the problems of data privacy and model offset in cross-modal retrieval model training are solved, and the effect of efficiently training of global cross-modal retrieval models in an unsupervised federated learning environment is achieved.
Patent Information
- Application Number
- CN202510179134.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art faces the problems of data privacy protection and model offset in cross-modal retrieval model training, especially in an unsupervised federated learning environment, which is difficult to fully utilize the relationship of local multimodal data and rely on high-quality annotations.
By performing unsupervised training locally, using cross-modal splicing features and local prototypes as constraints, iterative training of the adversarial generation network is generated to generate local augmented data, and sending local prototypes, generators and local cross-modal search models to the server for global training to form a global cross-modal search model.
On the premise of ensuring local data privacy, make full use of the multimodal relationship of local data, reduce dependence on high-quality annotations, improve the training efficiency and reliability of global cross-modal retrieval models, and effectively solve the model offset problem.
Smart Images

Figure CN120234609A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of federated learning technology, and in particular to a cross-modal retrieval model training method, an unsupervised federated learning system, and a program product. Background Art
[0002] In modern society, the booming multi-modal data, such as images, texts, and videos on the Internet, makes a robust cross-modal retrieval model a key requirement. Among various advanced technologies, the deep cross-modal hashing method that utilizes deep neural networks by using deep hashing technology has significantly improved the accuracy and efficiency of large-scale cross-modal retrieval. However, training a deep cross-modal hashing model still faces the problem of data privacy. Real-world data is usually discretely distributed among multiple clients, which can only be accessed by the clients, and the clients include mobile phones, personal computers, and Internet of Things devices. To meet the requirements of privacy protection, federated learning is introduced into cross-modal retrieval. This allows the server to obtain a relatively good performance model without the need for the original client data. However, it is still challenged by the global model drift problem caused by non-independent and identically distributed data: the quantity and category of data vary from client to client.
[0003] Currently, the joint cross-modal hashing method with adaptive feature enhancement, FedCAFE, is a supervised cross-modal retrieval method that introduces a novel adaptive feature enhancement module and a new weighted aggregation strategy. It makes full use of the rich global information carried in the global model to constrain the model during the local training process. The federated unsupervised cross-modal hashing method based on prototype transfer, PT-FUCH, protects local private data by exploring a unified global prototype for different clients without relying on any annotations. The global prototype is used to guide local cross-modal hashing learning and promote the alignment of the feature space, thereby alleviating the model bias caused by the differences in local multi-modal data distributions. In addition, it also includes an adaptive cross-modal knowledge distillation method that transfers valuable semantic knowledge from the global model of a specific modality to the local prototype learning process, reducing the risk of overfitting.
[0004] However, the FedCAFE method highly depends on high-quality annotations of data. It is difficult to have a sufficient amount of multi-modal data with high-quality annotations in real-world scenarios, resulting in low feasibility of this method in reality. And the guidance of the global prototype in PT-FUCH can only alleviate the model drift problem caused by non-independent and identically distributed data to a certain extent, and the multi-modal relationships of the real data in each client are not fully utilized. Summary of the Invention
[0005] In view of this, the embodiments of this application provide a cross-modal retrieval model training method, an unsupervised federated learning system, and a program product to eliminate or improve one or more defects existing in the prior art.
[0006] One aspect of the present application provides a method for training a cross-modal retrieval model, including:
[0007] Using the cross-modal splicing features corresponding to the local multi-modal data and the local prototypes obtained by clustering the cross-modal splicing features in advance as constraints, performing unsupervised training on the local adversarial generation network to train the adversarial generation network into a local generator for generating local enhanced data;
[0008] Sending the local local prototypes, the local generator, and the local cross-modal retrieval model pre-trained locally based on the feature vectors of various modalities corresponding to the local multi-modal data to the server, so that the server obtains global training data according to the received local prototypes and local generators, and enables the server to train the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0009] In some embodiments of the present application, before performing unsupervised training on the local adversarial generation network using the cross-modal splicing features corresponding to the local multi-modal data and the local prototypes obtained by clustering the cross-modal splicing features in advance, it further includes:
[0010] Dividing the local multi-modal data into respective training data pairs, where each training data pair contains first-modal data and second-modal data belonging to different modal types respectively;
[0011] Inputting the first-modal data in each training data pair into the feature extraction model corresponding to the first modality, and inputting the second-modal data in each training data pair into the feature extraction model corresponding to the second modality, so that the feature extraction model corresponding to the first modality outputs the first-modal feature vector corresponding to the first-modal data, and the feature extraction model corresponding to the second modality outputs the second-modal feature vector corresponding to the second-modal data;
[0012] Training the local cross-modal retrieval model using the first-modal feature vectors and the second-modal data corresponding to each training data pair and the hash loss, so that the local cross-modal retrieval model is used to respectively obtain the hash representations corresponding to the first-modal data and the second-modal data;
[0013] And respectively projecting and splicing the first-modal feature vectors and the second-modal data corresponding to each training data pair to respectively obtain the cross-modal splicing features corresponding to each training data pair;
[0014] Cluster and project the cross-modal splicing features corresponding to each of the training data pairs to obtain local prototypes corresponding to the local multi-modal data.
[0015] In some embodiments of the present application, using the cross-modal splicing features corresponding to the local multi-modal data and the local prototypes obtained by clustering the cross-modal splicing features in advance as constraints, performing unsupervised training on the local adversarial generation network to train the adversarial generation network into a local generator for generating local enhanced data, includes:
[0016] Performing multiple rounds of iterative training on the local adversarial generation network using the cross-modal splicing features corresponding to the local multi-modal data, and respectively executing preset training steps in each round;
[0017] Wherein, the training steps include:
[0018] If it is the first round of iteration currently, input the cross-modal splicing features corresponding to the local multi-modal data into the local adversarial generation network, so that the adversarial generation network outputs the local enhanced data of the current round; if it is not the first round of iteration currently, adjust the local adversarial generation network based on the training loss obtained in the previous round, and input the cross-modal splicing features corresponding to the local multi-modal data into the adjusted adversarial generation network, so that the adversarial generation network outputs the local enhanced data of the current round;
[0019] Assign pseudo-labels to the local enhanced data according to the local prototypes;
[0020] And, divide the local enhanced data into first-modal enhanced features and second-modal enhanced features according to the modal type respectively;
[0021] Based on the first-modal enhanced features and second-modal enhanced features corresponding to the local enhanced data, and the local prototypes having the same pseudo-labels as the local enhanced data, obtain the training loss with data quality constraints corresponding to the adversarial generation network in the next round.
[0022] In some embodiments of the present application, the training loss with data quality constraints is solved based on the training loss function of the adversarial generation network;
[0023] Wherein, the training loss function of the adversarial generation network includes: a generator training loss function and a discriminator training loss function;
[0024] The generator training loss function is solved based on the client loss and a preset generator loss; the discriminator training loss function is solved based on the client loss and a preset discriminator loss;
[0025] Among them, the client loss is solved based on the intra-class and inter-class contrast alignment loss, the distribution alignment loss, and the hash loss for training the local cross-modal retrieval model;
[0026] The intra-class and inter-class contrast alignment loss and the distribution alignment loss are respectively solved based on the local prototype with the same pseudo-label as the local augmented data, the first-modal augmented feature and the second-modal augmented feature corresponding to the local augmented data.
[0027] In some embodiments of the present application, the intra-class and inter-class contrast alignment loss is solved based on the indirect contrast loss and the direct contrast loss;
[0028] The indirect contrast loss is solved based on the indirect contrast losses corresponding to the first modality and the second modality respectively belonging to different modality types;
[0029] Among them, the indirect contrast loss corresponding to the first modality is solved based on the first-modal augmented feature, the local prototype with the same pseudo-label as the local augmented data to which the first-modal augmented feature belongs, and a preset first temperature coefficient;
[0030] The indirect contrast loss corresponding to the second modality is solved based on the second-modal augmented feature, the local prototype with the same pseudo-label as the local augmented data to which the second-modal augmented feature belongs, and the first temperature coefficient;
[0031] The direct contrast loss is solved based on a preset second temperature coefficient, the first-modal augmented feature and the second-modal augmented feature corresponding to the local augmented data.
[0032] In some embodiments of the present application, the distribution alignment loss is solved based on the loss of minimizing the difference between the pseudo-label distributions of the first-modal augmented feature and the second-modal augmented feature and a preset weight hyperparameter;
[0033] Among them, the loss of minimizing the difference between the pseudo-label distributions of the first-modal augmented feature and the second-modal augmented feature is solved based on the pseudo-label distribution of the first-modal augmented feature, the pseudo-label distribution of the second-modal augmented feature, and an average parameter; the value of the average parameter is the sum average of the pseudo-label distribution of the first-modal augmented feature and the pseudo-label distribution of the second-modal augmented feature;
[0034] The pseudo-label distribution of the first-modal augmented feature is solved based on the first-modal augmented feature and the local prototype with the same pseudo-label as the local augmented data to which the first-modal augmented feature belongs;
[0035] The pseudo-label distribution of the second-modal enhancement feature is solved based on the second-modal enhancement feature and the local prototype that has the same pseudo-label as the local enhancement data to which the second-modal enhancement feature belongs.
[0036] The second aspect of the present application provides a cross-modal retrieval model training method, including:
[0037] Receiving local data from each client device respectively in the current federated learning round, where each piece of the local data contains: the corresponding local cross-modal retrieval model, the local prototype, and the local generator; and the local cross-modal retrieval model, the local prototype, and the local generator are pre-obtained by the client device based on the cross-modal retrieval model training method provided in the foregoing first aspect;
[0038] Obtaining global training data according to the local prototypes and the local generators corresponding to the current federated learning round, and training the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0039] In some embodiments of the present application, the step of obtaining global training data according to the local prototypes and the local generators corresponding to the current federated learning round, and training the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data includes:
[0040] Clustering each local cross-modal retrieval model corresponding to the current federated learning round to obtain the corresponding global cross-modal retrieval model;
[0041] Clustering each local prototype corresponding to the current federated learning round to obtain the corresponding global prototype;
[0042] And, based on the proportion of the local multi-modal data corresponding to each source, generating the corresponding local enhancement data respectively through each global model and jointly constituting the corresponding global enhancement data;
[0043] Performing discrete value filtering on the global enhancement data based on the global prototype to obtain the corresponding global training data, and training the global cross-modal retrieval model using the global training data.
[0044] The third aspect of the present application provides a cross-modal retrieval model training device based on unsupervised federated learning, including:
[0045] An unsupervised training module, which is used to perform unsupervised training on a local adversarial generation network by using the cross-modal concatenated features corresponding to local multi-modal data and local prototypes obtained by clustering the cross-modal concatenated features in advance, so as to train the adversarial generation network into a local generator for generating local enhanced data;
[0046] A local data sending module, which is used to send the local local prototypes, the local generator, and a local cross-modal retrieval model pre-trained locally based on the feature vectors of various modalities corresponding to local multi-modal data to a server, so that the server can obtain global training data according to the received local prototypes and local generators, and enable the server to train a global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0047] The fourth aspect of the present application provides a cross-modal retrieval model training device based on unsupervised federated learning, including:
[0048] A local data receiving module, which is used to receive local data from each client device respectively in the current federated learning round. Each piece of local data includes: the corresponding local cross-modal retrieval model, the local prototype, and the local generator; and the local cross-modal retrieval model, the local prototype, and the local generator are pre-obtained by the client device based on the cross-modal retrieval model training method provided in the first aspect above;
[0049] A global model training module, which is used to obtain global training data according to the local prototypes and local generators corresponding to the current federated learning round, and train a global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0050] The fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the cross-modal retrieval model training method provided in the first aspect above, or implements the cross-modal retrieval model training method provided in the second aspect above.
[0051] The sixth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the cross-modal retrieval model training method provided in the first aspect above, or implements the cross-modal retrieval model training method provided in the second aspect above.
[0052] The seventh aspect of the present application provides an unsupervised federated learning system, including: a server, and each client device communicatively connected to the server respectively;
[0053] The client device includes a first memory, a first processor, and a first computer program stored on the first memory and executable on the processor. When the first processor executes the first computer program, it implements the cross-modal retrieval model training method provided in the foregoing first aspect;
[0054] The server includes a second memory, a second processor, and a second computer program stored on the second memory and executable on the processor. When the second processor executes the second computer program, it implements the cross-modal retrieval model training method provided in the foregoing first aspect.
[0055] The eighth aspect of the present application provides a computer program product, including a computer program, which when executed by a processor, implements the cross-modal retrieval model training method provided in the foregoing first aspect, or implements the cross-modal retrieval model training method provided in the foregoing second aspect.
[0056] The cross-modal retrieval model training method provided by the present application uses the cross-modal splicing features corresponding to local multi-modal data and the local prototypes obtained by clustering the cross-modal splicing features in advance as constraints to perform unsupervised training on the local adversarial generation network to train the adversarial generation network into a local generator for generating local enhanced data; sending the local prototypes, the local generator, and the local cross-modal retrieval model obtained by local training based on the feature vectors of various modalities corresponding to the local multi-modal data to the server, so that the server can obtain global training data according to the received local prototypes and local generators, and enabling the server to train the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data. During the process of training the global cross-modal retrieval model based on networked learning, it can realize unsupervised learning of the local generator based on multi-modal data, make full use of the multi-modal relationship of local data, and does not rely on high-quality data annotation. Furthermore, it can improve the efficiency and reliability of training the global cross-modal retrieval model while ensuring the privacy of local data, and can effectively solve the offset problem of the global cross-modal retrieval model.
[0057] The additional advantages, objectives, and features of the present application will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification and the drawings.
[0058] Those skilled in the art will understand that the objectives and advantages achievable with the present application are not limited to those specifically described above, and the above and other objectives achievable with the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:
[0060] Figure 1 It is a first process schematic diagram of the first cross-modal retrieval model training method in an embodiment of the present application.
[0061] Figure 2 It is a second process schematic diagram of the first cross-modal retrieval model training method in an embodiment of the present application.
[0062] Figure 3 It is an execution process schematic diagram of the training steps in the first cross-modal retrieval model training method in an embodiment of the present application.
[0063] Figure 4 It is a first process schematic diagram of the second cross-modal retrieval model training method in an embodiment of the present application.
[0064] Figure 5 It is a second process schematic diagram of the second cross-modal retrieval model training method in an embodiment of the present application.
[0065] Figure 6 It is a structural schematic diagram of the first cross-modal retrieval model training device in an embodiment of the present application.
[0066] Figure 7 It is a structural schematic diagram of the second cross-modal retrieval model training device in an embodiment of the present application.
[0067] Figure 8 It is an execution logic schematic diagram of the FedWGAN method for unsupervised cross-modal retrieval in an application example of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments of the present application and their descriptions are used to explain the present application, but do not limit the present application.
[0069] Here, it should also be noted that, in order to avoid obscuring the present application with unnecessary details, only the structures and / or processing steps closely related to the solution according to the present application are shown in the drawings, while other details less relevant to the present application are omitted.
[0070] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0071] Here, it should also be noted that, unless otherwise specified, the term "connection" in this document can refer not only to direct connection, but also to indirect connection with an intermediate.
[0072] Hereinafter, embodiments of the present application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0073] In order to improve the efficiency and reliability of global cross-modal retrieval model training while ensuring local data privacy, embodiments of the present application respectively provide a first cross-modal retrieval model training method, a first cross-modal retrieval model training device for executing the first cross-modal retrieval model training method, a second cross-modal retrieval model training method, a second cross-modal retrieval model training device for executing the second cross-modal retrieval model training method, an electronic device, a computer-readable storage medium, an unsupervised federated learning system, and a computer program product, which can make full use of the multi-modal relationships of local data and do not rely on high-quality data annotation during the process of global cross-modal retrieval model training based on networked learning to improve the efficiency and reliability of cross-modal retrieval model training while ensuring local data privacy.
[0074] Specifically, it will be described in detail through the following embodiments.
[0075] Based on this, an embodiment of the present application provides a first cross-modal retrieval model training method that can be implemented by a first cross-modal retrieval model training device, see Figure 1 , and the first cross-modal retrieval model training method specifically includes the following content:
[0076] Step 100: Constrained by the cross-modal splicing features corresponding to the local multi-modal data and the local prototypes obtained by clustering the cross-modal splicing features in advance, perform unsupervised training on the local adversarial generation network to train the adversarial generation network into a local generator for generating local enhanced data.
[0077] In one or more embodiments of the present application, local multimodal data refers to multimodal data local to the device where the first cross-modal retrieval model training device is located. This multimodal data can at least include data of two different modal types, namely first modal data and second modal data. Among them, one of the first modal data and the second modal data can be image data, and the other can be text data. Currently, since cross-modal refers to the interaction, understanding, and processing of information between different modalities. A modality refers to the form of perception or data, such as text, image, audio, video, etc. The research on cross-modal aims to achieve collaboration and interoperability between different modalities, such as cross-modal retrieval. Therefore, the first modal data and the second modal data can also include audio and / or video data, etc., specifically selected according to actual application requirements.
[0078] Among them, taking image data and text data as an example, in one or more embodiments of the present application, the first modality can refer to an image, and the second modality can refer to text. Correspondingly, the first modal data can be image data; the second modal data can be text data.
[0079] Correspondingly, the cross-modal splicing feature corresponding to the local multimodal data refers to the cross-modal splicing feature obtained according to the local multimodal data. Specifically, a first modal feature vector can be extracted from the first modal data in the multimodal data and a second modal feature vector can be extracted from the second modal data in the multimodal data, and then the cross-modal splicing feature is generated according to the first modal feature vector and the second modal feature vector. In one example, the first modal feature vector can be an image feature vector (which can also be simply referred to as an image feature); the second modal feature vector can be a text feature vector (which can also be simply referred to as a text feature).
[0080] It can be understood that the local prototype obtained by clustering the cross-modal splicing feature refers to: performing clustering processing on each of the cross-modal splicing features in a manner such as K-mans to obtain each clustering center, and then forming the clustering result data corresponding to each clustering center into a corresponding data set as the local prototype of the first cross-modal retrieval model training device. From the perspective of the first cross-modal retrieval model training device, the local prototype can be referred to as the local prototype.
[0081] In one or more embodiments of the present application, the adversarial generation network can adopt the WGAN (Wasserstein GAN) model. The WGAN model is an improved generative adversarial network GAN that solves the problems of the original GAN in training stability, loss function indicativeness, and sample diversity.
[0082] It should be noted that federated learning is a machine learning technology with privacy protection and distributed collaboration, aiming to train a model without the data leaving the local area. Its core idea is to distribute the model training tasks to multiple devices or nodes, and after local computing, send the model updates to the central server without the need to centralize the original data. Data generation is the process of synthesizing new data through algorithms or models, which is usually used to simulate actual data or generate data for training and testing machine learning models.
[0083] In addition, the local generator refers to the data generator local to the first cross-modal retrieval model training device, and can also be called the local generator or the WGAN-generator. In the federated learning scenario, the first cross-modal retrieval model training device can be set in the client device, and this client device is communicatively connected to the server in its federated learning system, and then sends the local data generated in this round of federated learning rounds through step 100 to the server through the following step 200.
[0084] Step 200: Send the local local prototype, the local generator, and the local cross-modal retrieval model locally trained based on the feature vectors of various modalities corresponding to the local multi-modal data to the server, so that the server can obtain global training data according to the received local prototypes and local generators, and enable the server to train the global cross-modal retrieval models corresponding to the local cross-modal retrieval models based on the global training data.
[0085] It can be understood that from the perspective of the first cross-modal retrieval model training device, the local cross-modal retrieval model can also be called the local cross-modal retrieval model. Specifically, an unsupervised contrastive cross-modal hashing (UCCH) model can be adopted. The UCCH model realizes learnable hashing operations through a momentum optimizer and proposes a cross-modal ranking learning loss (CRL) to utilize the discrimination of all negative pairs, thereby improving the retrieval performance. This method can better perform cross-modal hashing learning on unlabeled data and reduces the performance gap between contrastive learning and hashing.
[0086] In step 200, the server is used to perform: receiving local data from different sources in the current federated learning round, where each piece of the local data contains: the corresponding local cross-modal retrieval model, the local prototype, and the local generator; clustering the local cross-modal retrieval models corresponding to the current federated learning round to obtain the corresponding global cross-modal retrieval model; clustering the local prototypes corresponding to the current federated learning round to obtain the corresponding global prototype; and, based on the proportion of the local multi-modal data corresponding to each source, generating the corresponding local enhanced data through each of the global models and jointly constituting the corresponding global enhanced data; filtering the global enhanced data based on the global prototype to obtain the corresponding global training data, and training the global cross-modal retrieval model using the global training data.
[0087] As can be seen from the above description, the first cross-modal retrieval model training method provided by the embodiments of the present application can, in the process of training the global cross-modal retrieval model based on networked learning, achieve unsupervised learning of the local generator based on multi-modal data, can make full use of the multi-modal relationship of local data, and does not rely on high-quality data annotation. Furthermore, it can improve the efficiency and reliability of training the global cross-modal retrieval model while ensuring the privacy of local data, and can effectively solve the deviation problem of the global cross-modal retrieval model.
[0088] In order to further improve the effectiveness and applicability of cross-modal retrieval model training, in a first cross-modal retrieval model training method provided by the embodiments of the present application, see Figure 2 , before step 100 in the first cross-modal retrieval model training method, the following specific content is further included:
[0089] Step 010: Divide the local multi-modal data into each training data pair, where each training data pair contains first-modal data and second-modal data that belong to different modal types respectively.
[0090] Step 020: Input the first-modal data in each training data pair into the feature extraction model corresponding to the first modality, and input the second-modal data in each training data pair into the feature extraction model corresponding to the second modality, so that the feature extraction model corresponding to the first modality outputs the first-modal feature vector corresponding to the first-modal data, and the feature extraction model corresponding to the second modality outputs the second-modal feature vector corresponding to the second-modal data.
[0091] In an example, the first-modal feature vector is the image feature vector F I; The feature extraction model corresponding to the first modality can be selected as a convolutional neural network CNN or a CNN-F model. Among them, CNN-F refers to a fast version of the convolutional neural network architecture, that is, CNN-F (Fast), which comes from the paper "Return of the Devil in the Details: Delving Deep into Convolutional Nets. In BMVC, Chatfield, K.; Simonyan, K.; Vedaldi, A.; and Zisserman, A. 2014). The second modality feature vector is the text feature vector F T ; The feature extraction model corresponding to the second modality can be selected as the BERT (Bidirectional Encoder Representations from Transformers) model. The BERT model is a pre-trained language model based on the Transformer architecture. Through bidirectional encoder representations, it can utilize the left and right context information of the input sequence simultaneously, so as to understand the meaning of language more accurately.
[0092] Step 030: Use each of the training data to train the local cross-modal retrieval model for the corresponding first modality feature vector, the second modality data, and the hash loss, so that the local cross-modal retrieval model is used to obtain the hash representations corresponding to the first modality data and the second modality data respectively.
[0093] Among them, for the k-th client, its local cross-modal retrieval model can be written as UCCH k .
[0094] And, Step 040: Project and splice the first modality feature vector and the second modality data corresponding to each of the training data respectively, to obtain the cross-modal splicing features corresponding to each of the training data respectively.
[0095] Among them, a multi-layer perceptron MLP (Multi-Layer Perceptron) can be used to project the first modality feature vector and the second modality data corresponding to each of the training data. In an example, the image feature vector F I is input into the first multi-layer perceptron MLP I , and the image feature vector F T is input into the second multi-layer perceptron MLP T . The cross-modal splicing feature is F, and the following local prototype is C. For the k-th client, its local prototype can also be written as C k .
[0096] Step 050: Cluster and project the cross-modal concatenated features corresponding to each of the training data pairs to obtain local prototypes corresponding to the local multi-modal data.
[0097] To further improve the effectiveness and reliability of unsupervised training for the local local generator, in a first cross-modal retrieval model training method provided in an embodiment of the present application, refer to Figure 2 The step 100 in the first cross-modal retrieval model training method specifically includes the following content:
[0098] Step 110: Iteratively train the local adversarial generation network for multiple rounds using the cross-modal concatenated features corresponding to the local multi-modal data, and perform preset training steps in each round.
[0099] Among them, refer to Figure 3 The training steps in the step 110 specifically include the following content:
[0100] Step 111: If it is the first round of iteration, input the cross-modal concatenated features corresponding to the local multi-modal data into the local adversarial generation network so that the adversarial generation network outputs the local enhanced data for the current round; if it is not the first round of iteration, adjust the local adversarial generation network based on the training loss obtained in the previous round, and input the cross-modal concatenated features corresponding to the local multi-modal data into the adjusted adversarial generation network so that the adversarial generation network outputs the local enhanced data for the current round.
[0101] Among them, the local enhanced data (or local enhanced data) can be denoted as F g .
[0102] Step 112: Assign pseudo-labels to the local enhanced data according to the local prototypes.
[0103] And, step 113: Divide the local enhanced data into first-modal enhanced features and second-modal enhanced features according to the modal type.
[0104] Among them, in an example, the first-modal enhanced feature is an image enhanced feature The second-modal enhanced feature is a text enhanced feature
[0105] Step 114: Based on the first-modal enhanced feature and the second-modal enhanced feature corresponding to the local enhanced data, and the local prototype having the same pseudo-label as the local enhanced data, obtain the training loss with data quality constraints corresponding to the adversarial generation network in the next round.
[0106] To further implement unsupervised learning of the local generator based on multimodal data, which can make full use of the multimodal relationships of local data without relying on high-quality data annotation, in a first cross-modal retrieval model training method provided in an embodiment of the present application, the training loss with data quality constraints is solved based on the training loss function of the adversarial generation network;
[0107] Among them, the training loss function of the adversarial generation network includes: the generator training loss function L gen and the discriminator training loss function L disc ;
[0108] The generator training loss function L gen is solved based on the client loss L local and the preset generator loss L G ; The discriminator training loss function L disc is solved based on the client loss L local and the preset discriminator loss L D ;
[0109] Among them, the client loss L loca1 is solved based on the intra-class and inter-class contrast alignment loss L c , the distribution alignment loss L J and the hash loss L h used to train the local cross-modal retrieval model;
[0110] The intra-class and inter-class contrast alignment loss L c and the distribution alignment loss L J are respectively solved based on the local prototype with the same pseudo-label as the local enhanced data, the first-modal enhanced feature and the second-modal enhanced feature corresponding to the local enhanced data.
[0111] To further improve the effectiveness and reliability of the unsupervised learning of the local generator, in a first cross-modal retrieval model training method provided in an embodiment of the present application, the intra-class and inter-class contrast alignment loss L c is solved based on the indirect contrast loss L ind and the direct contrast loss L d ;
[0112] The indirect contrast loss L ind is solved based on the indirect contrast losses respectively corresponding to the first modality and the second modality belonging to different modality types; for example, based on the indirect contrast loss of the image modality and the indirect contrast loss of the text modality ;
[0113] Among them, the indirect contrast loss corresponding to the first modality is solved based on the first modality enhanced feature, the local prototype having the same pseudo-label as the local enhanced data to which the first modality enhanced feature belongs, and a preset first temperature coefficient; for example: based on the image modality enhanced feature and the image modality enhanced feature the local prototype having the same pseudo-label as the local enhanced data F g to which it belongs and a preset first temperature coefficient σ1 are solved to obtain
[0114] The indirect contrast loss corresponding to the second modality is solved based on the second modality enhanced feature, the local prototype having the same pseudo-label as the local enhanced data to which the second modality enhanced feature belongs, and the first temperature coefficient σ1; for example: based on the text modality enhanced feature and the text modality enhanced feature the local prototype having the same pseudo-label as the local enhanced data F g to which it belongs and a preset first temperature coefficient σ1 are solved to obtain
[0115] The direct contrast loss L d is solved based on a preset second temperature coefficient σ2, the first modality enhanced feature and the second modality enhanced feature corresponding to the local enhanced data; for example: based on a preset second temperature coefficient σ2, the local enhanced data F g the corresponding image modality enhanced feature and the text modality enhanced feature to obtain the direct contrast loss Ld.
[0116] In order to further improve the effectiveness and reliability of the unsupervised learning of the local generator, in a first cross-modal retrieval model training method provided in an embodiment of the present application, the distribution alignment loss L J is solved based on the loss of minimizing the difference between the pseudo-label distributions of the first modality enhanced feature and the second modality enhanced feature and a preset weight hyperparameter; for example: according to the image modality enhanced feature and the text modality enhanced feature and a preset weight hyperparameter α to obtain the loss of minimizing the difference
[0117] Among them, the loss for minimizing the difference between the pseudo-label distributions of the first-modal enhancement feature and the second-modal enhancement feature is solved based on the pseudo-label distribution of the first-modal enhancement feature, the pseudo-label distribution of the second-modal enhancement feature, and the average parameter; the value of the average parameter is the sum average of the pseudo-label distribution of the first-modal enhancement feature and the pseudo-label distribution of the second-modal enhancement feature; for example, according to the image-modal enhancement feature of the pseudo-label distribution the text-modal enhancement feature of the pseudo-label distribution and the average parameter Z, the loss for minimizing the difference is solved Among them, the average parameter Z is the pseudo-label distribution of the image-modal enhancement feature and the pseudo-label distribution of the text-modal enhancement feature of the sum average.
[0118] The pseudo-label distribution of the first-modal enhancement feature is solved based on the first-modal enhancement feature and the local prototype that has the same pseudo-label as the local enhanced data to which the first-modal enhancement feature belongs; for example, based on the image-modal enhancement feature and the local prototype that has the same pseudo-label as the local enhanced data to which the image-modal enhancement feature belongs, the pseudo-label distribution of the image-modal enhancement feature is solved
[0119] The pseudo-label distribution of the second-modal enhancement feature is solved based on the second-modal enhancement feature and the local prototype that has the same pseudo-label as the local enhanced data to which the second-modal enhancement feature belongs, for example, based on the text-modal enhancement feature and the local prototype that has the same pseudo-label as the local enhanced data to which the text-modal enhancement feature belongs, the pseudo-label distribution of the text-modal enhancement feature is solved
[0120] Furthermore, based on the foregoing first cross-modal retrieval model training method, the present application also provides a second cross-modal retrieval model training method that can be implemented by a second cross-modal retrieval model training device. This second cross-modal retrieval model training device can be set in a server, see Figure 4 , the second cross-modal retrieval model training method specifically includes the following content:
[0121] Step 300: Receive local data from each client device in the current round of federated learning. Each piece of local data contains: the corresponding local cross-modal retrieval model, the local prototype, and the local generator; and the local cross-modal retrieval model, the local prototype, and the local generator are pre-obtained by the client device based on the first cross-modal retrieval model training method.
[0122] It can be understood that the first cross-modal retrieval model training method can adopt the first cross-modal retrieval model training method executed by the first cross-modal retrieval model training device provided in the client device in the foregoing embodiments. Its specific implementation can refer to the embodiments of the first cross-modal retrieval model training method, which will not be elaborated here.
[0123] Step 400: Obtain global training data according to the local prototypes and local generators corresponding to the current round of federated learning, and train the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0124] As can be seen from the above description, the second cross-modal retrieval model training method provided in the embodiments of the present application can, in the process of training the global cross-modal retrieval model based on networked learning, achieve unsupervised learning of the local generator based on multi-modal data, make full use of the multi-modal relationship of local data, and do not rely on high-quality data annotation. Furthermore, it can improve the efficiency and reliability of global cross-modal retrieval model training while ensuring the privacy of local data, and can effectively solve the deviation problem of the global cross-modal retrieval model.
[0125] In order to further improve the effectiveness and reliability of aggregated global training, and effectively solve the deviation problem of the global cross-modal retrieval model, in the second cross-modal retrieval model training method provided in the embodiments of the present application, refer to Figure 5 , the specific content of step 400 includes the following:
[0126] Step 410: Cluster each local cross-modal retrieval model corresponding to the current round of federated learning to obtain the corresponding global cross-modal retrieval model.
[0127] Step 420: Cluster each local prototype corresponding to the current round of federated learning to obtain the corresponding global prototype.
[0128] And step 430: Based on the proportion of local multi-modal data corresponding to each source, generate corresponding local enhanced data through each global model respectively and jointly constitute the corresponding global enhanced data.
[0129] Step 440: Discretely filter the global enhanced data based on the global prototype to obtain corresponding global training data, and use the global training data to train the global cross-modal retrieval model.
[0130] From a software perspective, the present application also provides a first cross-modal retrieval model training device for executing all or part of the first cross-modal retrieval model training method. Refer to Figure 6 The first cross-modal retrieval model training device specifically includes the following:
[0131] An unsupervised training module 10, configured to perform unsupervised training on a local adversarial generation network with the cross-modal splicing features corresponding to local multi-modal data and local prototypes obtained by clustering the cross-modal splicing features in advance, so as to train the adversarial generation network into a local generator for generating local enhanced data.
[0132] A local data sending module 20, configured to send the local local prototype, the local generator, and a local cross-modal retrieval model locally trained based on the feature vectors of various modalities corresponding to local multi-modal data to a server, so that the server obtains global training data according to the received local prototypes and local generators, and enables the server to train a global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0133] The embodiment of the first cross-modal retrieval model training device provided by the present application can specifically be used to execute the processing flow of the embodiment of the first cross-modal retrieval model training method in the above embodiment, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiment of the first cross-modal retrieval model training method above.
[0134] Part of the cross-modal retrieval model training of the first cross-modal retrieval model training device can be completed in a client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. The present application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor for specific processing of cross-modal retrieval model training.
[0135] The above-mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side. In other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform having a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0136] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols not yet developed on the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also, for example, include RPC protocol (Remote Procedure Call Protocol) and REST protocol (Representational State Transfer) used on top of the above-mentioned protocols.
[0137] As can be seen from the above description, the first cross-modal retrieval model training device provided by the embodiments of this application can, during the process of implementing global cross-modal retrieval model training based on networked learning, achieve unsupervised learning of the local generator based on multi-modal data, can make full use of the multi-modal relationship of local data, and does not need to rely on high-quality data annotation. Furthermore, it can improve the efficiency and reliability of global cross-modal retrieval model training while ensuring the privacy of local data, and can effectively solve the offset problem of the global cross-modal retrieval model.
[0138] Perform discrete value filtering on the global enhanced data based on the global prototype to obtain corresponding global training data, and use the global training data to train the global cross-modal retrieval model.
[0139] At the software level, this application also provides a second cross-modal retrieval model training device for executing all or part of the second cross-modal retrieval model training method. See Figure 7 , the second cross-modal retrieval model training device specifically includes the following content:
[0140] The local data receiving module 30 is used to receive local data from each client device respectively in the current round of federated learning. Each of the local data contains: the corresponding local cross-modal retrieval model, the local prototype, and the local generator; and the local cross-modal retrieval model, the local prototype, and the local generator are pre-obtained by the client device based on the cross-modal retrieval model training method executable by the client.
[0141] The global model training module 40 is used to obtain global training data according to the respective local prototypes and the respective local generators corresponding to the current round of federated learning, and train the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
[0142] The embodiment of the second cross-modal retrieval model training device provided in this application can specifically be used to execute the processing flow of the embodiment of the second cross-modal retrieval model training method in the above embodiment, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiment of the second cross-modal retrieval model training method above.
[0143] The part of the second cross-modal retrieval model training device for cross-modal retrieval model training can be completed in the server.
[0144] The server can be communicatively connected to the client device to achieve data transmission with the server. The server can include a server on the task scheduling center side. In other implementation scenarios, it can also include a server of an intermediate platform, such as a server of a third-party server platform communicatively linked to the task scheduling center server. The server can include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0145] As can be seen from the above description, the second cross-modal retrieval model training device provided in the embodiment of this application can, in the process of training the global cross-modal retrieval model based on networked learning, achieve unsupervised learning of the local generator based on multi-modal data, can make full use of the multi-modal relationship of local data, does not need to rely on high-quality data annotation, and can thus improve the efficiency and reliability of global cross-modal retrieval model training while ensuring the privacy of local data, and can effectively solve the deviation problem of the global cross-modal retrieval model.
[0146] Furthermore, this application also provides an embodiment of an unsupervised federated learning system, including: a server, and each client device communicatively connected to the server respectively;
[0147] The client device includes a first memory, a first processor, and a first computer program stored on the first memory and executable on the processor. When the first processor executes the first computer program, the foregoing first cross-modal retrieval model training method is implemented;
[0148] The server includes a second memory, a second processor, and a second computer program stored on the second memory and executable on the processor. When the second processor executes the second computer program, the foregoing second modal retrieval model training method is implemented.
[0149] To further illustrate the above embodiments, the present application also provides an application example of a cross-modal retrieval model training method implemented using an unsupervised federated learning system, that is, a new FedWGAN method for unsupervised cross-modal retrieval. To alleviate the global model drift problem caused by non-independent and identically distributed data and make full use of the data of each client to obtain a global model with generalization and robustness, it is a means worthy of research to make full use of the information of the client to train the data generation model, because each client can learn a data generator containing the characteristics of its own dataset to assist in the training of the global model. First, a clustering method is used to obtain a pseudo-label and a batch of local prototypes for the real data, and then the class relationship and data distribution similarity between the real multi-modal data and the generated multi-modal data are used to further constrain the quality of the generated data to obtain a high-quality local generator. The generator trained locally is sent to the server to generate data. The local prototypes are also uploaded to the server to calculate the global cluster prototypes, which are used to filter out the outliers in the generated data in the server. Finally, the filtered data is input into the aggregated global retrieval model to ensure that the offset problem of the final global model is solved with non-independent and identically distributed generated data.
[0150] In this application example, the first modality is taken as the image modality and the second modality is taken as the text modality for illustration. The problem to be solved in this application example is essentially the problem of cross-modal retrieval in a federated scenario. Given a multi-modal dataset composed of N image-text pairs where I n and T n are the image and text of the nth sample respectively. Cross-modal hashing retrieval aims to obtain the hash representation of the image modality and the hash representation of the text modality where H represents the length of the final hash code. The vector similarity matching algorithm can be used for retrieval between B and B T respectively.
[0151] The overall local training process and global training process are shown in Table 1 below:
[0152] Table 1
[0153]
[0154]
[0155] See Figure 8 On each client, the images and texts are first input into CNN-F and BERT to extract their features F I and F T Then, the local prototype C is calculated using their concatenated result F and provided to the WGAN to generate synthetic features F g Then, the synthetic features F g are divided into and To improve the quality of and the WGAN is additionally constrained by two loss functions proposed in this application example: the intra-class and inter-class contrast loss between real data and generated data, and the JSD loss for data distribution alignment. After the WGAN and the retrieval model UCCH are locally trained k this application example sends them together with the local prototype C to the server to calculate the global prototype C G and the global UCCH model. The uploaded WGAN generator is used to create synthetic data and further filtered based on their minimum distance to the global prototype to remove outliers. The remaining synthetic data is used to train the global UCCH model on the server, and the trained global model is sent back to the client to start the next round of communication.
[0156] Among them, this application example is mainly divided into four modules: the WGAN local trainer module, the intra-class and inter-class contrast alignment module, the distribution alignment module supervised by local prototypes, and the global model correction module supervised by global prototypes. This application example proposes an abbreviation FedWGAN (Generating Synthetic Data for Unsupervised Federated Learning of Cross-Modal Retrieval). This abbreviation is composed of two parts, Fed and WGAN. Fed represents federated learning, and WGAN represents the core WGAN network method used in this article to implement an unsupervised cross-modal retrieval method guided by generated data in a federated scenario. This application example is described in detail in the form of each module:
[0157] 1) WGAN local trainer module
[0158] Since the private non-IID data of the client cannot be transmitted to the server, a data generator needs to be trained to simulate the original data and the generator is transmitted to the server. The generator should generate data with the same characteristics and distribution as the data on the client. Considering the stability and easy convergence of the training process, the Wasserstein Generative Adversarial Network (WGAN) is adopted to learn data generation. Similar to PT-FUCH, the pre-trained CNNF and Bert are used to extract the image feature F I (feature of the image, the feature of the image) and the text feature F T (feature of the text, the feature of the text). Then, after performing MLP projection, they are concatenated into F (concatenated feature):
[0159] F = concat(MLP I (F I ), MLP T (F T )) (1)
[0160] WGAN consists of a generator (G) and a discriminator (D). The input of WGAN is Gaussian noise ∈. The generator tries to generate as realistic samples as possible from the noise to deceive the discriminator. The traditional generator loss is defined as:
[0161]
[0162] refers to the expectation of the noise ∈ under the p ∈ distribution
[0163] where p ∈ represents the distribution of the noise ∈. The discriminator tries to maximize the score of the real samples F while minimizing the score of the generated samples G(∈).
[0164]
[0165] p data refers to the probability distribution of the real input data; F ~ p data means that F is sampled from the p data distribution;
[0166] The output of the generator (i.e., the locally enhanced data) is written as F g (i.e., the generated feature), which should have a similar and the same distribution as F. The locally enhanced data F g is divided into the generated image enhancement feature (i.e., the generated image feature) and the text enhancement feature (i.e., generated text feature). In addition to the above traditional WGAN loss function, three new loss constraints are designed in the following two parts to further improve the quality of the generated data.
[0167] 2) Intra-class and inter-class contrast alignment module
[0168] Since there are no data labels, first use K-means to cluster the fused feature F to obtain local class prototypes ( represents the feature representation of the M-th prototype in the k-th client, where M is the number of prototypes and k represents the k-th client. These prototypes are the anchors in the indirect constraint process, i.e.: use these prototypes to assign pseudo-labels separately for the generated local enhanced data F g separately.
[0169] In this way and the indirect contrast constraints within and between classes can be performed using pseudo-labels. Taking the image modality as an example, the image features assigned the same pseudo-label should be close to each other and far from other prototypes. The same constraint also applies to the text modality. This indirect loss can be expressed as:
[0170]
[0171] represents the i-th sample of the image enhancement feature (i.e., i-th generated image feature);
[0172] represents the i-th sample of the text enhancement feature (i.e., i-th generated text feature);
[0173] where represents the prototype on the k-th client that has the same pseudo-label as , Sim represents the cosine similarity function, σ1 is the temperature coefficient, and i represents the characteristics of the i-th sample. Summing them up obtains the indirect contrast loss, which encourages the features of different modalities to be close to the corresponding local prototypes.
[0174]
[0175] Another thing to note is that the image-text features are generated simultaneously. Therefore, it is necessary to ensure that they have the same label. For this purpose, a direct contrast loss constraint is designed:
[0176]
[0177] The pseudo-label of the i-th sample representing the image enhancement feature ;
[0178] The pseudo-label of the i-th sample representing the text enhancement feature ;
[0179] where σ2 is also a temperature coefficient. The total loss of the intra-class and inter-class contrast alignment module is:
[0180] L c = L ind + L d (8)
[0181] 3) Local prototype supervised distribution alignment module
[0182] To further align the generated image features and text features, a JSD-based loss is designed from the perspective of their pseudo-label distributions. First, calculate the pseudo-label distributions of the generated image features and text features and where C is the local prototype on each client. The similarity between the feature and the local prototype C is used to represent the distribution of the feature. Then use the JSD divergence to measure the and distribution difference between. Minimizing the difference between the two distributions enables the model to learn structural knowledge invariant to data augmentation. The calculation formula of this loss is:
[0183]
[0184] The same principle also applies to the characteristics of real data to improve the prototype for subsequent iterations. Finally, the local prototype supervised distribution alignment loss is obtained:
[0185]
[0186] (Q I ||Q T ) is a relatively common representation used to represent the calculation of the Q I between Q T JSD divergence;
[0187] where α is the weight hyperparameter, which gradually decreases from 1 to 0.5. This is because the generation effect of the generator is not ideal enough at the beginning and it will become better as the training progresses. So the value of α will be gradually increased as the generation effect improves.
[0188] According to the method of UCCH, use the hash loss L hAs the local unsupervised hash representation learning loss of the entire framework, the client loss L in this federated learning framework is aggregated local = L h + λL c + γL J , where λ and γ are hyperparameters used to control the weights of each loss. Incorporating L local into the training process of WGAN, the generator training loss and discriminator training loss can be respectively:
[0189] L gen = L local + L G (12)
[0190] L disc = L local + L D (13)
[0191] 4) Global model correction module supervised by global prototypes
[0192] Because the data generated by the local generator is to be used to assist the global model obtained after aggregating the local models. However, since the generated data is randomly generated from input noise, these synthetic data may not be consistent with the true data distribution. To improve the quality of the synthetic dataset, a global model correction module based on global prototype supervision is designed to improve the quality of the generated data and enhance the performance of the global model.
[0193] First, the generators G, local prototypes C, and local retrieval models from all clients are uploaded to the server. Then the server aggregates the retrieval model UCCH and applies K-means to the uploaded local prototypes to obtain a set of global prototypes On the other hand, according to the number of samples on each client, the corresponding generator generates data on the server, where μ is a proportional hyperparameter to determine the number of samples generated on the server, and D k represents the number of samples on the k-th client.
[0194] Then, the global prototypes are used to filter the synthetic data, and only the data is retained when the minimum distance from the synthetic data to the global prototype is less than or equal to the threshold:
[0195]
[0196] This represents the global data generated by the generator uploaded to the server on the server, where G here represents global, g represents generate, and i represents the i-th sample;
[0197] It represents the prototype of the m-th class globally. C represents the prototype, G represents the global, and m represents the clustering of samples of the m-th class.
[0198] Among them, θ is a learnable parameter to determine and refine the threshold for screening synthetic data. The screened data is used to train the global retrieval model. After the training is completed, the updated global retrieval model will be sent back to each client for subsequent training.
[0199] That is to say, the application example of this application proposes a new FedWGAN method for unsupervised cross-modal retrieval, which alleviates the data non-IID problem through a data generation strategy. It can benefit many existing federated learning methods; this application example designs two new loss constraints for WGAN to generate high-quality local client multi-modal data in an unsupervised scenario; this application example also designs a global outlier removal strategy on the server side to further improve the quality of the generated data.
[0200] In summary, the key point of this application example is to train a multi-modal data generator for each client that can generate multi-modal data close to the current client's features and data distribution in an unsupervised federated scenario by combining the inter-modal relationships between real features and generated features. And, after obtaining efficient data generators for each client, upload the local prototypes, generators, and cross-modal retrieval models of the clients to the server. Aggregate the local prototypes into global prototypes and the local cross-modal retrieval models into global cross-modal retrieval models. Combine the global model to screen the data generated by each client, and use the screened data to retrain the global model to achieve the purpose of eliminating model drift.
[0201] The embodiment of this application also provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the first cross-modal retrieval model training method or the second cross-modal retrieval model training method mentioned in the above embodiment. The processor and the memory can be connected through a bus or other means. Taking the bus connection as an example. The receiver can be connected to the processor and the memory in a wired or wireless manner.
[0202] The processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., such as chips, or combinations of the above types of chips.
[0203] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the first cross-modal retrieval model training method or the second cross-modal retrieval model training method in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor executes various functional applications and data processing of the processor, that is, implements the first cross-modal retrieval model training method or the second cross-modal retrieval model training method in the above method embodiments.
[0204] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0205] The one or more modules are stored in the memory and, when executed by the processor, execute the first cross-modal retrieval model training method or the second cross-modal retrieval model training method in the embodiments.
[0206] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.
[0207] As an implementation, the functions of the receiver and transmitter in the present application can be considered to be implemented by a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented by a dedicated processing chip, a processing circuit or a general-purpose chip.
[0208] As another implementation, it can be considered to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, receiver and transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, receiver and transmitter by executing the codes in the memory.
[0209] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing first cross-modal retrieval model training method or the second cross-modal retrieval model training method are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium well-known in the technical field.
[0210] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing first cross-modal retrieval model training method or the second cross-modal retrieval model training method are implemented.
[0211] Those of ordinary skill in the art should understand that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to execute in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0212] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0213] In the present application, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0214] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A cross-modal retrieval model training method, characterized in that: include: Using the cross-modal splicing features corresponding to the local multimodal data and the local prototypes obtained by clustering the cross-modal splicing features in advance as constraints, unsupervised training is performed on the local generative adversarial network to train the generative adversarial network as a local generator for generating local enhanced data; The local local prototype, the local generator, and the local cross-modal retrieval model obtained by locally training the feature vectors of various modalities corresponding to the local multimodal data are all sent to the server, so that the server obtains global training data according to the received local prototypes and local generators, and trains the global cross-modal retrieval model corresponding to each local cross-modal retrieval model based on the global training data.
2. The cross-modal retrieval model training method according to claim 1, characterized in that: Before performing unsupervised training on a local generative adversarial network using the cross-modal splicing features corresponding to the local multi-modal data and the local prototype obtained by clustering the cross-modal splicing features in advance as constraints, the method further includes: Dividing the local multimodal data into training data pairs, wherein each of the training data pairs includes first modal data and second modal data belonging to different modal types; Inputting the first modality data in each of the training data pairs into a feature extraction model corresponding to the first modality, and inputting the second modality data in each of the training data pairs into a feature extraction model corresponding to the second modality, so that the feature extraction model corresponding to the first modality outputs a first modality feature vector corresponding to the first modality data, and the feature extraction model corresponding to the second modality outputs a second modality feature vector corresponding to the second modality data; Using each of the training data to respectively correspond to the first modal feature vector and the second modal data and the hash loss to train a local cross-modal retrieval model, so that the local cross-modal retrieval model is used to respectively obtain the hash representations corresponding to the first modal data and the second modal data; And, projecting and splicing the first modal feature vector and the second modal data corresponding to each of the training data pairs, respectively, to obtain the cross-modal splicing features corresponding to each of the training data pairs; The cross-modal concatenation features corresponding to each of the training data pairs are clustered and projected to obtain a local prototype corresponding to the local multimodal data.
3. The cross-modal retrieval model training method according to claim 1, characterized in that: The method uses the cross-modal splicing features corresponding to the local multimodal data and the local prototype obtained by clustering the cross-modal splicing features in advance as constraints to perform unsupervised training on the local generative adversarial network to train the generative adversarial network as a local generator for generating local enhanced data, including: The local generative adversarial network is trained iteratively for multiple rounds using the cross-modal concatenation features corresponding to the local multimodal data, and the preset training steps are executed in each round. Wherein, the training step comprises: If the current iteration is the first round, the cross-modal splicing features corresponding to the local multimodal data are input into the local generative adversarial network, so that the generative adversarial network outputs the local enhanced data of the current round; if the current iteration is not the first round, the local generative adversarial network is adjusted based on the training loss obtained in the previous round, and the cross-modal splicing features corresponding to the local multimodal data are input into the adjusted generative adversarial network, so that the generative adversarial network outputs the local enhanced data of the current round; assigning a pseudo label to the local enhanced data according to the local prototype; And, dividing the local enhancement data into first modality enhancement features and second modality enhancement features according to modality types; Based on the first modality enhancement feature and the second modality enhancement feature corresponding to the local enhancement data, and the local prototype with the same pseudo label as the local enhancement data, the training loss with data quality constraint corresponding to the adversarial generative network in the next round is obtained.
4. The cross-modal retrieval model training method according to claim 3, characterized in that: The training loss with data quality constraints is solved based on the training loss function of the generative adversarial network; The training loss function of the adversarial generative network includes: a generator training loss function and a discriminator training loss function; The generator training loss function is solved based on the client loss and the preset generator loss; the discriminator training loss function is solved based on the client loss and the preset discriminator loss; The client loss is solved based on intra-class and inter-class comparison alignment loss, distribution alignment loss, and hash loss for training the local cross-modal retrieval model; The intra-class and inter-class comparison alignment losses and the distribution alignment loss are respectively solved based on the local prototype having the same pseudo label as the local enhancement data, the first modality enhancement feature and the second modality enhancement feature corresponding to the local enhancement data.
5. The cross-modal retrieval model training method according to claim 4, characterized in that: The intra-class and inter-class contrast alignment losses are solved based on the indirect contrast loss and the direct contrast loss; The indirect contrast loss is solved based on the indirect contrast loss corresponding to the first mode and the second mode, which belong to different modality types respectively; The indirect contrast loss corresponding to the first modality is solved based on the first modality enhancement feature, the local prototype having the same pseudo label as the local enhancement data to which the first modality enhancement feature belongs, and a preset first temperature coefficient; The indirect contrast loss corresponding to the second modality is solved based on the second modality enhancement feature, the local prototype having the same pseudo label as the local enhancement data to which the second modality enhancement feature belongs, and the first temperature coefficient; The direct contrast loss is solved based on a preset second temperature coefficient, a first modal enhancement feature and a second modal enhancement feature corresponding to the local enhancement data.
6. The cross-modal retrieval model training method according to claim 4, characterized in that: The distribution alignment loss is based on minimizing the difference between the pseudo-label distributions of the first modality enhancement feature and the second modality enhancement feature and solving a preset weight hyperparameter; Wherein, the difference minimization loss between the pseudo label distribution of the first modality enhancement feature and the second modality enhancement feature is solved based on the pseudo label distribution of the first modality enhancement feature, the pseudo label distribution of the second modality enhancement feature and an average parameter; the value of the average parameter is the summed average of the pseudo label distribution of the first modality enhancement feature and the pseudo label distribution of the second modality enhancement feature; The pseudo-label distribution of the first modality enhancement feature is solved based on the first modality enhancement feature and the local prototype having the same pseudo-label as the local enhancement data to which the first modality enhancement feature belongs; The pseudo-label distribution of the second modality enhancement feature is solved based on the second modality enhancement feature and the local prototype having the same pseudo-label as the local enhancement data to which the second modality enhancement feature belongs.
7. A cross-modal retrieval model training method, characterized in that: include: In the current federated learning round, local data is received from each client device respectively, wherein each of the local data includes: the local cross-modal retrieval model, the local prototype and the local generator that correspond to each other; and the local cross-modal retrieval model, the local prototype and the local generator are pre-acquired by the client device based on the cross-modal retrieval model training method according to any one of claims 1 to 6; Global training data is obtained according to each of the local prototypes and each of the local generators corresponding to the current federated learning round, and a global cross-modal retrieval model corresponding to each of the local cross-modal retrieval models is trained based on the global training data.
8. The cross-modal retrieval model training method according to claim 7, characterized in that: The acquiring global training data according to each of the local prototypes and each of the local generators corresponding to the current federated learning round, and training the global cross-modal retrieval model corresponding to each of the local cross-modal retrieval models based on the global training data, includes: Clustering the local cross-modal retrieval models corresponding to the current federated learning round to obtain a corresponding global cross-modal retrieval model; Clustering the local prototypes corresponding to the current federated learning round to obtain the corresponding global prototype; And, based on the proportion of the local multimodal data corresponding to each source, respectively generate the corresponding local enhancement data through each of the global models to jointly constitute the corresponding global enhancement data; Based on the global prototype, the global enhanced data is subjected to discrete value filtering to obtain corresponding global training data, and the global cross-modal retrieval model is trained using the global training data.
9. An unsupervised federated learning system, characterized in that: include: A server, and client devices respectively connected to the server for communication; The client device includes a first memory, a first processor, and a first computer program stored in the first memory and executable on the processor, wherein the first processor implements the cross-modal retrieval model training method according to any one of claims 1 to 6 when executing the first computer program; The server includes a second memory, a second processor, and a second computer program stored in the second memory and executable on the processor, and the second processor implements the cross-modal retrieval model training method as described in claim 7 or 8 when executing the second computer program.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the cross-modal retrieval model training method as described in any one of claims 1 to 6, or implements the cross-modal retrieval model training method as described in claim 7 or 8.
Citation Information
Patent Citations
Full convolution video description generation method based on self-optimization mechanism
CN111464881A
Cross-modal retrieval method, device and equipment and computer readable storage medium
CN112487217A
Unbalanced data-oriented federal cross-modal retrieval method and system
CN116244484A
Content augmentation with machine generated content to meet content gaps during interaction with target entities
US20230121711A1
Cross-modal hash retrieval method, terminal device, and storage medium
WO2022104540A1
Cited By
Cross-college practical training data sharing method based on data privacy
CN121277901A