A movie classification system based on multi-modal deep representation collaborative federated learning

CN118468098BActive Publication Date: 2026-08-21EAST CHINA UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410475258.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2026-08-21
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

首先,基于损失的联邦学习方法在处理客户端数据时,并未充分考虑电影多模态数据间的差异性和互补性,限制了模型的学习效果

Benefits of technology

[0014]本发明的有益效果是:本发明提出一种基于多模态深度表征协同联邦学习的电影分类系统,以克服当前多模态联邦学习方法在模态协同和特征提取方面的局限性。本发明充分利用全局多模态深度表征,旨在提高参与联邦训练各客户端上本地模型在电影多模态数据集上的分类性能。本发明引入多模态深度表征的概念,其包含模态固有表征与模态融合表征。通过构建关注模态间差异性和互补性的多模态协同损失,本发明能更全面地理解电影多模态数据,从而提升分类准确率。同时,本发明采用协同个性化参数聚合策略来解决不同客户端上的模型多模态融合能力差异的问题,通过聚合距离本地模态融合表征较远的其它客户端的参数,有效增强了各客户端上多模态模型的泛化能力,进一步提高了各客户端上本地模型的性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118468098B_ABST
    Figure CN118468098B_ABST
Patent Text Reader

Abstract

The application discloses a movie classification system based on multi-modal deep representation collaborative federated learning, and the method comprises the following steps: firstly, initializing local model parameters of each client with public model parameters; then, the client updates the local model parameters according to the total loss on the local movie multi-modal data, obtains the local multi-modal deep representation, and uploads the local multi-modal deep representation and the local model parameters to the server; the server calculates the global multi-modal deep representation, obtains the initialization local model parameters of each client in the next global round through personalized parameter aggregation, and distributes the local model parameters to each client; finally, each client replaces the local model parameters. Except for the initialization, the above training process is repeatedly executed until the local model parameters of each client converge. For the to-be-recognized movie data, only the category thereof is recognized in the local client. While protecting the privacy, the application effectively improves the classification accuracy of each client on the movie multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a film classification system based on multimodal deep representation collaborative federated learning, belonging to the field of film data classification technology. Background Technology

[0002] With the continuous development of the film industry, a large number of new films are released every year. Classifying films is a complex and interesting task, involving multiple dimensions and considerations. Traditional classification methods may rely on manual annotation or fixed rules, but these methods are often inefficient, subjective, and struggle to handle the large volume and diversity of film content. Machine learning algorithms can learn from massive amounts of data and extract useful features, thus more accurately identifying film categories. Compared to traditional rule-based or manual annotation methods, machine learning algorithms offer higher classification accuracy and better meet user needs. Machine learning algorithms can automatically adapt to new data and changes, continuously updating and optimizing classification models. This means that even if the content and style of a film change, machine learning algorithms can quickly adjust and adapt to these changes, maintaining classification accuracy. Film is a complex media form containing multiple modalities of information, including images, sound, and text. Traditional classification methods often focus only on a single modality, such as classifying based solely on image or text content, which may lead to incomplete and biased information. Multimodal fusion methods, on the other hand, can simultaneously utilize information from multiple modalities in a film, such as images, audio, and subtitles, thus gaining a more comprehensive understanding of the film's content. Because data from different modalities may have different characteristics, resolutions, and noise levels, direct fusion may introduce errors. Federated multimodal fusion methods, however, can align and normalize data from different modalities through appropriate preprocessing and feature extraction techniques, thereby reducing distributional differences and improving fusion results. In film classification tasks, large amounts of user data and privacy information are often involved. By employing federated learning, multimodal fusion can be performed without sharing the original data, thus protecting user privacy and data security. Federated multimodal fusion methods also offer better privacy protection and data security.

[0003] Current multimodal federated learning methods largely focus on model performance on the client side, failing to fully leverage the advantages of global movie multimodal data for collaborative learning and feature extraction. First, loss-based federated learning methods do not adequately consider the differences and complementarities among movie multimodal data when processing client-side data, limiting the model's learning effectiveness. Second, existing personalized federated learning methods primarily address the issue of data distribution differences on the client side, neglecting the differences in multimodal model fusion capabilities across different clients. These issues may lead to difficulties for the model when processing different modalities or even different categories of data. Summary of the Invention

[0004] This invention proposes a movie classification system based on multimodal deep representation collaborative federated learning. In each global round, each client updates its local model parameters and obtains its local multimodal deep representation based on the total loss on its local movie multimodal data, and uploads both the local model parameters and the local multimodal deep representation to the server. The total loss includes classification loss and modal collaboration loss, which in turn includes modal intrinsic loss and modal fusion loss. The server uses the local multimodal deep representations uploaded by each client to calculate the global multimodal deep representation, and aggregates personalized parameters to calculate the initial local model parameters for each client in the next global round, further improving the classification accuracy of each client on the movie multimodal data.

[0005] A film classification system based on multimodal deep representation collaborative federated learning, characterized by the following steps: S1: Initialize the local model parameters of each client using the common model parameters; S2: The client updates the local model parameters based on the total loss on the local movie multimodal data, obtains the local multimodal depth representation, and uploads the local multimodal depth representation and local model parameters to the server; S3: The server calculates the global multimodal depth representation and obtains the initial local model parameters of each client in the next global round through personalized parameter aggregation, and sends them to each client; S4: Each client replaces the local model parameters; S5: Repeat S2-S4 until the accuracy of each client's local model on the local movie multimodal data no longer improves in 3 rounds of training.

[0006] The technical solution adopted in this invention can be further refined. The local multimodal deep representation in step S2 consists of two parts: local modality intrinsic representation and local modality fusion representation. (Client-side) Local modal intrinsic representation The calculation method is as follows: , in, For client number, For category numbering, Modal numbering, for , It is the first The category number on each client is Modal numbering is The The inherent representation of individual modalities on a sample.

[0007] Local multimodal deep representation, characterized by: client Local modality fusion representation The calculation method is as follows: , in, For client number, For category numbering, For the first The client number Local modality fusion representations across categories for , It is the first The category number on each client is The Individual modality fusion representation on a sample.

[0008] The technical solution adopted in this invention can be further refined. The total loss in step S2 is defined as the weighted sum of classification loss and modal cooperation loss; Client Total losses The calculation method is as follows: , in, These are the weighting coefficients. For the client Modal intrinsic loss on For the client The modal fusion loss, or modal cooperative loss, consists of intrinsic modal loss and modal fusion loss. The classification loss is calculated as follows: , in, For client number, Let cross-entropy be the loss function. For the client Local model parameters on for For the client The first Multimodal data of a movie, For the client The first Category labels for each movie's multimodal data.

[0009] The technical solution adopted in this invention can be further refined. Modal inherent loss. The calculation method is as follows: , in, For client number, For L2 norm distance loss, For category numbering, The total number of categories in the movie multimodal data. Modal numbering, The modality number of the movie multimodal data. For local modal inherent characterization, The first one issued by the server The first mode Global modality inherent representations across each category.

[0010] The technical solution adopted in this invention can be further refined. Modal fusion loss. The calculation method is as follows: , in, For client number, The L2 norm distance loss function is used. For category numbering, The total number of categories in the movie multimodal data. For local modality fusion representation, The first one issued by the server Global modality fusion representation across categories.

[0011] The technical solution adopted in this invention can be further refined. In step S3, the global multimodal depth representation calculated by the server includes both the inherent global modality representation and the fused global modality representation. Global Modal Intrinsic Representation The formula is: , in, For client number, The total number of clients participating in federated training. This is an inherent representation of the local modality; Global modality fusion representation The formula is: , in, For client number, The total number of clients participating in federated training. This is a local modality fusion representation.

[0012] The technical solution adopted in this invention can be further refined. In the process of personalized parameter aggregation in step S3, the correlation between parameters of each client is first calculated by canonical correlation analysis distance matrix, and then the initialization parameters on each client in the next global round are calculated by personalized parameter aggregation; The calculation method for personalized parameter aggregation is as follows: , in The number of iterations. , The total number of clients participating in federated training. For client number, For the global round number, The number of clients aggregated. It is a set that stores the first... The client IDs that have completed the aggregation of personalized parameters in the previous step are now set. Storage and client The furthest distance in the canonical correlation analysis distance matrix The client's ID, For the client number in the set, For the first In each global round, the client The parameters of the pre-trained local model, when At that time, perform initialization operations. For the client The initialization aggregation parameter on was assigned a value , Indicates the number of the next global round. Initialize local model parameters on each client.

[0013] The technical solution adopted in this invention can be further refined. (Client) and client Canonical correlation analysis distance The calculation method is as follows: , in, For category numbering, and For the IDs of two different clients, The total number of categories in the movie multimodal data. This indicates the calculation of absolute value. and Clients respectively With the client Local modality fusion representation on and The results after canonical correlation analysis.

[0014] The beneficial effects of this invention are as follows: This invention proposes a movie classification system based on multimodal deep representation collaborative federated learning to overcome the limitations of current multimodal federated learning methods in modality collaboration and feature extraction. This invention fully utilizes global multimodal deep representations to improve the classification performance of local models on various clients participating in federated training on movie multimodal datasets. This invention introduces the concept of multimodal deep representations, which includes modality-intrinsic representations and modality fusion representations. By constructing a multimodal collaborative loss that focuses on the differences and complementarities between modalities, this invention can more comprehensively understand movie multimodal data, thereby improving classification accuracy. Simultaneously, this invention employs a collaborative personalized parameter aggregation strategy to address the problem of differences in multimodal fusion capabilities among models on different clients. By aggregating parameters from other clients that are far removed from the local modality fusion representation, the generalization ability of multimodal models on each client is effectively enhanced, further improving the performance of local models on each client. Attached Figure Description

[0015] Figure 1 This is a framework diagram of a movie classification system based on multimodal deep representation collaborative federated learning, as proposed in this invention.

[0016] Figure 2 This is a schematic diagram of the personalized parameter aggregation process in a movie classification system based on multimodal deep representation collaborative federated learning, according to the present invention. Detailed Implementation

[0017] The technical solution adopted in this invention can be further refined. The following is a detailed description of the implementation methods:

[0018] Step 1: Initialization phase. Initialize the local model parameters of each client using the common model parameters. The nth client initializes the same model parameters. .

[0019] Step 2: Client Local movie multimodal data Include Data, of which , This is the movie multimodal data on the nth client, and the data from each client contains... Multimodal data of movies. The first in The sample of multimodal data for the film is ,in For the nth client on the first The first sample of multimodal movie data Data for each modality, The number of label categories in the dataset on all clients is [number]. Client Upper The category labels for the multimodal data samples of movies are as follows: In each global round, each client retrieves its local movie multimodal data. Upward Local training rotation.

[0020] Step 3: When the global round At that time, the client Local training is performed on a local dataset using only classification loss, and local multimodal deep representations are obtained. Client The classification loss formula is: , in, For client number, For cross-entropy loss, For the client Model parameters on for For the training set Multimodal data of a movie, For the training set Category labels for each movie's multimodal data.

[0021] Step 4: Client In local dataset On the local multimodal depth representation, compute the local multimodal depth representation;

[0022] Step 4.1: Multimodal deep representation consists of two parts: modality-intrinsic representation and modality fusion representation. Modality-intrinsic representation captures the unique features of each modality on a specific task or dataset, providing insights for the client. The above category number is Movie multimodal data samples in local model The The average value of the feature extractor output for each modality. For the client in the current global round The parameters on; Client Local modal intrinsic representation The formula is: , in, For client number, For category numbering, Modal numbering, For the first On the first client The first mode Local modality inherent representations on each category for , It is the category number Modal numbering is The The inherent representation of individual modalities on a single data point.

[0023] Step 4.2: Modality fusion representation in multimodal deep representation is used to capture the correlation and complementarity between different modalities, providing insights for the client. The above category number is Data samples in the local model The average value of the output of the fusion layer; Client Local modality fusion representation The formula is: , in, For client number, For category numbering, For the first The client number Local modality fusion representations across categories for , It is the category number The Individual modality fusion representation on data points.

[0024] Step 5: When the global round At that time, the client Local training is performed on the local dataset using the total loss, which is a weighted sum of classification loss and modality collaboration loss, where modality collaboration loss consists of modality intrinsic loss and modality fusion loss.

[0025] Step 5.1: Calculate the modal intrinsic loss; Client Upper-mode intrinsic loss The formula is: , in, For client number, For L2 norm distance loss, For category numbering, The total number of categories in the movie multimodal dataset. Modal numbering, The number of modalities in the movie multimodal dataset. For local modal inherent characterization, The first one issued by the server The first mode Global modality inherent representations across each category.

[0026] Step 5.2: Calculate the modal fusion loss; Client Upper modal fusion loss The formula is: , in, For client number, For L2 norm distance loss, For category numbering, The total number of categories in the movie multimodal dataset. For local modality fusion representation, The first one issued by the server Global modality fusion representation across categories.

[0027] Step 5.3: Calculate the total loss, which is defined as the weighted sum of the classification loss and the modal collaboration loss, where the modal collaboration loss is the mean of the intrinsic modal loss and the modal fusion loss; Client Total losses The formula is: , in, These are the weighting coefficients. For the client calculated in step 5.1 Modal intrinsic loss on For the client calculated in step 5.2 Modal fusion loss on This is the classification loss calculated in step 3.

[0028] Step 6: Each client trains its local model parameters using the backpropagation algorithm based on the total loss. And compared with the local multimodal depth characterization calculated in steps 4.1 and 4.2. and Uploaded to the server, where For the client in the first The local model parameters at the end of the local training round.

[0029] Step 7: The server performs personalized parameter aggregation. The purpose of personalized parameter aggregation is to obtain the first... Initial parameters for each client in each global round .

[0030] Step 7.1: First, the server calculates the canonical correlation analysis distance matrix. , Each element The calculation formula is: , in, For category numbering, and For the IDs of two different clients, For category numbering, The total number of categories in the movie multimodal dataset. This indicates the calculation of absolute value. and Clients respectively With the client Local modality fusion representation on and The results after canonical correlation analysis.

[0031] Step 7.2: Specifically, the optimization objective of the canonical correlation analysis transformation is:

[0032] in and It is a transformation matrix and elements, It is a client Multimodal fusion characterization The One value, It is a client Multimodal fusion characterization The Values. yes and covariance, and yes and The variance of . This optimization problem is solved using the eigenvalue decomposition algorithm. The multimodal fusion representation of a client across all categories can be represented as: Based on the above optimization objectives, we can conclude that... , Multimodal fusion representation on two clients and The optimal linear transformation matrix between and ,in for The length of the vector. for The length of the vector. Then the linear transformation matrix. , Applied to respectively and Here, the number of eigenvalues ​​after transformation is set to 1, resulting in the transformed matrix. and The specific calculation method is as follows: , ,

[0033] Step 7.3: The server passes through The initial parameters for each client are calculated iteratively step by step. , Calculate the number of steps for the current iteration. Set a set. Store the The client IDs that have completed the aggregation of personalized parameters in the previous step are now set. Storage and client The furthest distance in the canonical correlation analysis distance matrix The client's ID, The client number in the set. When At that time, after each iterative calculation is completed, the client number that has completed the aggregation of personalized parameters will be... Add to collection middle; The formula for aggregating personalized parameters is: , in The number of iterations. , The total number of clients participating in federated training. For client number, For the global round number, The number of clients aggregated. For the first In each global round, the client The parameters of the pre-trained local model, when At that time, perform initialization operations. To initialize the aggregation parameter, it was assigned a value. , Indicates the number of the next global round. Initialize local model parameters on each client.

[0034] Step 8: The server calculates the global multimodal representation, which includes the global modality intrinsic representation and the global modality fusion representation. Global Modal Intrinsic Representation The formula is: , in, For client number, The total number of clients participating in federated training. This is an inherent representation of the local modality; Global modality fusion representation The formula is: , in, For client number, The total number of clients participating in federated training. This is a local modality fusion representation.

[0035] Step 9: The server will... Initialize local model parameters for each client in each global round. With global multimodal deep representation , Distribute to each client.

[0036] Step 10: Repeat steps 3 through 9 until the accuracy of each client's local model on the local movie multimodal data no longer improves in 3 rounds of training.

[0037] Step 11: Testing phase, client Local multimodal data of movies to be tested , which includes One modality. (The following is a list of modalities.) The input is fed into the local model, and the output for each category is taken from the last MLP layer of the local model. Predicted probability value , .in For the total number of categories in the movie multimodal dataset, the last one is... The category number corresponding to the prediction results of the data is .

[0038] Experimental Design

[0039] Experimental Dataset Selection: The MM-IMDB dataset used in this invention is for movie genre prediction. The MM-IMDB dataset is an extension of the MovieLens 20M dataset, collecting movie category, poster, and scene description information for each movie. The entire dataset contains ratings for 25,959 movies, corresponding to a multi-label classification task across 23 movie categories. This experiment uses both scene description information and poster data, representing text and image modalities respectively. The text modality uses the Google Word2Vec model to extract features; the final vocabulary contains 41,612 words, and all text is converted to lowercase before processing. Images are scaled proportionally and cropped as needed, with a size of 160×256 pixels. The VGG-16 model is used to extract image features.

[0040] The MM-IMDB dataset contains 15,552 samples as the training set, 2,608 samples as the validation set, and 7,799 samples as the test set. To simulate a real-world federated scenario, we merged the validation and test sets, resulting in 18,160 samples for the training set and 7,799 samples for the test set. In the experiment, five clients were set up, with the training and test sets evenly distributed across them. Each client had 3,632 samples for the local training set and 1,559 samples for the local test set. Additionally, the weight coefficient was set to 0.5, the learning rate to 0.01, the batch size to 128, and the global rounds to... Set to 5, local training rounds Set it to 6, and the optimizer to Adam.

[0041] The local model uses MaxoutMLP as the image feature extractor, MaxoutMLP as the speech feature extractor, and two MLP layers as the classifier. The number of neurons in these two MLP layers is set to 512 and 23, respectively. The fusion layer refers to the MLP layer with 512 neurons. The outputs of the image feature extractor and the speech feature extractor are concatenated and then input into the classifier. The final prediction result is obtained through the softmax function.

[0042] We use the Micro F1 Score to evaluate the model's performance on the MM-IMDB dataset, and the calculation formula is as follows: , , , Where TP represents the number of correctly classified positive samples, TN represents the number of correctly classified negative samples, FP represents the number of misclassified positive samples, and FN represents the number of misclassified positive samples. We evaluate the performance of the local models on each client by calculating the Micro F1 Score on the local test set.

[0043] Comparative experiment results when the number of clients is 5: Table 1 Comparative experimental results on the MM-IMDB dataset

[0044]

[0045] In the table of client training results, the first few columns correspond to the results on the local test set of each client, and the last column is the average of the results on the local test set of all clients.

[0046] As can be seen, the model achieves an average Micro F1 score of 62.10% on the MM-IMDB dataset, outperforming the three comparison methods: FedAvg, FedProx, and PerFedAvg, thus demonstrating the effectiveness of the invention. Furthermore, compared to other federated learning methods, the results of this invention are best on most clients of the MM-IMDB dataset, indicating that the invention possesses good stability.

Claims

1. A film classification system based on multimodal deep representation collaborative federated learning, characterized in that, Includes the following steps: S1: Initialize the local model parameters of each client using the common model parameters; S2: The client updates the local model parameters based on the total loss on the local movie multimodal data, obtains the local multimodal depth representation, and uploads the local multimodal depth representation and local model parameters to the server. The total loss is defined as the weighted sum of the classification loss and the modal collaboration loss. Client Total losses The calculation method is as follows: , in, These are the weighting coefficients. For the client Modal intrinsic loss on For the client The modal fusion loss, or modal cooperative loss, consists of intrinsic modal loss and modal fusion loss. The classification loss is calculated as follows: , in, For client number, Let cross-entropy be the loss function. For the client Local model parameters on for For the client The first Multimodal data of a movie, For the client The first Category labels for multimodal movie data; S3: The server calculates the global multimodal depth representation and obtains the initial local model parameters of each client in the next global round through personalized parameter aggregation, and distributes them to each client. The personalized parameter aggregation process is as follows: first, the correlation between the parameters of each client is calculated through the distance matrix of canonical correlation analysis, and then the initial parameters on each client in the next global round are calculated through personalized parameter aggregation. The calculation method for personalized parameter aggregation is as follows: , in The number of iterations. , The total number of clients participating in federated training. For client number, For the global round number, The number of clients aggregated. It is a set that stores the first... The client IDs that have completed the aggregation of personalized parameters in the previous step are now set. Storage and client The furthest distance in the canonical correlation analysis distance matrix The client's ID, For the client number in the set, For the first In each global round, the client The parameters of the pre-trained local model, when At that time, perform initialization operations. For the client The initialization aggregation parameter on was assigned a value , Indicates the number of the next global round. Initialize local model parameters on each client; S4: Each client replaces the local model parameters; S5: Repeat S2-S4 until the accuracy of each client's local model on the local movie multimodal data no longer improves in 3 rounds of training. The multimodal data refers to text modality and image modality.

2. The film classification system based on multimodal deep representation collaborative federated learning according to claim 1, characterized in that, The local multimodal deep representation in step S2 consists of two parts: local modality intrinsic representation and local modality fusion representation. (Client) Local modal intrinsic representation The calculation method is as follows: , in, For client number, For category numbering, Modal numbering, for , It is the first The category number on each client is Modal numbering is The The inherent representation of individual modalities on each sample.

3. A film classification system based on multimodal deep representation collaborative federated learning according to claim 2, characterized in that, Client Local modality fusion representation The calculation method is as follows: , in, For client number, For category numbering, For the first The first client Local modality fusion representations across categories for , It is the first The category number on each client is The Individual modality fusion representation on a sample.

4. A film classification system based on multimodal deep representation collaborative federated learning according to claim 1, characterized in that, The modal intrinsic loss The calculation method is as follows: , in, For client number, For L2 norm distance loss, For category numbering, The total number of categories in the movie multimodal data. Modal numbering, The modality number of the movie's multimodal data. The local modal inherent characterization as described in claim 2, The first one issued by the server The first mode Global modality inherent representations across each category.

5. A film classification system based on multimodal deep representation collaborative federated learning according to claim 1, characterized in that, The modality fusion loss The calculation method is as follows: , in, For client number, The L2 norm distance loss function is used. For category numbering, The total number of categories in the movie multimodal data. The local modal fusion characterization as described in claim 3, The first one issued by the server Global modality fusion representation across categories.

6. A film classification system based on multimodal deep representation collaborative federated learning according to claim 1, characterized in that, The global multimodal depth representation calculated by the server in step S3 includes both the global modal intrinsic representation and the global modal fusion representation. Global Modal Intrinsic Representation The calculation method is as follows: , in, For client number, The total number of clients participating in federated training. This is an inherent representation of the local modality; Global modality fusion representation The calculation method is as follows: , in, For client number, The total number of clients participating in federated training. This is a local modality fusion representation.

7. A film classification system based on multimodal deep representation collaborative federated learning according to claim 1, characterized in that, The client and client Canonical correlation analysis distance matrix The calculation method is as follows: , in, For category numbering, and For the IDs of two different clients, The total number of categories in the movie multimodal data. This indicates the calculation of absolute value. and Clients respectively With the client Local modality fusion representation on and The results after canonical correlation analysis.

Citation Information

Patent Citations

  • Personalized federal learning method, device and system based on global feature sharing

    CN116777015A

  • Image classification method based on federal knowledge distillation and ensemble learning

    CN117523291A