Film and television resource recommendation method and device based on multi-modal data, equipment and storage medium
By performing feature extraction and model training on multimodal data, the problems of low recommendation accuracy and low personalization in the existing film and television content recommendation system are solved, and more accurate film and television resource recommendations are achieved.
Patent Information
- Application Number
- CN202510404095.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing film and television content recommendation systems, multimodal data is not fully utilized, resulting in low recommendation accuracy and low personalization.
By extracting feature of multimodal data, using preset embedded vector models and deep neural network models, the target resource recommendation model is trained to achieve accurate and personalized recommendation of film and television resources.
It improves the accuracy and personalization capabilities of the film and television content recommendation system, and can better meet users' viewing needs.
Smart Images

Figure CN120408135A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of resource recommendation, and particularly to a method, apparatus, device and storage medium for recommending film and television resources based on multimodal data. Background Art
[0002] In traditional film and television content recommendation systems, recommendations are usually based on users' historical behaviors or metadata of content (such as genre, actors, directors, etc.). However, these methods often ignore the rich information contained in multimodal data (such as images, videos, texts, audios, etc.), and this information can improve the accuracy and personalization degree of recommendations. Therefore, how to solve the problems of low recommendation accuracy and low personalization degree in existing film and television content recommendation systems has become an urgent problem to be solved. Summary of the Invention
[0003] The main purpose of this application is to provide a method, apparatus, device and storage medium for recommending film and television resources based on multimodal data, aiming to solve the technical problems of low recommendation accuracy and low personalization degree in existing film and television content recommendation systems.
[0004] To achieve the above object, this application proposes a method for recommending film and television resources based on multimodal data, and the method for recommending film and television resources based on multimodal data includes:
[0005] Extract features from the multimodal data to obtain target cross-modal features;
[0006] Input the target cross-modal features into a preset embedding vector model to obtain target embedding vector data;
[0007] Train an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model;
[0008] Perform film and television resource recommendation based on the target resource recommendation model to obtain a film and television resource recommendation result.
[0009] In addition, to achieve the above object, this application also proposes a device for recommending film and television resources based on multimodal data, and the device for recommending film and television resources based on multimodal data includes:
[0010] An extraction module, configured to extract features from the multimodal data to obtain target cross-modal features;
[0011] An embedding module, configured to input the target cross-modal features into a preset embedding vector model to obtain target embedding vector data;
[0012] A training module, configured to train an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model;
[0013] A recommendation module for recommending video resources based on the target resource recommendation model to obtain a video resource recommendation result.
[0014] In addition, to achieve the above object, the present application also proposes a video resource recommendation device based on multimodal data. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the video resource recommendation method based on multimodal data as described above.
[0015] In addition, to achieve the above object, the present application also proposes a storage medium. The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the video resource recommendation method based on multimodal data as described above.
[0016] In addition, to achieve the above object, the present application also provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the steps of the video resource recommendation method based on multimodal data as described above. Description of the Drawings
[0017] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the video resource recommendation method based on multimodal data of the present application;
[0020] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the video resource recommendation method based on multimodal data of the present application;
[0021] Figure 3 It is a schematic brief flowchart of the video resource recommendation method based on multimodal data provided for Embodiment 1 of the present application;
[0022] Figure 4 It is a schematic module structure diagram of the video resource recommendation device based on multimodal data according to the embodiment of the present application;
[0023] Figure 5Schematic diagram of the device structure of the hardware operating environment involved in the method for recommending film and television resources based on multi-modal data in the embodiments of the present application.
[0024] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0025] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0026] To better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings of the specification and specific implementation manners.
[0027] The main solution of the embodiments of the present application is: extracting features from multi-modal data to obtain target cross-modal features; inputting the target cross-modal features into a preset embedding vector model to obtain target embedding vector data; training an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model; and performing film and television resource recommendation based on the target resource recommendation model to obtain a film and television resource recommendation result.
[0028] In traditional film and television content recommendation systems, recommendations are usually based on users' historical behaviors or metadata of content (such as genre, actors, directors, etc.). However, these methods often ignore the rich information contained in multi-modal data (such as images, videos, texts, audios, etc.), and this information can improve the accuracy and personalization of recommendations. Therefore, how to solve the problems of low recommendation accuracy and low personalization in existing film and television content recommendation systems has become an urgent problem to be solved.
[0029] The present application extracts features from multi-modal data to obtain target cross-modal features; inputs the target cross-modal features into a preset embedding vector model to obtain target embedding vector data; trains an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model; and performs film and television resource recommendation based on the target resource recommendation model to obtain a film and television resource recommendation result.
[0030] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a film and television resource recommendation device based on multi-modal data that can implement the above functions. Hereinafter, taking the film and television resource recommendation device based on multi-modal data as the execution subject as an example, this embodiment and the following embodiments will be described.
[0031] Based on this, the embodiments of the present application provide a method for recommending film and television resources based on multi-modal data, referring to Figure 1 ,Figure 1 This is a schematic flowchart of the first embodiment of the method for recommending film and television resources based on multi-modal data in this application.
[0032] In this embodiment, the method for recommending film and television resources based on multi-modal data includes steps S10 to S40:
[0033] Step S10, extract features from the multi-modal data to obtain target cross-modal features;
[0034] It should be noted that in this embodiment, rich features are extracted from various media types such as images, videos, texts, and audios, and these features are effectively fused into a unified vector space. Then, through specific algorithms and model architectures, the alignment and integration of cross-modal features are realized to ensure that the fused features can accurately express the multi-dimensional characteristics of film and television content. Furthermore, the subsequent training of the recommendation model can be optimized using these fused multi-modal feature data.
[0035] It can be understood that the multi-modal data includes but is not limited to data such as image frames, video clips, text plot descriptions, and audio clips, and the target cross-modal features include but are not limited to image and video features, text features, and audio features.
[0036] In specific implementation, static image frames and dynamic video clips are extracted from a film or television drama to obtain image frames and video clips. Then, text plot descriptions are obtained according to text data such as the plot summary, reviews, and subtitles of the film or television drama, and audio clips are obtained according to audio dialogues, background music, etc. in the film or television drama. Furthermore, features are extracted from the image frames, video clips, text plot descriptions, and audio clips to obtain image and video features, text features, and audio features. Among them, the image and video features include high-level semantic features and low-level visual features of the image and video, the text features include semantic features and sentiment analysis features, and the audio features include spectrogram features and frequency spectrum features.
[0037] Step S20, input the target cross-modal features into a preset embedding vector model to obtain target embedding vector data;
[0038] It can be understood that the preset embedding vector model refers to an embedding vector model used to map the features of different modalities to a unified low-dimensional space, and the target embedding vector data refers to the representation data of the input data of each modality in the embedding layer. Among them, the embedding vector model adopts a two-tower neural network architecture, with each modality having an independent encoder branch, and finally, feature alignment and fusion are performed through a shared embedding layer. For example: Text modality: Use the BERT model to convert text data into 768-dimensional feature vectors; Image modality: Use the ResNet model to convert image data into 1024-dimensional feature vectors; Audio modality: Use the AST model to convert audio data into 128-dimensional feature vectors; Shared embedding layer: Map the features of the above modalities to a unified 512-dimensional space and perform L2 normalization to ensure that the features of different modalities are aligned on the same scale.
[0039] In a specific implementation, in the deep neural network model for processing multimodal data in this embodiment, a shared embedding layer is created, which maps the features of different modalities to a unified low-dimensional space. By defining a general embedding vector, it is ensured that the input data of each modality can be effectively represented in this layer, that is, the target embedding vector data is obtained.
[0040] Step S30, train the initial resource recommendation model according to the target embedding vector data to obtain the target resource recommendation model;
[0041] It can be understood that the initial resource recommendation model refers to an untrained deep neural network model for processing multimodal data, and the target resource recommendation model refers to a trained deep neural network model for processing multimodal data. Among them, the initial resource recommendation model is a deep neural network model suitable for processing multimodal data, and specifically adopts a multimodal cross-attention fusion architecture. This model includes the following parts: Input branch: Each modality has an independent encoder branch, such as BERT for processing text, ResNet for processing images, and AST for processing audio; Fusion module: Through the cross-attention mechanism, interact the text features with the visual and audio features to ensure effective fusion of the features of different modalities; Output layer: Use a gated fusion unit (GFU) to dynamically weight the features of each modality to generate the final fused features.
[0042] In a specific implementation, train the untrained deep neural network model for processing multimodal data through the representation data of the input data of each modality in the embedding layer, and combine an optimizer and a loss function during the training process. Finally, obtain the trained deep neural network model for processing multimodal data, that is, the target resource recommendation model.
[0043] In a feasible implementation manner, step S30 may include steps A11 to A13:
[0044] Step A11: Divide the target embedded vector data into batches based on a preset batch division strategy to obtain at least one batch of model training data;
[0045] It should be noted that the preset batch division strategy refers to the strategy of dividing the target embedded vector data into multiple batches of model training data. For example, assume there are 16 samples (i.e., the target embedded vector data, numbered from 1 to 16), and the batch size is 4. After random shuffling, the data may become: sample numbers: 10, 3, 15, 7, 1, 12, 6, 14, 9, 4, 11, 2, 8, 16, 5, 13. According to the batch size of 4, the data is divided into multiple batches: the first batch: 10, 3, 15, 7; the second batch: 1, 12, 6, 14; the third batch: 9, 4, 11, 2; the fourth batch: 8, 16, 5, 13. The model training data refers to the embedded vector data used for model training.
[0046] In specific implementation, this embodiment inputs data in batches (batch) to accelerate the training process and improve the stability of model convergence. Then, according to the strategy of dividing the target embedded vector data into multiple batches of model training data, batch division processing is performed on the representation data of the input data of each modality in the embedding layer. Finally, at least one batch of embedded vector data for model training, that is, at least one batch of model training data, is obtained.
[0047] Step A12: Train the initial resource recommendation model according to at least one batch of model training data and a cross-modal loss function to obtain a trained resource recommendation model, where the cross-modal loss function is: Ltotal = αLintra + βLinter + γLcontrast + λLreg, Lintra is the intra-modal consistency loss function, Linter is the inter-modal alignment loss function, Lcontrast is the cross-modal contrast loss function, Lreg is the regularization term function, α is the weight for balancing the intra-modal consistency loss function, β is the weight for balancing the inter-modal alignment loss function, γ is the weight for balancing the cross-modal contrast loss function, and λ is the regularization coefficient;
[0048] It can be understood that the cross-modal loss function refers to the loss function used for cross-modal model training. The cross-modal loss function is determined based on cosine similarity, triplet loss function, etc. The trained resource recommendation model refers to the initial resource recommendation model after training.
[0049] In the specific implementation, the untrained deep neural network model for processing multimodal data is trained based on the embedded vector data for model training after batch division and combined with the loss function for cross-modal model training. During the training process, the embedded vector data for model training is input in batches to speed up the training process and improve the stability of model convergence. Finally, the trained initial resource recommendation model, i.e., the training resource recommendation model, is obtained.
[0050] It should be noted that in order to fuse multimodal data such as text, speech, and images into the same low-dimensional vector space, this embodiment combines intra-modal consistency, inter-modal alignment, and cross-modal contrast learning to obtain a multimodal fusion loss function: Ltotal = αLintra + βLinter + γLcontrast + λLreg. Specifically, the intra-modal consistency loss (Lintra) is used to ensure that similar samples within the same modality are close in the embedding space. Assuming that the number of modalities is MM (such as M = 3M = 3 corresponds to text, speech, and image), a triplet loss is used for each modality mm: Lintra = M1m = 1∑M(a, p, n)∑[∥ema-emp∥22-∥ema-emn∥22+margin], where ema, emp, emnema, emp, emn are the anchor points, positive samples, and negative sample embeddings of modality mm, and margin is a hyperparameter, [·] + = max(0, ·) [·] + = max(0, ·);
[0051] The inter-modality alignment loss (Linter) is used to enforce alignment of embeddings of the same instance in different modalities. For each instance ii, the cosine similarity between all modality pairs is calculated:
[0052] Linter=-C1i=1∑B(m,n)∈P∑∥emi∥∥eni∥emi·eni
[0053] Where P is the combination of all modality pairs (such as text-speech, text-image, speech-image), B is the batch size, and C = B × |P| is the normalization factor;
[0054] The cross-modal contrast loss (Lcontrast) is used to bring positive pairs closer and push negative pairs apart based on the InfoNCE loss: Where sim(u, v) = u·v / (∥u∥∥v∥)sim(u, v) = u·v / (∥u∥∥v∥) is the cosine similarity, ττ is the temperature parameter, and the denominator contains multimodal positive pairs of the same sample and negative pairs of different samples;
[0055] The regularization term (LregLreg) is used to prevent overfitting and constrain the embedding vector and model parameters: Lreg = B1i = 1∑B∥ei∥22+m = 1∑M∥θm∥22, where θm is the parameter of each modal encoder;
[0056] Among them, α, β, γ: weights for balancing intra-modal, inter-modal and contrastive losses (such as α = 1.0, β = 0.5, γ = 0.1); λ: regularization coefficient (such as λ = 1e-4); margin: margin of triplet loss (such as 0.5); τ: temperature parameter (such as 0.07).
[0057] It should be understood that the model in this embodiment adopts a multi-branch structure, with each modality (text, image, audio) having an independent encoder branch, and finally performing feature alignment and fusion through a shared embedding layer. The specific structure is as follows: Text branch: Using the BERT model, outputting the text feature vector h txt ; Image branch: Use the ResNet model to output the image feature vector h img ; Audio branch: Use AST model to output audio feature vector h aud ; Shared embedding layer: maps each modal feature to a unified low-dimensional space and outputs the fusion feature h fuse .
[0058] Among them, the data source of the loss function Lintra is the output of the state branch (h txt , h img , h aud ) to align features within the same modality; the data source of the loss function Linter is the output of the shared embedding layer (h fuse ) to align the same instance features of different modalities; the data source of the loss function Lcontrast is the output of the shared embedding layer (h fuse ) to bring the cross-modal positive sample pairs closer and push the negative sample pairs apart; the data source of the loss function Lreg is the parameters of each modal branch and the shared embedding layer to prevent overfitting and constrain the model parameters.
[0059] Step A13: reversely optimize the training resource recommendation model to obtain a target resource recommendation model.
[0060] It is understandable that the model optimization strategy refers to a strategy for pre-setting the optimization of the trained model, such as a strategy for optimizing model parameters through the back propagation algorithm during the training process.
[0061] In a specific implementation, the optimization of the model in this embodiment needs to be carried out under a sufficient number of epochs (the process in which the entire training data set is completely propagated forward and backward through the neural network once), and then determine whether the model converges to a satisfactory performance level through the number of model optimizations, that is, optimize the initial resource recommendation model after training through the backpropagation algorithm, and when the number of model optimizations meets the requirements, obtain the target resource recommendation model.
[0062] In a feasible implementation manner, step A13 may include steps B11 to B13:
[0063] Step B11, perform backward optimization on the training resource recommendation model to obtain the current model optimization times and the optimized training resource recommendation model;
[0064] It can be understood that the current model optimization times refer to the number of times of optimizing the model parameters currently.
[0065] In a specific implementation, based on the strategy of optimizing the trained model preset in advance, optimize the initial resource recommendation model after training, that is, optimize the model parameters through the backpropagation algorithm, and then determine the number of times of optimizing the model parameters currently and the optimized training resource recommendation model.
[0066] Step B12, when the current model optimization times are greater than or equal to the optimization times threshold, evaluate the optimized training resource recommendation model according to the model verification data to determine the model evaluation result;
[0067] It can be understood that the optimization times threshold refers to the critical value of the optimization times used to judge whether the model performance level meets the requirements, and the model evaluation result refers to the result of evaluating the accuracy and efficiency of the model.
[0068] In a specific implementation, when the number of times of optimizing the model parameters currently is greater than or equal to the critical value of the optimization times used to judge whether the model performance level meets the requirements, it indicates that the model converges to a satisfactory performance level, and then conduct a test evaluation on the model, that is, use the validation set to verify the performance of the model in the cross-modal task, input the validation set into the optimized training resource recommendation model to obtain the film and television resources output by the model, and then compare them with the film and television data that the user has watched historically, obtain the similarity between the film and television resources output by the model and the film and television data that the user has watched historically, and compare it with the preset similarity threshold, and finally determine the result of evaluating the accuracy and efficiency of the model according to the comparison result to determine the model evaluation result.
[0069] In a feasible implementation manner, step B12 may include steps C11 to C13:
[0070] Step C11: When the current number of model optimization times is greater than or equal to the optimization times threshold, input model verification data into the optimized training resource recommendation model to obtain target test recommendation data;
[0071] It can be understood that model verification data refers to the test data used to verify the accuracy and efficiency of the model, and target test recommendation data refers to the movie and TV recommendation data generated based on the model verification data.
[0072] In specific implementation, when the current number of times of optimizing model parameters is greater than or equal to the optimization times critical value used to judge whether the model performance level meets the requirements, it indicates that the model converges to a satisfactory performance level. Input the test data used to verify the accuracy and efficiency of the model into the optimized training resource recommendation model for model testing to obtain the movie and TV recommendation data generated based on the model verification data output by the model, that is, target test recommendation data.
[0073] Step C12: Compare the target test recommendation data with historical movie and TV data to obtain a target comparison result;
[0074] It can be understood that historical movie and TV data refers to the movie and TV data that users have watched historically. Historical movie and TV data includes images, videos, texts, audios, etc. The target comparison result refers to the similarity comparison result between the target test recommendation data and the historical movie and TV data.
[0075] In specific implementation, compare the movie and TV recommendation data generated based on the model verification data with the historical movie and TV data that users have watched historically, that is, compare the data such as images, videos, texts, and audios in the movie and TV recommendation data generated based on the model verification data and the historical movie and TV data that users have watched historically. Then determine the similarity between the movie and TV recommendation data generated based on the model verification data and the historical movie and TV data that users have watched historically. Finally, obtain the target comparison result. For example, use a convolutional neural network (CNN) to extract the image feature vectors in the generated movie and TV recommendation data and the historical movie and TV data that users have watched historically, and calculate the cosine similarity of the two feature vectors to obtain the image similarity; decompose the videos in the generated movie and TV recommendation data and the historical movie and TV data that users have watched historically into video frames, extract features for each frame, and then integrate the features through a time series model (such as LSTM) to obtain video feature vectors respectively. Finally, according to the weight coefficients of the image feature vectors and the video feature vectors, comprehensively calculate the overall similarity.
[0076] Step C13: When the target comparison result is a data similarity result, determine that the model evaluation result is a passed evaluation result.
[0077] It can be understood that the data similarity result refers to the comparison result where the similarity between the generated movie and TV drama recommendation data and the movie and TV drama data that the user has watched historically meets the requirements. The evaluation passed result refers to the evaluation result where the performance of the optimized training resource recommendation model meets the requirements. In this embodiment, a validation set can be used to verify the performance of the model in cross-modal tasks.
[0078] In a specific implementation, when the target comparison result is a comparison result where the similarity between the generated movie and TV drama recommendation data and the movie and TV drama data that the user has watched historically meets the requirements, it indicates that the performance of the model meets the requirements. Furthermore, it is determined that the model evaluation result is an evaluation passed result, that is, there is no need to adjust the model architecture or optimize the hyperparameters. On the contrary, when the target comparison result is a comparison result where the similarity between the generated movie and TV drama recommendation data and the movie and TV drama data that the user has watched historically does not meet the requirements, it indicates that the performance of the model does not meet the requirements, that is, it is necessary to adjust the model architecture or optimize the hyperparameters.
[0079] Step B13, when the model evaluation result is an evaluation passed result, determine the optimized training resource recommendation model as the target resource recommendation model.
[0080] It can be understood that when the model evaluation result is an evaluation passed result, it indicates that the accuracy and efficiency of the optimized training resource recommendation model meet the requirements, and there is no need to adjust the model architecture or optimize the hyperparameters. Furthermore, determine the optimized training resource recommendation model as the target resource recommendation model.
[0081] Step S40, perform movie and TV drama resource recommendation based on the target resource recommendation model to obtain a movie and TV drama resource recommendation result.
[0082] In a specific implementation, the movie and TV drama resource recommendation result refers to the result of recommending movie and TV drama resources based on the fused multi-modal data. By using a trained deep neural network model for processing multi-modal data and combining the user's multi-modal data for movie and TV drama resource recommendation, the movie and TV drama resource recommendation result is determined.
[0083] It should be noted that this embodiment adopts cross-modal embedding learning technology to ensure that features from different modalities can be aligned and fused in a unified feature space. The main steps are as follows: Design a deep neural network model suitable for processing multi-modal data, and define that the data of each modality enters the model through different input branches.
[0084] Create a shared embedding layer in the model, and this layer maps the features of different modalities to a unified low-dimensional space. By defining a general embedding vector, it is ensured that the input data of each modality can be effectively represented in this layer.
[0085] To achieve the alignment of different modal data features in the low-dimensional embedding space, mainly a cross-modal loss function is designed. This loss function can be based on cosine similarity, triplet loss function, etc., to ensure that similar data samples are closer in the embedding space.
[0086] Compile the designed model and select appropriate optimizers and loss functions. Use the Adam optimizer and cooperate with the previously defined cross-modal loss function to train the model.
[0087] Input the preprocessed data into the model for training. The data can be input in batches to accelerate the training process and improve the stability of model convergence.
[0088] During the training process, optimize the model parameters through the backpropagation algorithm so that the features of different modalities can be gradually aligned and integrated in the shared low-dimensional space. This process needs to be carried out under a sufficient number of epochs until the model converges to a satisfactory performance level.
[0089] Use the trained model to map the features of different modal data into the shared low-dimensional space. This means that the representation forms of each modal data will correspond to similar positions in the same embedding space, making the relationships between cross-modal data clearer and more comparable.
[0090] Evaluate the trained model. The validation set can be used to verify the performance of the model in cross-modal tasks. According to the evaluation results, it may be necessary to adjust the model architecture or optimize the hyperparameters to further improve the accuracy and efficiency of the model.
[0091] Finally, apply the fused multi-modal data to the film and television content recommendation system. These data can not only increase the depth of understanding of users' interests and preferences by the recommendation system, but also provide more accurate and personalized recommended content.
[0092] In this embodiment, target cross-modal features are obtained by extracting features from multi-modal data; inputting the target cross-modal features into a preset embedding vector model to obtain target embedding vector data; training an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model; and performing film and television resource recommendation based on the target resource recommendation model to obtain film and television resource recommendation results. Through advanced deep learning technologies, features from multiple media types (such as images, videos, texts, audios, etc.) are effectively integrated to improve the accuracy and personalization ability of the film and television content recommendation system. Specific algorithms and model architectures are designed for the alignment and integration of cross-modal features to ensure that the fused features can accurately reflect the multi-dimensional characteristics of film and television content, thereby better meeting the personalized viewing needs of users.
[0093] Based on the first embodiment of this application, in the second embodiment of this application, for the same or similar content as in the above-mentioned Embodiment 1, reference can be made to the above introduction, and it will not be repeated hereinafter. On this basis, please refer to Figure 2 In the method for recommending film and television resources based on multimodal data, step S10 of the method further includes steps S11 to S13:
[0094] Step S11, determining image data, text data, and audio data according to the multimodal data;
[0095] It can be understood that the image data refers to static image frames and dynamic video segments extracted from a film and television drama, the text data refers to text data such as the plot summary, reviews, and subtitles of the film and television drama, and the audio data refers to audio dialogues, background music, etc. in the film and television drama.
[0096] In a specific implementation, this embodiment obtains multimodal data from various data sources, specifically static image frames and dynamic video segments extracted from a film and television drama, extracts text data such as the plot summary, reviews, and subtitles of the film and television drama, and extracts audio dialogues, background music, etc. in the film and television drama, that is, image data, text data, and audio data.
[0097] Step S12, preprocessing the image data, the text data, and the audio data to obtain preprocessed image data, text data, and audio data;
[0098] It can be understood that this embodiment preprocesses each data type to ensure the accuracy and consistency of the data in the subsequent feature extraction stage, that is, normalizing the image pixel values to a specific range to eliminate the differences between different images, splitting the text into words or phrases, removing common but meaningless stop words, performing spectral analysis on the audio data, and finally obtaining preprocessed image data, text data, and audio data.
[0099] In a feasible implementation manner, step S02 may include steps D11 to D13:
[0100] Step D11, extracting image frames from the image data to obtain preprocessed image data;
[0101] It can be understood that extracting image frames from the image data, extracting key frames from the video as representative images, then normalizing the image pixel values to a specific range to eliminate the differences between different images, and finally obtaining preprocessed image data.
[0102] Step D12, performing text segmentation on the text data based on a preset text processing strategy to obtain preprocessed text data;
[0103] It can be understood that the preset text processing strategy refers to the pre-set strategy for processing text data, such as splitting text into words or phrases.
[0104] In a specific implementation, the text is split into words or phrases, common but meaningless stop words are removed, and finally the preprocessed text data is obtained.
[0105] Step D13, preprocess the audio data based on a preset audio processing strategy to obtain preprocessed audio data.
[0106] It can be understood that the preset audio processing strategy refers to the pre-set strategy for processing audio data.
[0107] In a specific implementation, spectral analysis is performed on audio dialogues, background music, etc. in a movie or TV drama, spectral analysis is carried out, Mel spectrogram or MFCC features are extracted, the audio is represented as a fixed-length feature vector, and finally the preprocessed audio data is obtained.
[0108] It should be noted that in this embodiment, each type of data is preprocessed to ensure the accuracy and consistency of the data in the subsequent feature extraction stage. Images and videos: Frame extraction is performed, key frames are extracted from the video as representative images, and then the pixel values of the images are normalized to a specific range to eliminate differences between different images, and a convolutional neural network (CNN) is used to extract high-level visual features, representing the images as fixed-length feature vectors; Text: The text is split into words or phrases, common but meaningless stop words are removed, and finally technologies such as Word2Vec and GloVe are used to convert the text into word vectors, representing the text as fixed-length feature vectors; Audio: Spectral analysis is performed, Mel spectrogram or MFCC features are extracted, and the audio is represented as a fixed-length feature vector.
[0109] Step S13, respectively extract features from the preprocessed image data, text data, and audio data to obtain target cross-modal features.
[0110] It can be understood that image semantic features, image visual features, text semantic features, text sentiment features, spectrogram features, and spectral features are respectively extracted based on the preprocessed image data, text data, and audio data, and then the image semantic features, image visual features, text semantic features, text sentiment features, spectrogram features, and spectral features are summarized to obtain the target cross-modal features.
[0111] In a feasible implementation manner, step S03 may include steps E11 to E14:
[0112] Step E11: Extract features from the preprocessed image data based on a preset convolutional neural network to obtain image semantic features and image visual features;
[0113] It can be understood that the preset convolutional neural network refers to a pre-set convolutional neural network (CNN), the image semantic features refer to the high-level semantic features of the image video, and the image visual features refer to the low-level visual features of the image video.
[0114] In specific implementation, this embodiment uses convolutional neural network (CNN) and video processing technology to extract visual features from images and videos. Specifically, a pre-trained CNN model (such as ResNet or VGG) is used to extract high-level semantic features and low-level visual features of the image video, including color, texture, shape, etc.
[0115] Step E12: Extract text semantic features from the preprocessed text data based on a preset word vector model, and perform a topic structure analysis on the preprocessed text data to obtain text sentiment features;
[0116] It can be understood that the preset word vector model refers to a vector model used to convert text into high dimensions, the text semantic features refer to the semantic features corresponding to the text data, and the text sentiment features refer to sentiment analysis features.
[0117] In specific implementation, natural language processing (NLP) technology is used to process text data to extract semantic features and sentiment analysis features. Specifically, a word vector model (such as Word2Vec or BERT) is used to convert text into a high-dimensional vector representation to capture the semantic similarity and context information between words. In addition, topic modeling technology (such as Latent Dirichlet Allocation, LDA) is used to discover the latent topic structure in the text to help deeply understand the complexity and relevance of the document content.
[0118] Step E13: Perform acoustic analysis on the preprocessed audio data to obtain an acoustic analysis result, and obtain spectrogram features and spectral features based on the acoustic analysis result;
[0119] It can be understood that the acoustic analysis result refers to the result of analyzing the preprocessed audio data through acoustic analysis technology. The spectrogram features refer to a two-dimensional representation form obtained by converting the audio signal from the time domain to the time-frequency domain, and the spectral features refer to specific frequency-related feature information extracted from the audio data.
[0120] In a specific implementation, acoustic analysis technology is used to perform acoustic analysis on the processed audio data, and spectrograms and spectral features are extracted from the analysis results. These features can effectively express the tone and emotional content of the audio.
[0121] Step E14, summarize the features according to the image semantic features, the image visual features, the text semantic features, the text emotional features, the spectrogram features, and the spectral features to obtain target cross-modal features.
[0122] In a specific implementation, this embodiment summarizes and processes the high-level semantic features and low-level visual features of image videos, the semantic features and sentiment analysis features corresponding to text data, spectrogram features, and spectral features to obtain target cross-modal features.
[0123] In this embodiment, image data, text data, and audio data are determined according to multi-modal data; the image data, the text data, and the audio data are preprocessed to obtain preprocessed image data, text data, and audio data; feature extraction is respectively performed on the preprocessed image data, text data, and audio data to obtain target cross-modal features. By extracting rich features from various media types such as images, videos, texts, and audios and effectively integrating these features into a unified vector space, the accuracy and personalization ability of the film and television content recommendation system are improved.
[0124] Exemplarily, to help understand the implementation process of the film and television resource recommendation method based on multi-modal data obtained by combining the above Embodiment 1, please refer to Figure 3 , Figure 3 A brief flow schematic diagram of a film and television resource recommendation method based on multi-modal data is provided. Specifically: rich features are extracted from various media types such as images, videos, texts, and audios, and these features are effectively integrated into a unified vector space. Through specific algorithms and model architectures, the alignment and integration of cross-modal features are realized to ensure that the fused features can accurately express the multi-dimensional characteristics of film and television content. Furthermore, the subsequent multi-modal feature data after fusion can be used to optimize the training of the recommendation model.
[0125] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the film and television resource recommendation method based on multi-modal data of this application. Any simple transformations in more forms based on this technical concept are within the protection scope of this application.
[0126] This application also provides a film and television resource recommendation device based on multi-modal data. Please refer to Figure 4 , the film and television resource recommendation device based on multi-modal data includes:
[0127] An extraction module 10 for extracting features from multimodal data to obtain target cross-modal features;
[0128] An embedding module 20 for inputting the target cross-modal features into a preset embedding vector model to obtain target embedding vector data;
[0129] A training module 30 for training an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model;
[0130] A recommendation module 40 for recommending film and television resources based on the target resource recommendation model to obtain a film and television resource recommendation result.
[0131] Optionally, the training module 30 is further configured to:
[0132] Perform batch division on the target embedding vector data based on a preset batch division strategy to obtain at least one batch of model training data;
[0133] Train the initial resource recommendation model according to at least one batch of model training data and a cross-modal loss function to obtain a trained resource recommendation model, where the cross-modal loss function is: Ltotal = αLintra + βLinter + γLcontrast + λLreg, Lintra is an intra-modal consistency loss function, Linter is an inter-modal alignment loss function, Lcontrast is a cross-modal contrast loss function, Lreg is a regularization term function, α is the weight for balancing the intra-modal consistency loss function, β is the weight for balancing the inter-modal alignment loss function, γ is the weight for balancing the cross-modal contrast loss function, and λ is a regularization coefficient;
[0134] Perform reverse optimization on the trained resource recommendation model to obtain a target resource recommendation model.
[0135] Optionally, the training module 30 is further configured to:
[0136] Perform reverse optimization on the trained resource recommendation model to obtain the current model optimization times and the optimized trained resource recommendation model;
[0137] When the current model optimization times is greater than or equal to an optimization times threshold, evaluate the optimized trained resource recommendation model according to model verification data to determine a model evaluation result;
[0138] When the model evaluation result is a passed evaluation result, determine the optimized trained resource recommendation model as the target resource recommendation model.
[0139] Optionally, the training module 30 is further configured to:
[0140] When the current number of model optimization times is greater than or equal to the optimization times threshold, input model verification data into the optimized training resource recommendation model to obtain target test recommendation data;
[0141] Compare the target test recommendation data with historical film and television data to obtain a target comparison result;
[0142] When the target comparison result is a data similarity result, determine that the model evaluation result is a passed evaluation result.
[0143] Optionally, the extraction module 10 is further configured to:
[0144] Determine image data, text data, and audio data according to multimodal data;
[0145] Preprocess the image data, the text data, and the audio data to obtain preprocessed image data, text data, and audio data;
[0146] Extract features from the preprocessed image data, text data, and audio data respectively to obtain target cross-modal features.
[0147] Optionally, the extraction module 10 is further configured to:
[0148] Extract image frames from the image data to obtain preprocessed image data;
[0149] Segment the text data based on a preset text processing strategy to obtain preprocessed text data;
[0150] Preprocess the audio data based on a preset audio processing strategy to obtain preprocessed audio data.
[0151] Optionally, the extraction module 10 is further configured to:
[0152] Extract features from the preprocessed image data based on a preset convolutional neural network to obtain image semantic features and image visual features;
[0153] Extract text semantic features from the preprocessed text data based on a preset word vector model, and perform topic structure analysis on the preprocessed text data to obtain text sentiment features;
[0154] Perform acoustic analysis on the preprocessed audio data to obtain an acoustic analysis result, and obtain spectrogram features and frequency spectrum features according to the acoustic analysis result;
[0155] Summarize the features according to the image semantic features, the image visual features, the text semantic features, the text sentiment features, the spectrogram features, and the spectrum features to obtain target cross-modal features.
[0156] The film and television resource recommendation device based on multimodal data provided by this application adopts the film and television resource recommendation method based on multimodal data in the above embodiment, and can solve the technical problems of low recommendation accuracy and low personalization degree in the existing film and television content recommendation system. Compared with the prior art, the beneficial effects of the film and television resource recommendation device based on multimodal data provided by this application are the same as those of the film and television resource recommendation method based on multimodal data provided by the above embodiment, and other technical features in the film and television resource recommendation device based on multimodal data are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.
[0157] This application provides a film and television resource recommendation device based on multimodal data. The film and television resource recommendation device based on multimodal data includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the film and television resource recommendation method in Embodiment 1 above.
[0158] Next, refer to Figure 5 , which shows a schematic structural diagram of a film and television resource recommendation device suitable for implementing the embodiments of this application. The film and television resource recommendation device based on multimodal data in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The film and television resource recommendation device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of this application.
[0159] As Figure 5As shown, the film and television resource recommendation device based on multimodal data may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the film and television resource recommendation device based on multimodal data are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the film and television resource recommendation device based on multimodal data to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a film and television resource recommendation device with various systems based on multimodal data, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.
[0160] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0161] The film and television resource recommendation device based on multimodal data provided by this application adopts the film and television resource recommendation method based on multimodal data in the above-mentioned embodiment, and can solve the technical problems of low recommendation accuracy and low personalization degree in the existing film and television content recommendation system. Compared with the prior art, the beneficial effects of the film and television resource recommendation device based on multimodal data provided by this application are the same as those of the film and television resource recommendation method based on multimodal data provided by the above-mentioned embodiment, and other technical features in the film and television resource recommendation device based on multimodal data are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0162] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0163] As described above, only the specific implementation manners of this application are provided, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
[0164] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the film and television resource recommendation method based on multimodal data in the above-mentioned embodiment.
[0165] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0166] The above computer-readable storage medium can be included in a multi-modal data-based film and television resource recommendation device; or it can exist independently without being assembled into a multi-modal data-based film and television resource recommendation device.
[0167] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by a multi-modal data-based film and television resource recommendation device, the multi-modal data-based film and television resource recommendation device is caused to: extract features from the multi-modal data to obtain target cross-modal features; input the target cross-modal features into a preset embedding vector model to obtain target embedding vector data; train an initial resource recommendation model based on the target embedding vector data to obtain a target resource recommendation model; and perform film and television resource recommendation based on the target resource recommendation model to obtain a film and television resource recommendation result.
[0168] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0170] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0171] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned method for recommending video resources based on multimodal data, and can solve the technical problems of low recommendation accuracy and low personalization degree of existing video content recommendation systems. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for recommending video resources based on multimodal data provided in the above embodiments, and will not be elaborated here.
[0172] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned method for recommending film and television resources based on multimodal data.
[0173] The computer program product provided by the present application can solve the technical problems of low recommendation accuracy and low personalization degree in the existing film and television content recommendation system. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the method for recommending film and television resources based on multimodal data provided by the above embodiments, and will not be elaborated here.
[0174] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A film and television resource recommendation method based on multi-modal data, characterized in that The film and television resource recommendation method based on multimodal data includes: Performing feature extraction on multimodal data to obtain target cross-modal features; Inputting the target cross-modal features into a preset embedding vector model to obtain target embedding vector data; Training an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model; Performing film and television resource recommendation based on the target resource recommendation model to obtain a film and television resource recommendation result.
2. The method according to claim 1, wherein The step of training the initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model includes: Performing batch division on the target embedding vector data based on a preset batch division strategy to obtain at least one batch of model training data; Training the initial resource recommendation model according to at least one batch of model training data and a cross-modal loss function to obtain a trained resource recommendation model, where the cross-modal loss function is: Ltotal = αLintra + βLinter + γLcontrast + λLreg, Lintra is an intra-modal consistency loss function, Linter is an inter-modal alignment loss function, Lcontrast is a cross-modal contrast loss function, Lreg is a regularization term function, α is the weight for balancing the intra-modal consistency loss function, β is the weight for balancing the inter-modal alignment loss function, γ is the weight for balancing the cross-modal contrast loss function, and λ is a regularization coefficient; Performing reverse optimization on the trained resource recommendation model to obtain a target resource recommendation model.
3. The method according to claim 2, wherein The step of performing reverse optimization on the trained resource recommendation model to obtain a target resource recommendation model includes: Performing reverse optimization on the trained resource recommendation model to obtain the current model optimization times and the optimized trained resource recommendation model; When the current model optimization times is greater than or equal to an optimization times threshold, evaluating the optimized trained resource recommendation model according to model verification data to determine a model evaluation result; When the model evaluation result is a passed evaluation result, determining the optimized trained resource recommendation model as the target resource recommendation model.
4. The method according to claim 3, wherein The step of evaluating the optimized trained resource recommendation model according to model verification data to determine a model evaluation result when the current model optimization times is greater than or equal to an optimization times threshold includes: When the current model optimization times is greater than or equal to an optimization times threshold, inputting the model verification data into the optimized trained resource recommendation model to obtain target test recommendation data; Comparing the target test recommendation data with historical film and television data to obtain a target comparison result; When the target comparison result is a data similarity result, determining the model evaluation result as a passed evaluation result.
5. The method according to claim 1, characterized in that, The step of performing feature extraction on multimodal data to obtain target cross-modal features includes: Determining image data, text data, and audio data according to multimodal data; Performing preprocessing on the image data, the text data, and the audio data to obtain preprocessed image data, text data, and audio data; Extract features from the preprocessed image data, text data, and audio data respectively to obtain target cross-modal features.
6. The method according to claim 5, characterized in that, The steps of preprocessing the image data, the text data, and the audio data to obtain preprocessed image data, text data, and audio data include: Extract image frames from the image data to obtain preprocessed image data; Segment the text data based on a preset text processing strategy to obtain preprocessed text data; Preprocess the audio data based on a preset audio processing strategy to obtain preprocessed audio data.
7. The method according to claim 5, characterized in that The steps of extracting features from the preprocessed image data, text data, and audio data respectively to obtain target cross-modal features include: Extract features from the preprocessed image data based on a preset convolutional neural network to obtain image semantic features and image visual features; Extract text semantic features from the preprocessed text data based on a preset word vector model, and perform topic structure analysis on the preprocessed text data to obtain text sentiment features; Perform acoustic analysis on the preprocessed audio data to obtain an acoustic analysis result, and obtain spectrogram features and frequency spectrum features according to the acoustic analysis result; Summarize the features according to the image semantic features, the image visual features, the text semantic features, the text sentiment features, the spectrogram features, and the frequency spectrum features to obtain target cross-modal features.
8. A film and television resource recommendation device based on multimodal data, characterized in that, The device includes: An extraction module for extracting features from multi-modal data to obtain target cross-modal features; An embedding module for inputting the target cross-modal features into a preset embedding vector model to obtain target embedding vector data; A training module for training an initial resource recommendation model according to the target embedding vector data to obtain a target resource recommendation model; A recommendation module for recommending movie and television resources based on the target resource recommendation model to obtain movie and television resource recommendation results.
9. A film and television resource recommendation device based on multimodal data, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the movie and television resource recommendation method based on multi-modal data according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the movie and television resource recommendation method based on multi-modal data according to any one of claims 1 to 7.