User generated content popularity prediction method based on retrieval enhancement prediction
By constructing a UGC knowledge vector library and a dual-path retrieval enhanced prediction method, the data and model coupling problem in UGC popularity prediction is solved, a more efficient and accurate UGC popularity prediction is achieved, and the versatility and robustness of the model are improved.
Patent Information
- Application Number
- CN202511270681.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing UGC popularity prediction methods find it difficult to fully model the complex semantic associations and external knowledge dependencies of UGC at the data level. Deep learning network training is difficult to capture the implicit cross-modal semantic associations in UGC content. There are also defects such as retrieval noise interference, inaccurate long text embedding and strong model coupling, which lead to prediction performance bottlenecks.
A retrieval-augmented prediction (RAP)-based method is adopted to build a UGC knowledge vector library, design a dual-path retrieval and rearrangement recall architecture, combine it with a deep learning model, and dynamically adjust the feature fusion weights to reduce the model's dependence on the retrieval library and enhance the model's versatility and generalization capabilities.
It effectively solves the performance bottleneck of existing prediction methods, improves the accuracy and robustness of UGC popularity prediction, reduces long text representation deviation and retrieval noise interference, adapts to the dynamic update of the retrieval library, and improves the versatility and generalization ability of the model.
Smart Images

Figure CN120763406A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information dissemination prediction, and in particular relates to a method for predicting the popularity of user-generated content based on retrieval-enhanced prediction. Background Art
[0002] In the rapidly developing digital age, continuous advancements in internet and mobile communication technologies have driven unprecedented growth in the production and dissemination of videos. User-generated content (UGC), a crucial way for people to access information and entertainment online, has become a key research area, with its causes and trends becoming increasingly popular. This research not only helps content producers enrich and innovate video content, but also helps internet platforms optimize resource allocation, enhance user experience, and ultimately drive the overall development of industry technology. UGC is characterized by its vast volume and diverse forms, encompassing multimodal data such as text, images, video, and audio. The explosive growth and complex nature of this data make accurately predicting its popularity a key technical challenge for business scenarios such as content recommendation, advertising placement, and public opinion monitoring.
[0003] Currently, existing methods for predicting the popularity of user-generated content (UGC) fall into the following categories: Traditional machine learning methods: As an important foundational tool, researchers have personalized and improved classic models to adapt them to prediction tasks. Common methods include linear regression, support vector machines (SVMs), decision trees, and random forests (RFs). Multimodal methods: Integrate multiple sources of information, including visual, audio, and textual descriptions of videos, user interactions, and social network relationships. Combining heterogeneous data sources through early and late fusion methods improves prediction accuracy and robustness. Variational autoencoders (VAEs): As generative models, they capture the latent representations of videos and user behaviors by learning probability distributions from data, helping to understand the complex dynamics of video dissemination. Knowledge graphs (KGs): Integrate video content, user behavior, and other information using structured semantic networks, mining deep data relationships through relational reasoning and path deduction. Graph neural networks (GNNs): Leveraging their ability to process graph-structured data, they can effectively represent and capture entity relationships. Entities such as videos and users can be easily embedded in graph structures. Heterogeneous graphs: Using different types of nodes and edges, they make entity connections clearer and relationships more complex, improving video popularity prediction. Information cascade prediction: The scale and popularity of dissemination are obtained by analyzing the information cascade process. The cascade process can represent the dissemination network formed by the video in the network through user behavior, and characterize the dynamics of information dissemination. Existing UGC popularity prediction methods have many limitations: at the data level, due to the limitations of the quality of UGC data and the explosive growth of long semantic information, it is difficult to fully model the complex semantic associations and external knowledge dependencies of UGC, resulting in a bottleneck in prediction performance. At the deep learning network training level, it is difficult to fully capture the cross-modal semantic associations implicit in UGC content, including the deep interactive relationships between multimodal features such as vision, hearing, and text. Even in the few methods that use retrieval libraries to enhance semantic understanding and UGC prediction tasks, there are the following technical limitations: Retrieval noise interference: If the vector retrieval database is of low quality, the Top-K information will be irrelevant to the target UGC content, which will introduce strong noise into the prediction model. Inaccurate embedding of long texts: Traditional single-path retrieval (dense vector similarity retrieval) often fails to accurately express semantic information when embedding long texts, reducing prediction reliability. Coupling defects: The retrieval module is strongly coupled with the prediction model, making it difficult for the model to adapt to dynamic updates of the vector retrieval library. It is easy to become overly dependent on the retrieval library, and it is unable to maintain basic prediction capabilities when the retrieval fails, thus losing its versatility and generalization capabilities.
[0004] These limitations collectively lead to performance bottlenecks in prediction models, making it difficult for traditional prediction methods based on a single deep learning network architecture to adapt. Therefore, this paper proposes a method for predicting the popularity of user-generated content based on Retrieval Augmented Prediction (RAP). Summary of the Invention
[0005] The purpose of the present invention is to provide a method for predicting the popularity of user-generated content based on retrieval-enhanced prediction, aiming to solve the problems raised in the above background technology.
[0006] The purpose of the present invention is achieved through the following technical solutions: A method for predicting popularity of user-generated content based on retrieval-enhanced prediction comprises the following steps: Step 1: Data collection; Collect multimodal data on UGC websites, including video streams, titles, comment content, covers, and popularity behavior indicators such as likes, views, and comments; Step 2: Data preprocessing; The collected raw data is cleaned, duplicates are removed, modality loss is processed, and text and images are standardized. The samples are then divided into a UGC knowledge vector library set, a training set, and a test set. Step 3: Feature construction; Extract three types of modal features: video, text, and image, and generate multimodal embedding representations through modality amplification and embedding transformation; Step 4: Construction of UGC knowledge vector library; For text modality, all text features are concatenated and long text embedding is generated to form the UCG knowledge vector library; Step 5: Two-way retrieval and rearrangement recall; For the text embedding of UGC samples, we perform dense vector similarity search and word frequency-based BM25 search in the UGC knowledge vector library. We then use the RFF algorithm to dual-channel fusion recall of the Top-K results to obtain the final recall set: Step 6: Retrieve the enhanced prediction results to assist downstream network UGC popularity prediction; Combined with a deep learning model, model training and popularity prediction are performed based on the input multimodal features and recall sets, and a mechanism is provided in the downstream network to dynamically adjust the feature fusion weights according to the confidence of the retrieval results.
[0007] Furthermore, in step 3, modality expansion includes: Use the llava model to process the video cover image to text and generate text modal features: ; in Indicates the The cover features of the sample are amplified by the llava model to generate text modal features; is the total number of words in the current sample, It is The first sample token, denotes the position of the word in the current sample is the th; For the video modality, take 10 frames at equal intervals and convert them into image modalities; Convert text and image modality information into embeddings through the jina-clip-v2 model: ; ; ; ; ; where , , , , are the embeddings of , , , , , denotes the video modality information of the th sample, denotes the title information of the th sample, denotes the comment information of the th sample, denotes the cover information of the th sample; is a text feature embedding function based on the jina-clip-v2 model; is an image feature embedding function based on the jina-clip-v2 model.
[0008] Further, in step 4, the construction of the UGC knowledge vector library includes: First, concatenate the text features of the title, cover augmented text, and comments, and then use the jina-clip-v2 model to generate long text embeddings. The relevant formula is as follows: ; ; where is the text embedding of the UGC sample; represents concatenation; is the text embedding in the UGC knowledge vector library after concatenation; , , correspond to the The title text features of each sample, the text features after modal amplification of the cover features, and the comment text features.
[0009] Furthermore, in step 5, dense vector similarity retrieval is achieved by calculating cosine similarity, and the formula is as follows: ; in Text embedding representing UGC samples and the first j vectors The cosine similarity of is the total number of documents in the UGC knowledge vector library; After sorting in descending order by cosine similarity, recall the Top-J results and record them as a set .
[0010] Furthermore, in step 5, the calculation formula for the relevance score of the BM25 search based on word frequency is as follows: ; ; in Score for BM25; is the first The text content of a document; For query The terms in is the inverse document frequency; For terms exist The frequency in and To adjust the parameters; For Documents length; is the average length of all documents in the UCG knowledge vector library; To include terms The number of documents; After sorting in descending order by BM25 score, recall the Top-J results and record them as a set .
[0011] Furthermore, in step 5, the fusion score of the RFF algorithm is: ; in Score for fusion; is the first k vectors; For the search method, For dense vector similarity retrieval, It is BM25 search based on word frequency; is the smoothing hyperparameter; for Ranking in the corresponding search results: for Similarity retrieval results in dense vectors Ranking in for Corresponding documents Results in BM25 Ranking in according to After sorting in descending order, take the top-K with the highest scores , and get the final recall set: ; in is the recall set, i.e., the RAP result; are the first K vectors sorted by score from high to low.
[0012] Furthermore, in step 6, the supervised learning paradigm is adopted in the model training phase, AdamW is selected as the optimizer, and a dynamic learning rate decay strategy is set; the input of the model is the feature representation after multimodal fusion, and the output is a continuous popularity score. The training goal is to minimize the error between the predicted value and the true value.
[0013] Furthermore, the prediction performance of step 6 was evaluated by normalized mean square error and Spearman rank correlation coefficient.
[0014] Compared with the prior art, the present invention has the following beneficial effects: This paper effectively addresses the performance bottleneck of existing UGC popularity prediction by constructing a UGC knowledge vector library, designing a retrieval-augmented prediction (RAP) method, and employing a dual-path retrieval, high-speed recall, and retrieval content decoupling and noise reduction architecture. The dual-path retrieval and rearranged recall architecture address the problem of representation bias in long texts. A decoupling noise reduction approach in the downstream deep learning network module, which dynamically adjusts feature fusion weights based on confidence, addresses retrieval noise interference. The overall architectural design overcomes the drawback of strong model coupling, enabling the model to adapt to dynamic updates to the retrieval library and improving its versatility and generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 Flow chart of the method of the present invention.
[0016] Figure 2 Build a schematic diagram for the feature.
[0017] Figure 3 Schematic diagram of dual-path retrieval and rearrangement recall.
[0018] Figure 4 Flowchart for retrieval enhanced prediction results (RAP results) to assist downstream network UGC popularity prediction. DETAILED DESCRIPTION
[0019] In order to have a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention is now described in detail below, but it should not be understood as limiting the scope of implementation of the present invention.
[0020] The present invention provides a method for predicting the popularity of user-generated content based on retrieval-enhanced prediction, the flow chart of which is as follows: Figure 1 As shown, the method includes the following steps: Step 1: Data collection; Carry out systematic data collection on UGC websites, regularly capture relevant information about UGC content, and the collection scope covers video content and its multimodal characteristics. Specifically, it includes the following dimensions: UGC content video stream: used to extract visual features; UGC content video titles: usually contain keywords, hot topics, etc., and have high semantic value; UGC content comments: reflect user interaction and are rich in audience emotions and attitudes; UGC content video cover: rich in visual information, has a great impact on click-through rate; The number of likes, views, and comments on UGC content: As core behavioral indicators of the popularity of UGC content, they serve as obvious supervisory signals.
[0021] Step 2: Data preprocessing; In order to improve the effectiveness and generalization ability of model training, the original data must be preprocessed first. The specific process is as follows: Data cleaning: remove records with incorrect formats, unify data encoding formats (such as UTF-8), field names and units; Duplicate removal: remove duplicate content to avoid bias in the model; Modality missing processing: remove samples with missing modalities (such as no title, no cover, no comments, etc.) to reduce noise interference; Text and image standardization: perform word segmentation and stop word removal on text fields.
[0022] After data cleaning and standardization, the samples are divided into three parts according to the three-stage requirements of UCG knowledge vector library construction, training, and testing: UGC knowledge vector library collection (50%): serves as the basic feature library for matching and retrieval, supporting RAP retrieval calculations; Training set (40%): used for model learning and parameter optimization; Test set (10%): used for model performance evaluation.
[0023] The overall data partitioning ensures randomness and distribution consistency, avoiding the training and testing deviation (data leakage) phenomenon.
[0024] Step 3: Feature construction; like Figure 2 As shown in Figure 1, the multimodal feature information extracted from UGC includes three categories: video, text, and image. The specific process of feature construction is as follows: Modality Amplification: For video covers in the image modality, we use the llava model to perform image-to-text processing to achieve modality amplification and increase the diversity of modal representations. For video modality, we take 10 frames at equal intervals and convert them into image modality.
[0025] Embedding conversion: Convert text and image modality information into embeddings through the jina-clip-v2 model.
[0026] The definitions of the variables involved are shown in Equations 1 to 10: Formula 1: ; in Indicates the Video modality information of samples; is the number of video frames, It is The first sample frame.
[0027] Formula 2: ; in Indicates the The title information of each sample; is the number of title words, It is The first sample tokens.
[0028] Formula 3: ; in Indicates the Comment information of samples; represents splicing; is the number of comments, It is The first sample Comments.
[0029] Formula 4: ; in Indicates the Cover information of each sample; represents the set of real numbers; Represents the height of the image; Represents the width of the image; Represents the number of channels of the image.
[0030] Formula 5: ; in Indicates the The cover features of the sample are amplified by the llava model to generate text modal features; is the total number of words in the current sample, It is The first sample tokens, Refers to the position of the word in the current sample indivual.
[0031] Formula 6: ; Formula 7: ; Formula 8: ; Formula 9: ; Formula 10: ; in 、 、 、 、 They are 、 、 、 、 Embedding It is a text feature embedding function based on the jina-clip-v2 model; It is an image feature embedding function based on the jina-clip-v2 model.
[0032] Step 4: Construction of UGC knowledge vector library; The UCG knowledge vector library is constructed for text modality, and the specific process is as follows: In order to reduce the computational complexity of the final prediction model, all text features are first concatenated, and then the jina-clip-v2 model is used to generate long text embeddings. The relevant formula is as follows: Formula 11: ; Formula 12: ; in Text embeddings for UGC samples (used for training and retrieval queries); It is the text embedding in the concatenated UCG knowledge vector library; 、 、 Corresponding to the first The title text features of each sample, the text features after modal amplification of the cover features, and the comment text features.
[0033] Step 5: Two-way retrieval and rearrangement recall; Text embedding for UGC samples (in Indicates the samples, is the dimension of the dense vector), in the UGC knowledge vector library ( is the total number of documents in the UGC knowledge vector library) and perform the following retrieval and fusion operations (for intuitive representation, Figure 3 (shown in human-readable form): (1) Dense vector similarity retrieval ( Figure 3 (Retriever branch): calculate The cosine similarity with all vectors in the UCG knowledge vector library is as follows: Formula 13: ; in express and the first j vectors The cosine similarity of .
[0034] After sorting in descending order by cosine similarity, recall the Top-J results and record them as a set .
[0035] (2) BM25 search based on word frequency ( Figure 3 BM25 branch line): The relevance score is calculated using the BM25 algorithm based on word frequency statistics. The relevant formula is as follows: Equation 14: ; Equation 15: ; in Score for BM25; is the first The text content of a document; For query The terms in is the inverse document frequency; For terms exist The frequency in and For adjustment parameters (usually , ); For Documents length; is the average length of all documents in the UCG knowledge vector library; is the total number of documents in the UGC knowledge vector library; To include terms The number of documents.
[0036] After sorting in descending order by BM25 score, recall the Top-J results and record them as a set .
[0037] (3) RFF (Rank Fusion Function) algorithm dual-path fusion recall ( Figure 3 RFF branch line): The dense vector similarity retrieval results and BM25 retrieval results are input into the RFF algorithm to calculate the fusion score. The formula is as follows: Equation 16: ; in Score for fusion; is the first k vectors; For the search method, For dense vector similarity retrieval, It is BM25 search based on word frequency; is the smoothing hyperparameter; for Ranking in the corresponding search results: for Similarity retrieval results in dense vectors Rank in (set to ), for Corresponding documents Results in BM25 Rank in (set to ).
[0038] according to After sorting in descending order, take the top-K with the highest scores , and get the final recall set: Equation 17: ; in is the recall set (i.e., RAP results); are the first K vectors sorted by score from high to low, .
[0039] Step 6: Retrieve enhanced prediction results (RAP results) to assist downstream network UGC popularity prediction; Combining the UGC popularity prediction method based on dual-path retrieval and re-ranking recall architecture, we use existing deep learning network models suitable for UGC popularity prediction (such as attention mechanism networks, multimodal fusion networks) or existing machine learning models (such as CatBoost, SVR) to build a more powerful end-to-end deep learning model (the core of which is the downstream deep learning network module) for UGC popularity prediction. The specific process is as follows ( Figure 4 ): Training phase: The input of the model is the feature representation after multimodal fusion. It is shown as follows: For each video modality in the training set , in the UGC knowledge vector library Perform a search based on a dual-path search and rearrangement recall architecture to obtain the Top-K set ; Find its characteristic mean : Equation 18: ; For each title mode in the training set , in the UGC knowledge vector library Perform a search based on a dual-path search and rearrangement recall architecture to obtain the Top-K set ; Find its characteristic mean : Equation 19: ; For each review modality in the training set , in the UGC knowledge vector library Perform a search based on a dual-path search and rearrangement recall architecture to obtain the Top-K set ; Find its characteristic mean : Equation 20: ; For each cover mode in the training set , in the UGC knowledge vector library Perform a search based on a dual-path search and rearrangement recall architecture to obtain the Top-K set ; Find its characteristic mean : Equation 21: ; For each review modality in the training set In the UGC knowledge vector library , a retrieval based on a two-way retrieval and rearrangement recall architecture is performed to obtain a Top-K set ; and the feature mean value thereof is : Equation 22: ; For a general deep learning prediction network , the prediction result is: Equation 23: ; wherein ; The output is a continuous popularity score.
[0040] The training target is to minimize the error between the prediction value and the true value (UGC content likes, plays, and comments in step 1). The target is: Equation 24: ; wherein MSE is the mean square error, n is the batch size, is the popularity label value of the i-th sample, is the popularity prediction value of the i-th sample.
[0041] The standard supervised learning paradigm is adopted, the optimizer is AdamW, and a dynamic learning rate decay strategy is set to improve the convergence stability. Through this process, a trained deep learning model is finally obtained.
[0042] Prediction phase: The trained deep learning model outputs accurate UGC content popularity prediction values based on the input target sample multi-modal features.
[0043] Because the UGC popularity prediction method based on a two-way retrieval and rearrangement recall architecture is used, the network parameter quantity of the deep learning system can be effectively reduced, and more accurate prediction results can be obtained with less computation. The RAP method not only captures deep semantic associations, but also retains the accuracy of keyword matching, enabling the prediction model to make more reliable inferences based on the true popularity of similar content. At the same time, the relevant cases retrieved provide intuitive evidence for the prediction results, enhancing the credibility of the model. At the same time, the RAP method requires the addition of a mechanism for dynamically adjusting feature fusion weights based on retrieval result confidence in the downstream deep learning network module to avoid introducing noise - if the quality of the vector retrieval library is poor, the information of Top-K is not related to the target UGC content, which will introduce strong noise. The retrieval result confidence dynamic adjustment feature fusion weight is implemented in . The weight of each retrieval library vector is Similarity: The greater the similarity, the higher the weight. For example, the weight for: Equation 25: ; The calculation methods for other vectors are similar.
[0044] The specific implementation of the present invention is described in detail below with reference to specific embodiments.
[0045] Example 1: Practical application verification of the present invention in the UGC popularity prediction scenario; 1. Retrieval accuracy verification experiment; To verify the RFF algorithm's dual-path fusion recall advantage in retrieval accuracy, we compared it with retrieval based solely on dense vector similarity (ES) and retrieval based on word frequency (BM25). Evaluation metrics include Recall@K and MRR@K. Recall@K represents the proportion of relevant results in the search results of the top K UGC knowledge vector libraries; MRR@K represents the mean reciprocal ranking, which is the reciprocal position of the first relevant result in the search results of the top K UGC knowledge vector libraries. This is shown in Equation 26: Equation 26: ; in It is the set of UGC content to be retrieved in the UGC knowledge vector library; is the number of the set; It is The position of the first relevant result in the query results of the content to be retrieved (if not in the previous is recorded as 0).
[0046] The results are shown in Table 1: Table 1 Search results
[0047] As can be seen from Table 1, RFF (RFF algorithm dual-path fusion recall) significantly outperforms ES and BM25 in all indicators. Among them, MRR@10 is improved by 6.76% compared with ES, and Recall@1 is improved by 6.96% compared with ES. This effectively solves the problem of poor embedding accuracy of long texts and significantly enhances retrieval accuracy.
[0048] 2. Computation time comparison experiment; To verify the computational efficiency advantage of the RFF algorithm's dual-path fusion recall, we compared its computational time with that of the large language model Qwen3-Beranker-0.6B. The experimental condition was Top-J = 1000 Recall (the top 1000 results).
[0049] The results are shown in Table 2: Table 2 Calculation time results
[0050] As can be seen from Table 2, the computation time of RFF (RFF algorithm dual-path fusion recall) (1.2858s) is much lower than that of the large language model (34.3237s), significantly reducing the computational cost and showing a clear efficiency advantage.
[0051] 3. Downstream network prediction performance verification experiment; To verify the performance improvement of the present invention (RAP method) on popularity prediction models, experiments were conducted on the MicroLens-100k-Dataset (with a training set to test set ratio of 0.8:0.2). Comparison models included the TMALL model (a transductive model used in multimodal domains), CatBoost (an algorithm based on gradient boosting decision trees (GBDT)), SVR (a regression version of the support vector machine (SVM), a support vector regression model (SVR) that uses a Gaussian kernel function in this experiment, suitable for nonlinear data), and a version combined with RAP. Evaluation metrics were NMSE and SRC. NMSE is the normalized mean square error, as shown in Equation 27: Equation 27: ; Where MSE is the mean square error; is the variance of the true value, , is the mean of the true values; is the true value; is the predicted value; is the number of samples.
[0052] SRC is the Spearman rank correlation coefficient, which is used to measure the monotonic correlation between two variables. It is calculated based on the Pearson correlation coefficient based on the ranking of the variables. As shown in Formula 28: Equation 28: ; in , represents the difference between the true value ranking and the predicted value ranking; is the number of samples.
[0053] The results are shown in Table 3: Table 3 Downstream network experimental results
[0054] As can be seen from Table 3, after adding the RAP method, the NMSE of all baseline models is significantly reduced (the prediction error is reduced), and the SRC is significantly increased (the correlation between the true value and the predicted value is enhanced), indicating that the RAP method can effectively assist the downstream network to improve the performance of UGC popularity prediction.
[0055] The above are only preferred embodiments of the present invention. It should be pointed out that for those skilled in the art, several variations and improvements can be made without departing from the concept of the present invention. These should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent.
Claims
1. A method for predicting popularity of user-generated content based on retrieval-enhanced prediction, characterized in that: The following steps are involved: Step 1: Data collection; Collect multimodal data on UGC websites, including video streams, titles, comment content, covers, and popularity behavior indicators such as likes, views, and comments; Step 2: Data preprocessing; The collected raw data is cleaned, duplicates are removed, modality loss is processed, and text and images are standardized. The samples are then divided into a UGC knowledge vector library set, a training set, and a test set. Step 3: Feature construction; Extract three types of modal features: video, text, and image, and generate multimodal embedding representations through modality amplification and embedding transformation; Step 4: Construction of UGC knowledge vector library; For text modality, all text features are concatenated and long text embedding is generated to form the UCG knowledge vector library; Step 5: Two-way search and rearrangement recall; For the text embedding of UGC samples, we perform dense vector similarity search and word frequency-based BM25 search in the UGC knowledge vector library. We then use the RFF algorithm to dual-channel fusion recall of the Top-K results to obtain the final recall set: Step 6: Retrieve the enhanced prediction results to assist downstream network UGC popularity prediction; Combined with a deep learning model, model training and popularity prediction are performed based on the input multimodal features and recall sets, and a mechanism is provided in the downstream network to dynamically adjust the feature fusion weights according to the confidence of the retrieval results.
2. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 1, characterized in that: In step 3, modality amplification includes: Use the llava model to process the video cover image to text and generate text modal features: ; in Indicates the The cover features of the sample are amplified by the llava model to generate text modal features; is the total number of words in the current sample, It is The first sample tokens, Refers to the position of the word in the current sample indivual; For the video modality, 10 frames are taken at equal intervals and converted into image modality; Convert text and image modality information into embeddings using the jina-clip-v2 model: ; ; ; ; ; in 、 、 、 、 They are 、 、 、 、 Embedded, Indicates the Video modality information of samples, Indicates the The header information of each sample, Indicates the Sample review information, Indicates the Cover information of each sample; It is a text feature embedding function based on the jina-clip-v2 model; It is an image feature embedding function based on the jina-clip-v2 model.
3. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 2, characterized in that: In step 4, the construction of the UGC knowledge vector library includes: First, we concatenate the text features of the title, cover augmented text, and comments, and then use the jina-clip-v2 model to generate long text embeddings. The relevant formula is as follows: ; ; in Text embedding for UGC samples; represents splicing; It is the text embedding in the concatenated UCG knowledge vector library; 、 、 Corresponding to the first The title text features of each sample, the text features after modal amplification of the cover features, and the comment text features.
4. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 3, characterized in that: In step 5, dense vector similarity retrieval is achieved by calculating cosine similarity, and the formula is as follows: ; in Text embedding representing UGC samples and the first j vectors The cosine similarity of is the total number of documents in the UGC knowledge vector library; After sorting in descending order by cosine similarity, recall the Top-J results and record them as a set .
5. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 4, characterized in that: In step 5, the calculation formula for the relevance score of the BM25 search based on word frequency is as follows: ; ; in Score for BM25; is the first The text content of a document; For query The terms in is the inverse document frequency; For terms exist The frequency in and is the adjustment parameter; For Documents length; is the average length of all documents in the UCG knowledge vector library; To include terms The number of documents; After sorting in descending order by BM25 score, recall the Top-J results and record them as a set .
6. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 5, characterized in that: In step 5, the fusion score of the RFF algorithm is: ; in Score for fusion; is the first k vectors; For the search method, For dense vector similarity retrieval, It is BM25 search based on word frequency; is the smoothing hyperparameter; for Ranking in the corresponding search results: for Similarity retrieval results in dense vectors Ranking in for Corresponding documents Results in BM25 Ranking in according to After sorting in descending order, take the top-K with the highest scores , and get the final recall set: ; in is the recall set, i.e., the RAP result; are the first K vectors sorted by score from high to low.
7. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 1, characterized in that: In step 6, the supervised learning paradigm is adopted in the model training phase, AdamW is selected as the optimizer, and a dynamic learning rate decay strategy is set; the input of the model is the feature representation after multimodal fusion, and the output is a continuous popularity score. The training goal is to minimize the error between the predicted value and the true value.
8. The method for predicting popularity of user-generated content based on retrieval-enhanced prediction according to claim 1, characterized in that: The prediction performance of step 6 was evaluated by normalized mean square error and Spearman rank correlation coefficient.
Citation Information
Patent Citations
Method and device for predicting popularity of user generated content in social network
CN113139134A
Social media popularity prediction method based on three-mode fusion expert model
CN118469096A
Multimodal social media popularity prediction method based on hypergraph retrieval enhancement
CN118690069A
Large language model knowledge base question answering system based on multi-path fusion recall retrieval algorithm
CN120144773A
Method and system for determining a popularity of online content
WO2013064505A1
Cited By
User participation degree prediction method based on distillation multi-modal retrieval enhancement
CN121434461A
User engagement prediction method based on distillation multimodal retrieval enhancement
CN121434461B