Short video personalized recommendation method and system integrating theme and emotion

CN122796239APending Publication Date: 2026-09-22HUAQIAO UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611257900.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

若推荐模型未充分利用视频主题类别、视频标签、视频描述等内容信息,则难以从内容层面对用户兴趣进行准确刻画

Benefits of technology

[0040](1)本发明通过将短视频的主题类别、视频标签、视频描述和情感极性引入推荐模型,相比仅依赖用户行为数据和视频ID的推荐方法,能够更充分地利用短视频内容语义信息,从内容层面对用户兴趣进行准确刻画;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122796239A_ABST
    Figure CN122796239A_ABST
Patent Text Reader

Abstract

The application discloses a short video personalized recommendation method and system integrating themes and emotions, and relates to the technical field of video recommendation.The method comprises the following steps: obtaining a historical interactive short video sequence of a target user and a candidate short video set, and analyzing to obtain video semantic information; constructing a recommendation model, encoding the semantic information into a semantic feature vector, adding the semantic feature vector and a video ID embedding after adjusting by a learnable weight to obtain a fusion feature, adding the fusion feature and a position embedding, and then encoding by a Transform to obtain a user interest feature, calculating a matching score of the candidate fusion feature by dot product, and generating a recommendation list according to the matching score.The application introduces theme categories, video tags, video descriptions and emotional polarity information of short videos, combines a learnable weight fusion mechanism and sequence interest modeling, and improves the content representation capability and recommendation accuracy of short videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video recommendation technology, specifically to a method and system for personalized short video recommendation that incorporates themes and emotions. Background Technology

[0002] With the rapid development of mobile internet and short video platforms, short videos have become an important content carrier for users to obtain information, enjoy leisure and entertainment, express opinions, and engage in social interaction. Short videos typically contain multi-source information, including visual images, audio, subtitles, video titles, and user interactions, and are characterized by rich content formats, rapid dissemination, and high user engagement. Faced with a massive amount of short video content, users find it difficult to quickly find content that matches their interests and preferences. Therefore, personalized recommendation technology has become an important technical means for short video platforms to achieve content distribution and alleviate information overload.

[0003] Currently, mainstream personalized recommendation methods for short videos primarily rely on modeling users' historical interaction behaviors, such as clicks, views, dwell times, and likes. These methods typically learn collaborative filtering relationships between users and videos by embedding video IDs and predicting the video a user is likely to interact with next based on their historical interaction sequences. However, existing personalized recommendation methods for short videos still have the following shortcomings.

[0004] First, existing recommendation methods lack sufficient semantic understanding of the video content itself. Current methods typically focus on using users' historical interaction behavior and video IDs to learn the relationship between users and videos. However, short videos themselves contain rich visual, audio, and textual semantic information. If the recommendation model does not fully utilize content information such as video topic categories, video tags, and video descriptions, it will be difficult to accurately characterize user interests at the content level.

[0005] Second, existing recommendation methods do not adequately utilize the topic and tag information of short videos. Short video content typically covers different topics such as lifestyle, entertainment, education, games, technology, and social hot topics, and different users have different interests and preferences for different topics. If the recommendation model cannot effectively utilize semantic information such as video topic categories, video tags, and video descriptions, it will be difficult to accurately characterize user interests at the video content level.

[0006] Third, existing recommendation methods are insufficient in modeling the sentiment polarity of short videos. Short videos not only contain objective content but also emotional tendencies such as positive, negative, and neutral. Users' interest in videos is not only related to the video's topic but may also be related to the emotional atmosphere and sentiment expressed in the video. If the recommendation model ignores sentiment polarity, it will be difficult to accurately capture users' emotional preferences, affecting the accuracy and rationality of the recommendation results.

[0007] Fourth, existing methods for fusing video semantic features and video ID features are relatively simple. Although some recommendation methods incorporate text features such as video titles, tags, or descriptions, they typically use direct concatenation, simple addition, or fixed-weight fusion methods, making it difficult to adaptively adjust the contribution ratio of different features according to the importance of different semantic information. Summary of the Invention

[0008] To address the aforementioned issues, this invention proposes a personalized recommendation method and system for short videos that integrates themes and emotions. By introducing video semantic information and adaptively fusing it with video ID embedding features using learnable weight vectors, and combining it with Transformer sequence modeling to capture dynamic changes in user interests, the method improves the representation ability of short video content and the accuracy of recommendations.

[0009] On the one hand, a personalized recommendation method for short videos that incorporates themes and emotions includes:

[0010] S1. Obtain the target user's historical interaction short video sequence and candidate short video set, perform content semantic analysis on the short videos and candidate short videos in the historical interaction short video sequence, and obtain the video semantic information of each short video;

[0011] S2, construct and train a personalized short video recommendation model to obtain a trained personalized short video recommendation model. The personalized short video recommendation model includes a video semantic feature extraction and encoding module, a video semantic and ID feature fusion module, a sequence interest modeling module, and a prediction matching module.

[0012] The video semantic feature extraction and encoding module encodes the video semantic information of each short video into a video semantic feature vector through semantic representation; the video semantic and ID feature fusion module obtains the video ID embedding features of each short video, adjusts the weights of the video semantic feature vectors through a learnable weight vector, and adds the weighted video semantic features and video ID embedding features to obtain fused video features; the sequence interest modeling module adds the fused video features and the position embedding vectors and inputs them into a Transformer encoder for sequence modeling, taking the output vector of the last position in the sequence as the user interest feature; the prediction matching module calculates the prediction preference score between the user interest feature and the fused video features of the candidate short videos through dot product as the matching score, and generates a personalized short video recommendation list based on the matching score;

[0013] S3: Input the semantic information of each short video of the target user into the trained short video personalized recommendation model to obtain the recommendation score of the candidate short videos, and generate a personalized short video recommendation list based on the recommendation score.

[0014] Furthermore, in S1, video semantic information The calculation formula is as follows:

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] ;

[0020] in, Indicates the topic category; Indicates video tags; Indicates emotional polarity; This indicates a description of the video content; Indicates the first A short video; Represents a thematic semantic field; This represents a semantic field for the tag; Representing sentiment semantic fields and This represents a content description field; , , and These represent the deterministic text formatting rules for the corresponding semantic fields; This indicates a structured semantic template generation rule that connects semantic fields according to a preset field order. This indicates a string concatenation operation.

[0021] Furthermore, in S2, the video semantic information is encoded into a video semantic feature vector through semantic representation, and the calculation formula is as follows:

[0022] ;

[0023] ;

[0024] ;

[0025] ;

[0026] The Tokenizer function represents the word segmentation and serialization processing function, which is used to convert the video semantic information of the i-th short video into a Token number sequence according to the vocabulary corresponding to the pre-trained semantic representation model, and to truncate or pad according to the maximum Token length, while generating the corresponding attention mask. Represents the token number sequence ; Indicates an attention mask; The maximum token length representing the semantic information of the video; The parameter is Pre-trained semantic representation model; This represents the hidden state matrix output by the last layer of the pre-trained semantic representation model; Indicates the position of the last valid token; This represents the pooling operation that extracts the hidden state of the last valid token based on the attention mask; Indicates the first The video semantic feature vector of a short video; The dimension representing the semantic feature vector of the video; , and This represents the hidden state vector output by the i-th short video in the last layer of the pre-trained semantic representation model at the positions corresponding to the 1st, 2nd, and maximum token lengths. Represents the space of real numbers with the maximum number of rows and columns of the longest token; This represents the token position index in the token number sequence; Let represent a d-dimensional real vector space.

[0027] Furthermore, in S2, the location embedding vector and the fused video features have the same vector dimension, which is used to characterize the temporal sequence of each short video in the user's historical interaction sequence.

[0028] Furthermore, in S2, the calculation formula for fused video features is as follows:

[0029] ;

[0030] ;

[0031] ;

[0032] ;

[0033] in, Represents the video ID embedding matrix; Indicates the video ID embedding feature; This indicates that the video ID embedding feature is retrieved from the video ID embedding matrix based on the video ID. Indicates the video ID; This represents a learnable weight vector with the same dimension as the video semantic feature vector; Indicates the first The fusion video features of short videos; Represents the semantic feature vector of the video; Represents the globally learnable weight vector; Indicates the first Adjustment weights corresponding to each semantic feature dimension; This represents element-wise multiplication; Indicates The elements in the matrix are diagonal matrices composed of diagonal elements; Let N represent the space of real matrix elements with N rows and d columns.

[0034] On the other hand, personalized recommendation systems for short videos that incorporate themes and emotions include:

[0035] The semantic information acquisition module is used to acquire the target user's historical interaction short video sequence and candidate short video set, perform content semantic analysis on the short videos and candidate short videos in the historical interaction short video sequence, and obtain the video semantic information of each short video.

[0036] The training module is used to build and train a personalized short video recommendation model to obtain a trained personalized short video recommendation model. The personalized short video recommendation model includes a video semantic feature extraction and encoding module, a video semantic and ID feature fusion module, a sequence interest modeling module, and a prediction and matching module.

[0037] The video semantic feature extraction and encoding module encodes the video semantic information of each short video into a video semantic feature vector through semantic representation; the video semantic and ID feature fusion module obtains the video ID embedding features of each short video, adjusts the weights of the video semantic feature vectors through a learnable weight vector, and adds the weighted video semantic features and video ID embedding features to obtain fused video features; the sequence interest modeling module adds the fused video features and the position embedding vectors and inputs them into a Transformer encoder for sequence modeling, taking the output vector of the last position in the sequence as the user interest feature; the prediction matching module calculates the prediction preference score between the user interest feature and the fused video features of the candidate short videos through dot product as the matching score, and generates a personalized short video recommendation list based on the matching score;

[0038] The recommendation module is used to input the semantic information of each short video of the target user into the trained short video personalized recommendation model, obtain the recommendation score of the candidate short videos, and generate a personalized short video recommendation list based on the recommendation score.

[0039] The present invention adopts the above technical solution and has the following beneficial effects:

[0040] (1) By introducing the topic category, video tag, video description and sentiment polarity of short videos into the recommendation model, this invention can make fuller use of the semantic information of short video content and accurately characterize user interests from the content level, compared with the recommendation method that only relies on user behavior data and video ID.

[0041] (2) The present invention adjusts the weight of the video semantic feature vector element by element through the learnable weight vector, and adds the weighted video semantic features with the video ID embedding features to obtain the fused video features, so that the model can adaptively adjust the contribution ratio of each semantic feature dimension according to the training process, avoiding the problem of insufficient adaptability of fixed weight or simple splicing fusion method.

[0042] (3) This invention performs sequence modeling by adding the fused video features and the location embedding vector and inputting them into the Transformer encoder. The output vector of the last position in the sequence is taken as the user interest feature. This enables the model to use the location embedding to represent the temporal relationship of each short video in the sequence. The multi-head self-attention mechanism captures the dynamic change of user interest over time, thereby improving the user interest modeling capability. The matching score (predicted preference score, recommendation score) between the user interest feature and the candidate short video feature is calculated by dot product to generate recommendation results, thereby improving the recommendation accuracy. Attached Figure Description

[0043] Figure 1 This is a flowchart of the personalized short video recommendation method incorporating themes and emotions according to an embodiment of the present invention;

[0044] Figure 2 This is a model structure diagram of user and candidate video matching scores in an embodiment of the present invention;

[0045] Figure 3 This is a diagram of a personalized short video recommendation system that incorporates themes and emotions, as described in an embodiment of the present invention. Detailed Implementation

[0046] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0047] like Figure 1 As shown, the present invention integrates thematic and emotional personalized recommendation methods for short videos, including:

[0048] S1. Obtain the target user's historical interaction short video sequence and candidate short video set, perform content semantic analysis on the short videos and candidate short videos in the historical interaction short video sequence, and obtain video semantic information.

[0049] Specifically, video semantic information The calculation formula is as follows:

[0050] ;

[0051] ;

[0052] ;

[0053] ;

[0054] ;

[0055] in, Indicates the topic category; Indicates video tags; Indicates emotional polarity; This indicates a description of the video content; Indicates the first A short video; Represents a thematic semantic field; This represents a semantic field for the tag; Representing sentiment semantic fields and This represents a content description field; , , and These represent the deterministic text formatting rules for the corresponding semantic fields; This indicates a structured semantic template generation rule that connects semantic fields according to a preset field order. This indicates a string concatenation operation.

[0056] Specifically, the video semantic information includes topic categories, video tags, video descriptions, and sentiment polarity. These are generated by at least one multimodal large model analyzing the visual, audio, and textual information of the short video. The visual information includes the short video's keyframe sequence; the audio information includes the short video's audio text or acoustic features; and the textual information includes the video title, subtitle text, speech recognition text, or other descriptive text. The topic categories, video tags, and video descriptions are obtained by the multimodal large model through cross-modal semantic reasoning based on the short video's keyframe sequence and audio text; the sentiment polarity is obtained by the multimodal large model through sentiment reasoning based on the short video's facial keyframes, acoustic features, and textual information.

[0057] Specifically, in this embodiment, the topic category, video tag, and video description are generated by the Qwen3-VL-4B model combined with retrieval enhancement methods; the sentiment polarity is generated by the Qwen2.5-Omni-3B model after sentiment inference pre-training and sentiment prediction fine-tuning.

[0058] S2, Construct and train a personalized recommendation model for short videos to obtain a trained personalized recommendation model for short videos. The personalized recommendation model for short videos includes a video semantic feature extraction and encoding module, a video semantic and ID feature fusion module, a sequence interest modeling module, and a prediction and matching module.

[0059] The video semantic feature extraction and encoding module encodes video semantic information into video semantic feature vectors through semantic representation; the video semantic and ID feature fusion module obtains video ID embedding features, adjusts the weights of the video semantic feature vectors using learnable weight vectors, and adds the weighted video semantic features and video ID embedding features to obtain fused video features; the sequence interest modeling module adds the fused video features and position embedding vectors and inputs them into a Transformer encoder for sequence modeling, taking the output vector of the last position in the sequence as the user interest feature; the prediction matching module calculates the prediction preference score between the user interest feature and the fused video features of the candidate short video through dot product as the matching score, and generates personalized short video recommendation results based on the matching score.

[0060] Specifically, based on the obtained video semantic information corresponding to each short video, the topic category, video tag, sentiment polarity, and video description are first constructed into structured video semantic text according to a preset semantic template. Then, a pre-trained semantic representation model is used to encode this into a video semantic feature vector. Video ID embedding features include the video ID embedding features of each short video in the historical interactive short video sequence, and the video ID embedding features of each candidate short video in the candidate short video set. Both types of videos query the corresponding ID embedding features from the same video ID embedding matrix based on their respective video IDs, and use the same feature fusion method. The video semantic and ID feature fusion module obtains the video ID embedding features of each historical interactive short video and each candidate short video, respectively. It adjusts the corresponding video semantic feature vectors through a globally shared learnable weight vector, and then adds the two together to obtain the fused video features of each short video.

[0061] Specifically, both historical interactive short videos and candidate short videos are analyzed separately as individual videos to obtain the video semantic information corresponding to each short video. In stage S1, the two types of videos are not aggregated, and they are not aggregated together later. Subsequently, the fusion features of historical interactive short videos are combined into a feature sequence according to the interaction order to model user interests, and the fusion features of candidate short videos are used to match user interest features.

[0062] Specifically, the location embedding vector refers to the vector representing the position of a historical interactive short video in the user's behavior sequence, such as the 1st, 2nd, and 3rd interaction positions. It represents the chronological order in which the videos appear. The sequence interest modeling module combines the fused video features of each historical interactive short video with its corresponding location embedding to form a feature sequence, which is then input into the Transformer encoder; the output vector of the last sequence position is taken as the user interest feature. In this embodiment, the dot product is used to calculate the predicted preference score, not to calculate the user interest feature. The user interest feature has already been obtained by the Transformer encoder, and the prediction matching module then performs a dot product between this user interest feature and the fused video features of each candidate short video to obtain the recommendation score for each candidate short video.

[0063] Specifically, the semantic information of the video is encoded into a video semantic feature vector through semantic representation, and the calculation formula is as follows:

[0064] ;

[0065] ;

[0066] ;

[0067] ;

[0068] The Tokenizer function represents the word segmentation and serialization processing function, which is used to convert the video semantic information of the i-th short video into a Token number sequence according to the vocabulary corresponding to the pre-trained semantic representation model, and to truncate or pad according to the maximum Token length, while generating the corresponding attention mask. Represents the token number sequence ; Indicates an attention mask; The maximum token length representing the semantic information of the video; The parameter is Pre-trained semantic representation model; This represents the hidden state matrix output by the last layer of the pre-trained semantic representation model; Indicates the position of the last valid token; This represents the pooling operation that extracts the hidden state of the last valid token based on the attention mask; Indicates the first The video semantic feature vector of a short video; The dimension representing the semantic feature vector of the video; , and This represents the hidden state vector output by the i-th short video in the last layer of the pre-trained semantic representation model at the positions corresponding to the 1st, 2nd, and maximum token lengths. Represents the space of real numbers with the maximum number of rows and columns of the longest token; This represents the token position index in the token number sequence; Let represent a d-dimensional real vector space.

[0069] Specifically, the location embedding vector and the fused video features have the same vector dimension, which is used to characterize the temporal sequence of each short video in the user's historical interaction sequence.

[0070] Specifically, the calculation formula for fusing video features is as follows:

[0071] ;

[0072] ;

[0073] ;

[0074] ;

[0075] in, Represents the video ID embedding matrix; Indicates the video ID embedding feature; This indicates that the video ID embedding feature is retrieved from the video ID embedding matrix based on the video ID. Indicates the video ID; This represents a globally shared, learnable weight vector with the same dimension as the video semantic feature vector. Indicates the first The fusion video features of short videos; Represents the semantic feature vector of the video; Represents the globally learnable weight vector; Indicates the first Adjustment weights corresponding to each semantic feature dimension; This represents element-wise multiplication; Indicates The elements in the matrix are diagonal matrices composed of diagonal elements; Let N represent the space of real matrix elements with N rows and d columns.

[0076] Specifically, in this embodiment, the structured semantic text is input into the general semantic representation model to obtain the video semantic feature vector, which can be represented as:

[0077] ;

[0078] in, Represents the semantic feature vector of the video. This represents a semantic representation model. This represents structured semantic text. The semantic feature acquisition unit in this embodiment is used to acquire the video semantic feature vector output by the semantic representation model, and the video semantic feature vector is represented as:

[0079] ;

[0080] Specifically, the semantic representation model is the Qwen3-Embedding-0.6B model; topic categories, video tags, video descriptions, and sentiment polarities are generated by analyzing the visual, audio, and textual information of the short video using at least one multimodal large model.

[0081] Specifically, in this embodiment, for the target user (u), its historical interaction short video sequence is represented as follows:

[0082] ;

[0083] in, Indicates user A series of historical interactive short videos, , , , This represents short videos that a user has interacted with in chronological order. During training, training samples are constructed from the user's historical sequences using a sliding window, utilizing the previous videos in the sequence. The video prediction One video, represented as:

[0084] ;

[0085] For the first in the user behavior sequence For each video, the fused feature vector of the video is added to the location embedding to obtain the input representation of the sequence encoder:

[0086] ;

[0087] in, Indicates video The fused feature vector, This represents the position embedding vector.

[0088] Specifically, the sequence interest modeling module uses a Transformer encoder to model short video sequences of users' historical interactions; The calculation process of the layer Transformer encoder is represented as follows:

[0089] ;

[0090] in, Indicates the first Layer input representation, This indicates a multi-head self-attention layer or a feedforward neural network layer. Representation layer normalization.

[0091] Specifically, the single-head attention calculation in the multi-head self-attention mechanism is represented as follows:

[0092] ;

[0093] ;

[0094] in, Represents the query matrix. Represents the key matrix. Represents a value matrix, They represent the learnable mapping matrices, This represents the dimension of the key vector.

[0095] Specifically, the computation of the feedforward neural network in this embodiment is represented as follows:

[0096] ;

[0097] in, , This is the weight matrix. , This is a bias term.

[0098] Specifically, user interest characteristics are processed through After the Transformer encoder, the last time step of the last layer's output is taken as the user interest feature, denoted as... The prediction matching module calculates the predicted preference score between user interest features and candidate short video features through dot product, expressed as:

[0099] ;

[0100] in, Indicates user's opinion on short videos The predicted preference score, Z represents the user's interest feature. This represents the fused video features of candidate short video i.

[0101] Specifically, the cross-entropy loss in this embodiment is expressed as:

[0102] ;

[0103] in, This represents the model training loss. Represents the training sample set, Indicates user Real interactive video The positive sample pairs formed Represents the set of all candidate videos. Indicates the candidate video index; This represents the fused video features of candidate short video j.

[0104] Specifically, the personalized short video recommendation list sorts candidate short videos in descending order of their recommendation scores and selects the top-scoring videos. The short videos are obtained; the recommendation score of the selected short videos is determined by the predicted preference score between the user's interest features and the fused video features of the candidate short videos.

[0105] Specifically, Figure 2 This embodiment demonstrates the calculation process for the user-candidate video matching score. Based on the user's historical behavior sequence, the ID features (after embedding) and semantic features (after extracting topic and sentiment information using Qwen Encoder) of each item are integrated through a learnable weight fusion layer, and then input into a multi-layer Transformer Encoder to capture temporal dependencies to generate the user's final interest features. On the right, the ID and semantic features of the candidate videos are processed in the same way and fused to obtain the candidate video features. Finally, the user interest features and the candidate video features are multiplied by a dot product to output the matching score.

[0106] S3: Input the target user's historical interaction short video sequence and candidate short video set into the trained short video personalized recommendation model to obtain the recommendation score of the candidate short videos, and generate a personalized short video recommendation list based on the recommendation score.

[0107] Specifically, the semantic information and video ID of each short video in the obtained historical interactive short video sequence, as well as the semantic information and video ID of each candidate short video in the candidate short video set, are input into the trained short video personalized recommendation model to obtain the recommendation score of each candidate short video. The candidate short videos are then sorted in descending order according to their recommendation scores, and the top K candidate short videos with the highest scores are selected to generate a personalized short video recommendation list. In this embodiment, the model's prediction and matching module outputs the recommendation score of each candidate short video, and the personalized short video recommendation list is generated by subsequent recommendation steps based on the recommendation scores and the selection of the top K candidate short videos.

[0108] Specifically, to verify the effectiveness of the short video personalized recommendation method incorporating themes and emotions described in this embodiment, the publicly available short video recommendation dataset MicroLens-100k was selected as the experimental dataset. This dataset originates from real short video social platforms and contains a large number of real-time interaction records between users and short videos. It exhibits high data sparsity and obvious time-series characteristics, effectively simulating the matching relationship between user historical behavior sequences and candidate short videos in short video recommendation scenarios. During the experiment, the target user's historical interaction short video sequence was used as the model input, and the short video that the user would interact with in the next moment was used as the prediction target. The model generates user interest features based on the target user's historical interaction short video sequence and calculates the matching score between the user interest features and the candidate short video features. Finally, the candidate short videos are sorted in descending order based on the matching score to obtain a personalized short video recommendation list.

[0109] Specifically, for the target user (u), their historical interaction short video sequence is represented as follows:

[0110]

[0111] in, Indicates user A series of historical interactive short videos, , , , This refers to the short videos that users have interacted with in chronological order. This indicates the length of the historical interactive short video sequence.

[0112] Specifically, in the experiment, the short video of the user's next real interaction was recorded as... The model predicts the recommendation score of each candidate short video in the candidate short video set based on the historical interaction short video sequence, and determines whether the actual interactive short video appears in the top K recommendation results. This embodiment was tested in a WSL2 environment on a Windows 10 64-bit operating system. The hardware platform configuration included an NVIDIA RTX A4000 GPU, an Intel Xeon E5-1607 v4 processor, and 16GB of system memory; the software environment was based on Python 3.10.19, primarily using deep learning libraries such as PyTorch 2.8.0 and Transformers 4.57.1.

[0113] This embodiment uses hit rate HR@K and normalized loss cumulative gain NDCG@K as evaluation metrics. HR@K measures whether genuine interactive short videos appear at the top of the recommendation list. In this context, NDCG@K is used to further measure the ranking position of the real-interaction short video in the recommendation list. In this embodiment, Set to 10 and 20, calculate HR@10, NDCG@10, HR@20, and NDCG@20 respectively. The calculation expression for HR@K is as follows:

[0114] ;

[0115] in, This represents the set of users in the test set. This represents the total number of users in the test set. This refers to one of the users; For indicator functions, when the condition The value is 1 when it is true, and 0 otherwise. The calculation expression for NDCG@K is as follows:

[0116] ;

[0117] in, This represents the set of users in the test set. This represents the total number of users in the test set. This refers to one of the users; For indicator functions, when the condition The value is 1 when it is true, and 0 otherwise.

[0118] If a genuine interactive short video does not appear in the top K recommendation results, the NDCG@K value for that user is 0. The higher the HR@K, the stronger the model's ability to recall genuine interactive short videos to the top (K) positions of the recommendation list; the higher the NDCG@K, the more the model can not only hit genuine interactive short videos, but also rank them higher in the list.

[0119] To verify the recommendation performance of the method in this embodiment compared to existing recommendation methods, the method in this embodiment is compared with SASRec, VRAgent-R1, LinkedOut, and DFF. SASRec is a sequence recommendation model based on the Transformer self-attention mechanism, which mainly predicts the user's next interest by modeling the dependencies in the user's historical interaction sequence; VRAgent-R1 is a recommendation method based on a large language model agent, which uses the reasoning ability of a large model to analyze user interests and candidate video information; LinkedOut uses a multimodal large model to perform semantic modeling of video content and introduces external knowledge representations to enhance semantic understanding in video recommendation; DFF uses a multimodal large model to extract semantic features of videos and integrates visual and textual information through a feature fusion module to achieve video recommendation. The comparative experimental results are shown in Table 1.

[0120] Table 1. Comparative Experiment Results;

[0121]

[0122] As can be seen from Table 1, the method in this embodiment achieved good experimental results on the four evaluation indicators: HR@10, NDCG@10, HR@20, and NDCG@20.

[0123] Compared to the SASRec method, the HR@10 of this embodiment improves from 0.0909 to 0.1041, HR@20 from 0.1278 to 0.1492, and NDCG@20 from 0.0610 to 0.0695. These results demonstrate that relying solely on video ID information from users' historical interaction sequences is insufficient to fully characterize short video content. This embodiment, by further incorporating semantic information such as topic categories, video tags, video descriptions, and sentiment polarity on top of video ID embedding features, enhances the model's understanding of candidate short video content, thereby improving the matching degree between recommendation results and user interests. Compared to recommendation methods based on large models or semantic enhancement, such as VRAgent-R1, LinkedOut, and DFF, this embodiment achieves superior results in both HR@10 and HR@20, indicating a stronger ability to recall potentially interesting short videos from the candidate short video set. The method in this embodiment also achieved better results in terms of NDCG@10 and NDCG@20 metrics, indicating that the method can not only improve the recommendation hit rate, but also make the ranking of real interactive short videos higher in the recommendation list.

[0124] like Figure 3 As shown, this embodiment also discloses a personalized short video recommendation system that incorporates themes and emotions, including:

[0125] The semantic information acquisition module 31 is used to acquire the target user's historical interaction short video sequence and candidate short video set, and to perform content semantic analysis on the short videos and candidate short videos in the historical interaction short video sequence to obtain video semantic information.

[0126] Training module 32 is used to construct and train a personalized recommendation model for short videos to obtain a trained personalized recommendation model for short videos. The personalized recommendation model for short videos includes a video semantic feature extraction and encoding module, a video semantic and ID feature fusion module, a sequence interest modeling module, and a prediction matching module.

[0127] The video semantic feature extraction and encoding module encodes video semantic information into video semantic feature vectors through semantic representation; the video semantic and ID feature fusion module obtains video ID embedding features, adjusts the weights of the video semantic feature vectors using learnable weight vectors, and adds the weighted video semantic features and video ID embedding features to obtain fused video features; the sequence interest modeling module adds the fused video features and position embedding vectors and inputs them into a Transformer encoder for sequence modeling, taking the output vector of the last position in the sequence as the user interest feature; the prediction matching module calculates the prediction preference score between the user interest feature and the fused video features of the candidate short video through dot product as the matching score, and generates personalized short video recommendation results based on the matching score;

[0128] The recommendation module 33 is used to input the target user's historical interaction short video sequence and candidate short video set into the trained short video personalized recommendation model, obtain the recommendation score of the candidate short video, and generate a personalized short video recommendation list based on the recommendation score.

[0129] The specific implementation of a short video personalized recommendation system that incorporates themes and emotions is the same as the short video personalized recommendation method that incorporates themes and emotions, and will not be described again in this embodiment.

[0130] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.

Claims

1. A method for personalized recommendation of short videos that incorporates themes and emotions, characterized in that, Includes the following steps: S1. Obtain the target user's historical interaction short video sequence and candidate short video set, perform content semantic analysis on the short videos and candidate short videos in the historical interaction short video sequence, and obtain the video semantic information of each short video; S2, construct and train a personalized short video recommendation model to obtain a trained personalized short video recommendation model. The personalized short video recommendation model includes a video semantic feature extraction and encoding module, a video semantic and ID feature fusion module, a sequence interest modeling module, and a prediction matching module. The video semantic feature extraction and encoding module encodes the video semantic information of each short video into a video semantic feature vector through semantic representation; the video semantic and ID feature fusion module obtains the video ID embedding features of each short video, adjusts the weights of the video semantic feature vectors using a learnable weight vector, and adds the weighted video semantic features and video ID embedding features to obtain the fused video features; sequence The interest modeling module adds the fused video features and the location embedding vector and inputs them into the Transformer encoder for sequence modeling, taking the output vector of the last position in the sequence as the user interest feature; the prediction matching module calculates the prediction preference score between the user interest feature and the fused video features of the candidate short video through dot product as the matching score, and generates a personalized short video recommendation list based on the matching score. S3: Input the semantic information of each short video of the target user into the trained short video personalized recommendation model to obtain the recommendation score of the candidate short videos, and generate a personalized short video recommendation list based on the recommendation score.

2. The personalized short video recommendation method incorporating themes and emotions according to claim 1, characterized in that, In S1, video semantic information The calculation formula is as follows: ; ; ; ; ; in, Indicates the topic category; Indicates video tags; Indicates emotional polarity; This indicates a description of the video content; Indicates the first A short video; Represents a thematic semantic field; This represents a semantic field for the tag; Representing sentiment semantic fields and This represents a content description field; , , and These represent the deterministic text formatting rules for the corresponding semantic fields; This indicates a structured semantic template generation rule that connects semantic fields according to a preset field order. This indicates a string concatenation operation.

3. The personalized short video recommendation method incorporating themes and emotions according to claim 1, characterized in that, In S2, the video semantic information is encoded into a video semantic feature vector through semantic representation, and the calculation formula is as follows: ; ; ; ; Wherein, Tokenizer() represents the word segmentation and serialization processing function, which is used to convert the video semantic information of the i-th short video into a token number sequence according to the vocabulary corresponding to the pre-trained semantic representation model; Represents the token number sequence ; Indicates an attention mask; The maximum token length representing the semantic information of the video; The parameter is Pre-trained semantic representation model; This represents the hidden state matrix output by the last layer of the pre-trained semantic representation model; Indicates the position of the last valid token; This represents the pooling operation that extracts the hidden state of the last valid token based on the attention mask; Indicates the first The video semantic feature vector of a short video; The dimension representing the semantic feature vector of the video; , and This represents the hidden state vector output by the i-th short video in the last layer of the pre-trained semantic representation model at the positions corresponding to the 1st, 2nd, and maximum token lengths. Represents the space of real numbers with the maximum number of rows and columns of the longest token; This represents the token position index in the token number sequence; Let represent a d-dimensional real vector space.

4. The personalized recommendation method for short videos incorporating themes and emotions according to claim 1, characterized in that, In S2, the location embedding vector and the fused video features have the same vector dimension, which is used to characterize the temporal sequence of each short video in the user's historical interaction sequence.

5. The personalized recommendation method for short videos incorporating themes and emotions according to claim 1, characterized in that, In S2, the calculation formula for fused video features is as follows: ; ; ; ; in, Represents the video ID embedding matrix; Indicates the video ID embedding feature; This indicates that the video ID embedding feature is retrieved from the video ID embedding matrix based on the video ID. Indicates the video ID; This represents a learnable weight vector with the same dimension as the video semantic feature vector; Indicates the first Adjustment weights corresponding to each semantic feature dimension; Indicates the first The fusion video features of short videos; Represents the semantic feature vector of the video; This represents element-wise multiplication; Indicates The elements in the matrix are diagonal matrices composed of diagonal elements; Let N represent the space of real matrix elements with N rows and d columns.

6. A personalized recommendation system for short videos that integrates themes and emotions, characterized in that, include: The semantic information acquisition module is used to acquire the target user's historical interaction short video sequence and candidate short video set, perform content semantic analysis on the short videos and candidate short videos in the historical interaction short video sequence, and obtain the video semantic information of each short video. The training module is used to build and train a personalized short video recommendation model to obtain a trained personalized short video recommendation model. The personalized short video recommendation model includes a video semantic feature extraction and encoding module, a video semantic and ID feature fusion module, a sequence interest modeling module, and a prediction and matching module. The video semantic feature extraction and encoding module encodes the video semantic information of each short video into a video semantic feature vector through semantic representation; the video semantic and ID feature fusion module obtains the video ID embedding features of each short video, adjusts the weights of the video semantic feature vectors using a learnable weight vector, and adds the weighted video semantic features and video ID embedding features to obtain the fused video features; sequence The interest modeling module adds the fused video features and the location embedding vector and inputs them into the Transformer encoder for sequence modeling, taking the output vector of the last position in the sequence as the user interest feature; the prediction matching module calculates the prediction preference score between the user interest feature and the fused video features of the candidate short video through dot product as the matching score, and generates a personalized short video recommendation list based on the matching score. The recommendation module is used to input the semantic information of each short video of the target user into the trained short video personalized recommendation model, obtain the recommendation score of the candidate short videos, and generate a personalized short video recommendation list based on the recommendation score.