A method for intelligent online video content recommendation that incorporates a learning interest model

By constructing dynamic interest vectors and deep semantic analysis, and combining user behavior and content features, a multi-dimensional relationship is established in the online video recommendation system. This solves the problems of personalization and real-time performance of interest models in existing systems, and achieves efficient and accurate content delivery, thereby improving learning efficiency and user satisfaction.

CN121000906BActive Publication Date: 2026-01-30SHENZHEN NEWVANE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511502857.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-30
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing online video recommendation systems struggle to build personalized learning interest models, fail to deeply integrate multi-source behavioral signals, and cannot perceive interest drift in real time. This results in homogenized and lagging recommendation results, making it impossible to adapt to the dynamic changes in user interests and impacting learning efficiency and user experience.

Method used

By constructing dynamic interest vectors and integrating user historical viewing records, search keywords, and interactive behavior data, an interest profile vector set is generated, and interest dimension clustering and weight allocation are performed. Deep semantic analysis is conducted on online video content to generate a video content feature vector set. A multi-dimensional relationship between user interests and video content is established, a similarity calculation model is used to calculate the matching degree, and a push sequence is generated through a multi-objective optimization strategy. User feedback is collected in real time for closed-loop optimization.

Benefits of technology

It achieves accurate characterization of user interests and in-depth content matching, improves the relevance and satisfaction of recommendations, ensures that the pushed content meets the user's current learning stage and exploration needs, and improves learning efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121000906B_ABST
    Figure CN121000906B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent online video content push method combining a learning interest model, comprising the following steps: constructing dynamic interest vectors based on multi-source user behavior data, generating a user interest profile vector set, and performing interest dimension clustering and weight allocation; generating a video content feature vector set based on the clustering results of the user interest profile vector set; establishing a multi-dimensional correlation between the user interest profile vector set and the video content feature vector set, and outputting a user-video matching confidence matrix; transforming the user-video matching confidence matrix into a push sequence and distributing it based on a multi-objective optimization strategy; and collecting user feedback on the pushed videos in real time to achieve closed-loop optimization. This invention has the following advantages and effects: it can achieve accurate perception and deep semantic matching of users' dynamic learning interests, significantly improving the accuracy, timeliness, and user satisfaction of content distribution, thereby optimizing learning efficiency and experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing technology, and in particular to an intelligent method for pushing online video content that combines a learning interest model. Background Technology

[0002] With the rapid development of information technology and the explosive growth of digital media resources, online video has become one of the core media for people to engage in self-directed learning, skills enhancement, and knowledge expansion. Especially in the fields of education, vocational training, and science popularization, video content, with its vivid and intuitive characteristics, meets users' needs for fragmented and personalized learning. However, the immense abundance of platform content has also brought the serious challenge of "information overload." This makes it crucial for users to efficiently and accurately discover content that aligns with their interests and cognitive levels from a vast amount of resources, significantly impacting their learning experience and effectiveness. Existing video recommendation systems, such as algorithms based on collaborative filtering, primarily analyze historical behavioral data of user groups to mine similarities and make recommendations. While this method can uncover potential interests to some extent, it inherently relies on group behavior patterns and struggles to address the uniqueness and evolution of individual interests. Especially for learners in their rapid growth phase, their interests may dynamically change as their cognitive boundaries expand, making collaborative filtering prone to homogenization and lag in recommendation results. On the other hand, content-based recommendation methods mainly rely on a shallow matching of video metadata (such as titles and tags) with the user's historical viewing content. They lack an understanding of the video's knowledge content, emotional tone, and logical depth, and cannot distinguish whether the user is attracted by the title or is genuinely interested in the essence of the content. This can easily lead to the trap of repeatedly recommending the same content.

[0003] While some advanced systems have attempted to incorporate multimodal data processing, a common problem remains: they fail to construct a personalized learning interest model capable of deeply integrating multi-source behavioral signals, perceiving interest shifts in real time, and dynamically matching them with the deep semantic features of videos. Specifically, traditional methods have the following limitations: First, in user modeling, most systems fail to organically integrate heterogeneous behavioral data from multiple sources, such as viewing history, search queries, likes, favorites, and comments, ignoring the differentiated intensity of intent and emotional investment behind different behaviors (for example, a deep search may represent true interest more than a random click), resulting in a one-sided and static user interest profile. Second, in content understanding, existing methods are often limited to surface-level tags, lacking multi-dimensional in-depth analysis of the thematic context, knowledge structure complexity, and emotional tendencies inherent in video content, leading to a mismatch between content representation and users' deep-seated learning motivations. Finally, in terms of feedback mechanisms, traditional systems often adopt a one-way push model, failing to establish an efficient and sensitive closed-loop optimization system. They cannot quickly adjust interest models and push strategies based on real-time user feedback (such as completion rate and interaction frequency), making it difficult to adapt to the continuous evolution of user interests. This core issue directly leads to poor recommendation performance: users are easily bombarded with seemingly relevant but actually repetitive or superficial content, while truly in-depth content tailored to their current learning stage and exploration needs remains largely unavailable. This not only reduces learning efficiency but also gradually erodes users' trust in the platform's recommendation function. Therefore, there is an urgent need in the field for a recommendation method that can truly understand users' dynamic learning interests and intelligently and accurately match them with the content of video knowledge, in order to overcome the bottlenecks of existing systems. Summary of the Invention

[0004] The purpose of this invention is to provide an intelligent online video content recommendation method that incorporates a learning interest model to address the problems mentioned in the background section.

[0005] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0006] A method for intelligently recommending online video content that incorporates a learning interest model includes the following steps:

[0007] S100. Construct dynamic interest vectors based on multi-source user behavior data, extract and fuse multimodal features from user historical viewing records, search keywords, and interactive behavior data to generate a user interest profile vector set, and perform interest dimension clustering and weight allocation on the user interest profile vector set.

[0008] S200. Based on the clustering results of the user interest profile vector set, perform deep semantic analysis and tag mapping on the video content in the online video library, extract the theme features, knowledge structure features and sentiment features of the video content, and generate a video content feature vector set.

[0009] S300. Establish a multi-dimensional association between the user interest profile vector set and the video content feature vector set, use a similarity calculation model to perform vector space alignment and matching degree calculation, and output the user-video matching confidence matrix.

[0010] S400: Based on a multi-objective optimization strategy, the user-video matching confidence matrix is ​​transformed into a push sequence, and the push sequence is sent to the user terminal in real time through the content delivery network interface;

[0011] The S500 collects user feedback on pushed videos in real time and dynamically adjusts the user interest profile vector set and push strategy based on the feedback data to achieve closed-loop optimization.

[0012] By adopting the above technical solutions, and constructing dynamic interest vectors and fusing multi-source user behavior data, deep mining and precise characterization of user interests are achieved, significantly improving the comprehensiveness and timeliness of the interest model. Through multimodal feature extraction and fusion, multi-dimensional user interest tendencies can be captured from viewing records, search keywords, and interactive behaviors, overcoming the limitations of traditional methods that rely on single behavioral data. Furthermore, through interest dimension clustering and weight allocation, the interest model possesses hierarchical and priority differentiation capabilities, better reflecting the user's true interest structure. Based on deep semantic analysis, multi-dimensional feature extraction of video content in terms of theme, knowledge structure, and sentiment tendencies makes content representation richer and more accurate, avoiding the problems caused by shallow tag matching. The system minimizes errors; by establishing a multi-dimensional correlation between user interests and video content and using a similarity calculation model for matching, it achieves higher-precision user-video alignment, effectively improving the relevance and satisfaction of recommendations; by generating push sequences using a multi-objective optimization strategy, it not only considers the matching degree but also takes into account the diversity and real-time nature of pushes, ensuring that users receive content that is both relevant to their interests and has exploratory value; by collecting user feedback in real time and dynamically adjusting the interest model and push strategy, a closed-loop optimization mechanism is constructed, enabling the system to adapt to the dynamic changes in user interests, continuously optimize the push effect, and ultimately achieve personalized, intelligent, and highly adaptable video content push, significantly improving users' learning efficiency, content acquisition experience, and platform stickiness.

[0013] A further setting is that S100 specifically includes the following steps:

[0014] Acquire multi-source user behavior data; wherein, the multi-source user behavior data includes user historical viewing records, search keyword sets, and interaction behavior data sets; the user historical viewing records include identification information and viewing duration of at least one viewed video; the search keyword set includes keyword sequences entered by the user in one or more search sessions; the interaction behavior data set includes at least one of the user's likes, favorites, comments, and shares of videos;

[0015] Multimodal feature extraction is performed on the multi-source user behavior data to generate viewing behavior feature vectors, search intent feature vectors, and interaction behavior feature vectors, respectively.

[0016] The viewing behavior feature vector, search intent feature vector, and interaction behavior feature vector are fused across modalities to generate a user interest profile vector set. The cross-modal feature fusion adopts a combination of self-attention mechanism and fully connected layer to align and weight the feature vectors of different types to generate a unified user interest representation vector.

[0017] The user interest profile vector set is subjected to interest dimension clustering analysis. An unsupervised clustering algorithm is used to divide user interests into multiple interest dimensions, and an initial weight is assigned to each interest dimension. The initial weight is determined based on the frequency and timeliness of the user behavior data contained in the interest dimension.

[0018] By adopting the above technical solutions and acquiring multi-source user behavior data, including historical viewing records, search keywords, and interactive behaviors, we can comprehensively capture users' interest expressions in different scenarios, enhancing the data foundation of the interest model. Through multimodal feature extraction of multi-source behavioral data—specifically viewing behavior feature vectors, search intent feature vectors, and interactive behavior feature vectors—different types of behavioral data can be effectively transformed into computable feature representations, providing high-quality input for subsequent fusion. Cross-modal feature fusion, employing a combination of self-attention mechanisms and fully connected layers, achieves alignment and weighted concatenation of different feature vectors, preserving the uniqueness of each modality while achieving information complementarity and enhancement, generating a more unified and expressive user interest representation vector. By performing interest dimension clustering analysis on the user interest profile vector set and assigning initial weights, the interest model becomes structured and interpretable. Furthermore, the weight settings, based on behavior frequency and timeliness, ensure the accuracy and dynamism of interest representation, ultimately improving the depth of user interest characterization and the targeting of push notifications.

[0019] A further step involves extracting multimodal features from the multi-source user behavior data to generate viewing behavior feature vectors, search intent feature vectors, and interaction behavior feature vectors, specifically including the following steps:

[0020] The viewing behavior feature vector is obtained by embedding and encoding video identifiers and viewing durations in the user's historical viewing records and by temporal modeling; the search intent feature vector is obtained by semantic encoding and attention weighting of the search keyword sequence; and the interaction behavior feature vector is obtained by one-hot encoding and weighting of different types of interaction behaviors.

[0021] By adopting the above technical solutions, and through embedding encoding and temporal modeling of viewing behavior feature vectors, it is possible to effectively capture the sequence patterns and duration information in the user's viewing history, thereby more accurately reflecting the user's continuous interests and attention preferences. By semantically encoding and attention-weighting the search keyword sequence, the ability to represent search intent is enhanced, enabling the identification of the user's core query intent and focus of interest, avoiding the one-sidedness of keyword matching. By performing one-hot encoding and weight allocation on interactive behaviors, different interaction types can contribute to the interest vector in a differentiated manner, reflecting the user's emotional investment and preference intensity for the content, further enriching the dimensions and accuracy of user interest representation, and improving the fineness and reliability of interest feature extraction as a whole.

[0022] A further setting is that S200 specifically includes the following steps:

[0023] Acquire multimodal video content data from an online video library; wherein, the multimodal video content data includes a video metadata set, a subtitle text set, and an audio feature set; the video metadata set includes the title, description, category tag, uploader information, and publication time of at least one video; the subtitle text set includes the transcribed text corresponding to the video or manually added subtitle text; the audio feature set includes acoustic features, speech emotion features, and background music features extracted from the video audio stream;

[0024] Multi-dimensional feature extraction is performed on the multimodal video content data to generate topic feature vectors, knowledge structure feature vectors, and sentiment tendency feature vectors, respectively.

[0025] The topic feature vector, knowledge structure feature vector, and sentiment tendency feature vector are fused across modalities to generate a video content feature vector set. The cross-modal feature fusion adopts a feature weighting and concatenation method based on an attention mechanism to align and integrate feature vectors from different sources to generate a unified video content representation vector.

[0026] Based on the clustering results of the user interest profile vector set, the video content feature vector set is labeled; wherein, the label mapping process matches the video content feature vector with the interest dimension, and assigns one or more interest tags and their matching degree weights to each video.

[0027] By adopting the above technical solutions, multimodal video content data, including video metadata, subtitle text, and audio features, is acquired, providing a rich data source for deep semantic analysis of videos and overcoming the problem of insufficient content representation relying solely on metadata. Multi-dimensional feature extraction from the multimodal video content data generates topic feature vectors, knowledge structure feature vectors, and sentiment feature vectors, achieving a comprehensive characterization of video content in terms of topic distribution, knowledge depth, and emotional tone. This makes the content representation more aligned with the characteristics of learning videos and the deep needs of users. Cross-modal feature fusion using an attention mechanism for feature weighting and concatenation effectively integrates features from different sources, highlighting important information and generating a unified and expressive video content representation vector. Tag mapping of video content based on user interest clustering results assigns interest tags and matching weights to videos, achieving a structured association between content and interest dimensions. This lays the foundation for subsequent accurate matching and significantly improves the depth of content understanding and the accuracy of content recommendations.

[0028] A further step involves extracting multi-dimensional features from the multimodal video content data to generate topic feature vectors, knowledge structure feature vectors, and sentiment tendency feature vectors, specifically including the following steps:

[0029] The topic feature vector is obtained by topic modeling and keyword extraction of the subtitle text set; the knowledge structure feature vector is obtained by analyzing the knowledge point distribution and logical relationship of the video metadata set; and the sentiment tendency feature vector is obtained by sentiment analysis of the audio feature set and the subtitle text set.

[0030] By employing the above technical solutions, the core themes and content focus of videos can be accurately captured through theme modeling and keyword extraction of subtitle text, enhancing the accuracy of content semantic representation. By analyzing the distribution and logical relationships of knowledge points in video metadata to generate knowledge structure feature vectors, the knowledge depth, hierarchical structure, and logical flow of videos can be identified, enabling the recommendation system to match users' knowledge needs and cognitive levels. Furthermore, by performing sentiment analysis on audio and subtitle text to generate sentiment tendency feature vectors, the emotional tone of videos and potential user emotional responses can be perceived, ensuring that recommended content is not only thematically relevant but also emotionally aligned with user preferences, thus comprehensively improving the dimensionality of video feature extraction.

[0031] A further setting is that S300 specifically includes the following steps:

[0032] The user interest profile vector set and the video content feature vector set are subjected to vector space alignment processing; wherein, the vector space alignment processing includes performing dimension mapping and scale normalization on each user interest profile vector in the user interest profile vector set and each video content feature vector in the video content feature vector set, so that the user interest profile vector set and the video content feature vector set are in the same vector space;

[0033] Based on the aligned user interest profile vector set and the video content feature vector set, a multimodal similarity calculation model is used to calculate the matching degree value between each user interest profile vector and each video content feature vector to generate a preliminary matching degree matrix; wherein, the multimodal similarity calculation model integrates two measurement methods, cosine similarity and Euclidean distance, and generates the preliminary matching degree matrix through a weighted linear combination;

[0034] The initial matching degree matrix is ​​calibrated with confidence to generate a calibrated matching degree matrix;

[0035] The calibrated matching degree matrix is ​​subjected to threshold filtering and sorting optimization to generate the final user-video matching confidence matrix. The threshold filtering is used to remove user-video matching pairs with matching confidence scores lower than a preset threshold, and then the remaining matching pairs are sorted in descending order of matching degree value.

[0036] By employing the aforementioned technical solutions, the user interest profile vector set and video content feature vector set are aligned in a vector space, including dimension mapping and scale normalization, ensuring their comparability within the same vector space and providing a foundation for similarity calculation. A multimodal similarity calculation model is used to fuse cosine similarity and Euclidean distance metrics, generating a preliminary matching matrix through a weighted linear combination. This approach balances directional similarity and distance proximity of the vectors, making the matching calculation more comprehensive and robust. Confidence calibration of the preliminary matching matrix further enhances the reliability and accuracy of the matching results. Threshold filtering and sorting optimization generate the final user-video matching confidence matrix, eliminating low-confidence matches and highlighting highly relevant content, effectively improving the quality of the pushed content.

[0037] A further step involves using a multimodal similarity calculation model to calculate the matching degree value between each user interest profile vector and each video content feature vector, based on the aligned user interest profile vector set and the video content feature vector set, to generate a preliminary matching degree matrix. This process includes the following steps:

[0038] The cosine similarity and Euclidean distance between each user interest profile vector and each video content feature vector are calculated separately, and the Euclidean distance is normalized. The cosine similarity and the normalized Euclidean distance are then fused using a weighted linear combination to generate a preliminary matching degree matrix. The weighted linear combination uses preset weight coefficients, which include a cosine similarity weight coefficient and an Euclidean distance weight coefficient, and the cosine similarity weight coefficient is greater than the Euclidean distance weight coefficient.

[0039] By adopting the above technical solution, cosine similarity and Euclidean distance values ​​are calculated separately, and the Euclidean distance is normalized, making the two measurement methods numerically comparable. By fusing the two similarity values ​​through a weighted linear combination, and setting the cosine similarity weight coefficient to be greater than the Euclidean distance weight coefficient, the dominant role of directional similarity in matching is emphasized, while also taking distance information into account. This makes the matching degree calculation more in line with the actual needs of user-video matching, generating a more accurate and stable preliminary matching degree matrix, and providing high-quality input for subsequent calibration and sorting.

[0040] A further step involves performing confidence calibration on the initial matching matrix to generate a calibrated matching matrix, specifically including the following steps:

[0041] Obtain a set of user historical feedback data and a set of video popularity data; wherein, the set of user historical feedback data is extracted based on the user's historical viewing records to show user skipping behavior and full viewing behavior data for videos; the set of video popularity data includes the number of times the video has been played, shared, favorited, and commented on;

[0042] Based on the user historical feedback data set and the video popularity data set, the feedback calibration coefficient and popularity calibration coefficient of each user-video matching pair are calculated respectively; wherein, the feedback calibration coefficient is obtained by weighted averaging of the corresponding user's historical feedback behavior on the same type of video, and the popularity calibration coefficient is obtained by normalizing and weighted fusion of the number of times the corresponding video is played, shared, favorited, and commented.

[0043] The feedback calibration coefficient and the popularity calibration coefficient are fused together by a weighted combination to generate a comprehensive calibration coefficient for each user-video matching pair;

[0044] Each matching degree value in the preliminary matching degree matrix is ​​multiplied by the corresponding comprehensive calibration coefficient to obtain the calibrated matching degree value, and a calibrated matching degree matrix is ​​generated.

[0045] The calibrated matching degree matrix is ​​smoothed; wherein the smoothing process employs a moving average algorithm or a Gaussian filtering algorithm to eliminate abnormal fluctuations in the matching degree values.

[0046] By adopting the above technical solution, rich contextual information is provided for matching calibration by acquiring historical user feedback data and video popularity data. The matching degree is corrected from both individual user behavior and group behavior dimensions by calculating feedback calibration coefficients and popularity calibration coefficients, making the calibration process both personalized and socially validated. A comprehensive calibration coefficient is generated through weighted combination, achieving effective fusion of multi-source calibration signals. Multiplying the initial matching degree value by the comprehensive calibration coefficient generates a calibrated matching degree matrix, significantly improving the confidence level and actual relevance of the matching degree. Smoothing the calibrated matrix eliminates abnormal fluctuations, making the matching results more stable and reliable, ultimately improving the accuracy of the push system and user trust.

[0047] A further setting is that S400 specifically includes the following steps:

[0048] The user-video matching confidence matrix is ​​processed to generate an initial push sequence. Then, based on a diversity control strategy, the initial push sequence is processed to enhance content diversity, generating a diversity-enhanced push sequence. The diversity control strategy includes limiting the number of videos belonging to the same interest dimension in the initial push sequence and introducing videos across interest dimensions to improve the content coverage of the push sequence.

[0049] The diversity enhancement push sequence is delivered to the user terminal in real time through the content delivery network interface.

[0050] By employing the above technical solutions, the push sequence generation process based on the user-video matching confidence matrix ensures that highly matched content is prioritized. Enhancing the content diversity of the initial sequence through a diversity control strategy—including limiting the number of videos within the same interest dimension and introducing videos across interest dimensions—effectively avoids homogenization and repetition in recommendation results, improving the breadth and exploratory nature of the push sequence's content coverage. This allows users to access a wider range of knowledge domains, satisfying their potential interests and expansion needs. Real-time generation of the sequence via a content distribution network interface guarantees the timeliness and efficiency of the push, improving the user experience.

[0051] A further setting is that S500 specifically includes the following steps:

[0052] Real-time collection of user feedback behavior data on pushed videos; wherein, the feedback behavior data includes viewing completion rate, interaction behavior type and frequency, and explicit feedback rating data; the viewing completion rate is the ratio of the user's actual viewing time to the total video duration; the interaction behavior type includes at least one of liking, collecting, commenting, and sharing; the explicit feedback rating data is the numerical value of the user's active rating of the video;

[0053] The feedback behavior data is subjected to multi-dimensional quantization processing to generate a feedback intensity vector set; wherein, the feedback intensity vector set includes the feedback intensity value corresponding to each pushed video, and the feedback intensity value is obtained by assigning different weights to different types of feedback behavior and summing them up by weight;

[0054] Based on the feedback intensity vector set and the user-video matching confidence matrix, the user interest profile vector set is dynamically updated; wherein, the dynamic update process includes adjusting the weights of the corresponding user interest dimensions according to the feedback intensity values, and recalculating and normalizing the user interest profile vectors.

[0055] Based on the updated user interest profile vector set, an updated user-video matching confidence matrix and push sequence are generated, and the updated push sequence is sent to the user terminal in real time through the content delivery network interface.

[0056] By adopting the above technical solution, real-time collection of user feedback behavior data, including viewing completion rate, interaction type, and explicit rating, comprehensively captures users' actual reactions to pushed content. Multi-dimensional quantification of feedback data generates a feedback intensity vector set, unifying different feedback behaviors into calculable indicators to provide a basis for model updates. Based on the feedback intensity vector and the original matching matrix, the user interest profile vector set is dynamically updated, including adjusting the weights of interest dimensions and recalculating interest vectors, enabling the interest model to quickly adapt to changes in user interests and maintain timeliness and accuracy. A new matching matrix and push sequence are generated based on the updated interest model and delivered to the user's terminal in real time, achieving closed-loop optimization and continuous iteration of the push strategy. Ultimately, this gives the recommendation system the ability to evolve and continuously improve push performance.

[0057] In summary, the present invention has the following beneficial effects: it can achieve accurate perception and deep semantic matching of users' dynamic learning interests, significantly improve the accuracy, timeliness and user satisfaction of content distribution, thereby optimizing learning efficiency and experience. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the main process of an embodiment;

[0059] Figure 2 This is a flowchart illustrating S100 in the embodiment;

[0060] Figure 3 This is a flowchart illustrating S200 in the embodiment;

[0061] Figure 4 This is a flowchart illustrating S300 in the embodiment;

[0062] Figure 5 This is a flowchart illustrating S400 in the embodiment;

[0063] Figure 6 This is a flowchart illustrating S500 in the embodiment. Detailed Implementation

[0064] The present invention will be further described in detail below with reference to the accompanying drawings.

[0065] like Figures 1 to 6 As shown;

[0066] This embodiment discloses an intelligent online video content recommendation method that combines a learning interest model, including the following steps:

[0067] S100. Construct dynamic interest vectors based on multi-source user behavior data. Extract and fuse multimodal features from user historical viewing records, search keywords, and interaction behavior data to generate a user interest profile vector set. Then, cluster and assign weights to the user interest profile vector set based on interest dimensions. Specific steps include:

[0068] Real-time acquisition of multi-source user behavior data is the foundation of the entire intelligent push method. In practice, multi-source user behavior data is first collected through user terminal interfaces and server log systems. This includes user viewing history, search keyword sets, and interaction behavior data sets. User viewing history includes identification information (such as video ID) and viewing duration (in seconds) of at least one viewed video. The search keyword set originates from keyword sequences entered by the user in multiple search sessions, with each sequence containing the keyword and its timestamp. The interaction behavior data set includes at least one of the user's actions such as liking, favorites, comments, and sharing of videos. This data is stored in a structured database to ensure real-time performance and completeness. The collection process employs a mechanism combining periodic and event-triggered methods: periodically collecting data from each... Execution is performed every minute, and event-triggered data is updated instantly based on real-time user interactions (such as clicking search or video operations). After data collection, preprocessing is performed, including noise reduction (removing invalid or duplicate records) and normalization (unifying the numerical range of different data sources to the [0,1] interval) to eliminate noise interference and improve the accuracy of subsequent processing.

[0069] Next, multimodal feature extraction is performed on multi-source user behavior data to generate viewing behavior feature vectors, search intent feature vectors, and interaction behavior feature vectors. The viewing behavior feature vector is achieved through embedding encoding and temporal modeling of video identifiers and viewing durations from the user's historical viewing records: First, embedding encoding techniques (such as Word2Vec) are used to map video identifiers to a 256-dimensional dense vector space to capture the semantic relationships of the video content; simultaneously, viewing duration is quantified as a time-weighted factor (e.g., videos with a viewing duration exceeding 80% are given higher weight), and temporal modeling methods (such as LSTM) are used to process the user's viewing sequence to generate a temporal feature vector. The specific formula is as follows: ;in, Mathematical symbols representing the feature vector of viewing behavior; The embedding vector for the video identifier. For the weighting coefficients based on viewing time, specifically ;in, Indicates the actual viewing time. The total video duration is represented by a 128-dimensional viewing behavior feature vector output by the LSTM network. The search intent feature vector is generated through semantic encoding and attention weighting of the search keyword sequence: semantic encoding uses a pre-trained language model (such as FastText) to transform the keyword sequence into a word vector sequence; the attention weighting mechanism assigns weights based on the importance of the keywords (e.g., high-frequency or recent keywords are given higher weights), as shown in the formula: ;in , representing the semantic encoding vector of the keyword. For attention weights, specifically ;in, Indicates an index; This represents the indexed score of all keywords, calculated based on keyword frequency and timeliness. Indicates the normalized denominator; It is an index variable that iterates through every keyword in the keyword set. The interactive behavior feature vector is achieved through one-hot encoding and weight allocation for different types of interactive behaviors: one-hot encoding encodes likes, favorites, comments, and shares into sparse vectors, with each operation corresponding to an independent dimension; weight allocation is based on operation type and frequency (for example, comments are assigned a weight of 0.4 due to high user engagement, likes 0.3, favorites 0.2, and shares 0.1), using the following formula: ;in, Indicates the type of interactive behavior; This indicates the preset weighting coefficient; The output is a binary encoded vector. After feature extraction, the dimensions of each vector are unified to 128 to ensure compatibility for subsequent fusion. Then, the generated viewing behavior feature vector, search intent feature vector, and interaction behavior feature vector are fused across modalities to generate a user interest profile vector set. The fusion process uses a combination of self-attention mechanism and fully connected layers: First, the self-attention mechanism aligns feature vectors from different sources and calculates the correlation weights between features (e.g., evaluating the interaction importance of viewing behavior and search intent through a multi-head attention layer). After alignment, the feature vectors are integrated into a temporary fusion vector through weighted concatenation (e.g., the weight coefficients are dynamically adjusted based on the feature variance). Next, a fully connected layer (containing a ReLU activation function) performs a non-linear transformation and dimensionality reduction on the temporary fusion vector, outputting a unified user interest representation vector with a dimension of 256. The fusion process is executed in real time on a distributed computing framework (such as TensorFlow), with the time consumption controlled within 100ms to ensure low latency.

[0070] Finally, interest dimension clustering analysis is performed on the user interest profile vector set. An unsupervised clustering algorithm (such as K-means) is used to divide user interests into multiple interest dimensions (e.g., the default number of clusters is 8, corresponding to topics such as "technology," "entertainment," and "education"): First, the clustering algorithm iteratively optimizes the cluster centers based on Euclidean distance to measure vector similarity; then, an initial weight is assigned to each interest dimension. The weight calculation is based on the frequency and timeliness of the user behavior data contained in that interest dimension, as shown in the formula: ;in, Indicates the initial weights; Normalized values ​​representing the frequency of user behavior data; This represents the frequency weighting factor, with a value of 0.6. As a timeliness factor, based on the time decay function ,in, Indicates the time difference in the occurrence of the behavior. This represents the attenuation coefficient. Clustering results are stored in an in-memory database, supporting rapid retrieval and updates. The entire process takes 10 minutes, and combined with historical data backtesting (the last 30 days), it ensures the timeliness and accuracy of user interest profiles, laying the foundation for subsequent video recommendations.

[0071] S200. Based on the clustering results of the user interest profile vector set, perform deep semantic analysis and tag mapping on the video content in the online video library, extract the theme features, knowledge structure features, and sentiment features of the video content, and generate a video content feature vector set; specifically including the following steps:

[0072] The first step in implementation is to acquire multimodal video content data from an online video library. This data encompasses three main categories: video metadata sets, subtitle text sets, and audio feature sets. The video metadata sets include the complete title of each video, detailed description text, and a multi-level classification tag system (such as "technology"). AI The video audio stream contains machine learning data, uploader identity information, and a posting timestamp accurate to the second; the subtitle text set is derived from transcribed text automatically generated by the video speech recognition module or finely edited subtitle text added manually, containing time-aligned sentence sequences; the audio feature set is extracted from the video audio stream using Mel spectrum analysis tools, specifically including 128-dimensional Mel frequency cepstral coefficient acoustic features, speech emotion classification features based on convolutional neural networks (such as being divided into eight emotion types including excitement, calm, and serious), and the rhythmic patterns and instrument composition features of the background music.

[0073] When extracting multi-dimensional features from the aforementioned multimodal video content data, a parallel processing pipeline is employed: The topic feature vector generation module first performs LDA topic modeling on the subtitle text set to identify the core topic distribution of the video. Simultaneously, it extracts keyword feature vectors using a TF-IDF weighted algorithm, and finally fuses them into a 256-dimensional topic feature vector through a fully connected layer. The knowledge structure feature vector construction module focuses on analyzing the topological relationships of knowledge points in the video metadata set. It uses knowledge graph embedding technology (such as TransE) to map conceptual entities (such as "linear regression" and "gradient descent") in the video description as knowledge nodes, and parses their logical dependencies to form a directed graph structure. Finally, it generates a 512-dimensional knowledge structure feature vector representing the depth of knowledge through a graph neural network (GNN). The sentiment tendency feature vector generation channel employs a dual-path processing mechanism: the audio path performs temporal pooling processing on the speech sentiment features to extract key inflection point features from the sentiment intensity waveform; the text path performs sentiment polarity analysis on the subtitle text (using a BERT model to output positive / negative scores). The two feature paths are integrated into a 128-dimensional sentiment tendency feature vector through a gated fusion unit. In the cross-modal feature fusion stage, topic feature vectors, knowledge structure feature vectors, and sentiment feature vectors are input into a feature weighting system based on a multi-head attention mechanism. This system first calculates the correlation matrix between different feature modalities: for example, when a video shows a high-intensity "excitement" type of sentiment feature, its interaction weight with the "entertainment" type of topic feature is automatically increased. The aligned feature tensors are then dynamically weighted and concatenated (the weight coefficients for knowledge structure features are set to 0.5 by default, topic features 0.3, and sentiment features 0.2), generating a 1024-dimensional temporary fusion vector. This is then processed through two fully connected layers with ReLU activation functions for dimensionality reduction, outputting a 256-dimensional unified video content representation vector. This vector encapsulates the video's semantic core, knowledge depth, and sentiment tone, among other comprehensive attributes. This process runs on a distributed computing framework, with single-video processing time controlled within 200 milliseconds. Based on the clustering results of the user interest profile vector set (such as the eight established interest dimensions like "programming education" and "film and television entertainment"), a label mapping operation is performed on the video content feature vector set. A cosine similarity matrix is ​​established between the center vector of the interest dimension and the feature vector of the video content. When the similarity between a video and the "Programming Education" dimension exceeds the threshold of 0.75, the interest tag is automatically associated and a matching weight is calculated. The formula for calculating the weight is: ;in, Indicates the matching degree weight; Represents the feature vector of video content; Represents the center vector of the interest dimension; The similarity function is represented by cosine similarity. This represents the sum of the exponential similarities between the current video and the center vectors of all eight interest dimensions; This represents the historical frequency factor of this dimension within the user group. The mapping results form a structured tag storage system, with each video carrying 1-3 interest tags and their precise matching weights, providing content feature basis for subsequent push decisions. The above process is executed in a rolling cycle of 10 minutes, with incremental feature extraction triggered in real time when new content is added to the video library.

[0074] S300. Establish a multi-dimensional association between the user interest profile vector set and the video content feature vector set, use a similarity calculation model to perform vector space alignment and matching degree calculation, and output a user-video matching confidence matrix; specifically including the following steps:

[0075] During implementation, the first step is vector space alignment. This step aims to map each user interest profile vector in the user interest profile vector set to each video content feature vector in the video content feature vector set into the same vector space, eliminating dimensionality differences and scale inconsistencies. The specific process is as follows: The user interest profile vector set is a set of feature vectors extracted from multi-source user behavior data (such as historical viewing records, search keywords, and interaction behaviors), with each vector having 256 dimensions; the video content feature vector set originates from multimodal video content data (such as video metadata, subtitle text, and audio features), with each vector also having 256 dimensions. To align the two, the system employs dimensionality mapping technology, using Principal Component Analysis (PCA) to reduce the high-dimensional vectors to a unified 128-dimensional space, ensuring the effectiveness of subsequent similarity calculations. Simultaneously, scale normalization is performed using the min-max normalization method, unifying the value range of each feature vector to a uniform dimensionality. This process, executed on a distributed computing framework, takes less than 50 milliseconds to ensure real-time performance. After alignment, the user interest profile vector set and the video content feature vector set are compatible, laying the foundation for matching degree calculation. Next, based on the aligned vector set, a preliminary matching degree matrix is ​​generated using a multimodal similarity calculation model. The multimodal similarity calculation model integrates cosine similarity and Euclidean distance, calculating the matching degree value between each user interest profile vector and each video content feature vector through a weighted linear combination. Specifically, for each pair of user-video vectors, the cosine similarity value and Euclidean distance value are calculated separately. The cosine similarity value measures the directional consistency between vectors; the Euclidean distance value represents the spatial distance between vectors. To unify the units, the Euclidean distance value needs to be normalized, and after normalization, a matching degree value is generated through a weighted linear combination. ;in, This represents the user-video matching score. The higher the score, the better the video j matches the user i's interest profile. The weighting coefficients representing cosine similarity; Represents the cosine similarity value; The weighting coefficients represent the Euclidean distance; This represents the normalized Euclidean distance value. Preferably, , To ensure consistency in direction, the combined matching score should be prioritized. The calculation results of all user-video pairs form a preliminary matching degree matrix, with the dimension being the number of users × the number of videos. This process is executed in parallel under GPU acceleration, with a single batch processing of 1000 vector pairs taking approximately 10 milliseconds, ensuring high efficiency.

[0076] Subsequently, confidence calibration is performed to improve the accuracy and robustness of the matching matrix. Calibration is based on a set of historical user feedback data and a set of video popularity data, generating a comprehensive calibration coefficient to correct the initial matching value. The set of historical user feedback data includes user skipping behavior (e.g., completion rate below 20%), complete viewing behavior (completion rate above 80%), and partial viewing behavior (completion rate between 20% and 80%), extracted from users' historical viewing records. The set of video popularity data includes the number of times a video is played, shared, favorited, and commented, queried in real-time from a database. Specific steps: First, the feedback calibration coefficient for each user-video matching pair is calculated. This coefficient is based on a weighted average of the corresponding user's historical feedback behavior on similar videos, specifically expressed by the formula: ;in, Indicates the feedback calibration coefficient; Indicates the first Records of user's historical feedback behavior; This indicates the total number of selected user historical feedback behavior records; This represents the historical feedback behavior value (e.g., skipping behavior is assigned a value of 0, watching the whole thing is assigned a value of 1, and partial watching behavior is assigned a value using linear interpolation). This represents the time-sensitivity weight, whose value is based on the time decay function. Calculated; where, This represents the attenuation coefficient, with a value of 0.1. This represents the time difference of the behavior. Next, the popularity calibration coefficient is calculated: the number of times the video is played, shared, favorited, and commented are normalized and then weighted and merged; preferably, the weights are preset as follows: play count weight coefficient 0.4, share count weight coefficient 0.3, favorite count weight coefficient 0.2, and comment count weight coefficient 0.1. Then, a comprehensive calibration coefficient is generated through fusion: ;in, Indicates the overall calibration coefficient; Indicates the popularity calibration coefficient; Represents the fusion weight, with values ​​ranging from 0 to 10. The focus is on user feedback. Then, each matching value in the initial matching matrix is ​​multiplied by its corresponding comprehensive calibration coefficient: ;in, The calibration results in the corrected matching score. Next, a calibrated matching score matrix is ​​generated. Finally, a smoothing process is performed using a moving average algorithm: a moving average with a window size of 3 is applied to the matrix rows (user dimension) to eliminate abnormal fluctuations. The calibration process is executed in real-time in an in-memory database with a latency of no more than 20 milliseconds.

[0077] Finally, threshold filtering and sorting optimization are performed to output a user-video matching confidence matrix. A preset threshold filtering is applied to the calibrated matching matrix: entries with a matching score below 0.5 are directly removed to ensure only high-confidence matching pairs are retained. The remaining matching pairs are sorted in descending order of matching score. After optimization, the matrix is ​​converted into a structured data format and stored in a distributed cache for fast retrieval. The entire process takes approximately [time period missing]. The system operates on a second-by-second basis, combining real-time data streams to ensure the timeliness and accuracy of push notification decisions. After outputting the user-video matching confidence matrix, the system seamlessly transitions to the push sequence generation stage, completing closed-loop optimization.

[0078] S400. Based on a multi-objective optimization strategy, the user-video matching confidence matrix is ​​transformed into a push sequence, and the push sequence is sent to the user terminal in real time through a content delivery network interface; specifically including the following steps:

[0079] The system first processes the user-video matching confidence matrix to generate an initial push sequence. The initial push sequence is generated based on the matching scores in the user-video matching confidence matrix, selecting the top scores in descending order. The initial recommendation list consists of 10 videos, of which 10 videos make up the 10 videos. The system's preset number of push notifications is typically between 10 and 20. During the sorting process, the system prioritizes video entries with a matching confidence score higher than a set threshold (e.g., 0.7) to ensure that the pushed content is highly relevant to the user's interests. The initial push sequence is stored in a distributed cache in a structured data format, including fields such as video identifier, matching score, and video metadata, for easy subsequent processing.

[0080] Next, the system enhances the content diversity of the initial push sequence based on a diversity control strategy, generating a diversity-enhanced push sequence. The core objective of the diversity control strategy is to improve the breadth of content coverage in the push sequence while maintaining recommendation relevance, avoiding excessive concentration of push content on a specific interest dimension, thereby enhancing user experience and exploration. In practice, the system first performs cluster analysis on the videos in the initial push sequence according to their respective interest dimensions, identifying the dominant interest dimensions in the sequence (e.g., "science and education" videos account for more than 60%). Subsequently, the system applies a quantity limit mechanism, setting an upper limit on the number of videos within the same interest dimension (e.g., no more than 50% of the total number of pushes), and removing videos exceeding the limit based on their matching degree from low to high. Simultaneously, the system introduces videos across interest dimensions to supplement sequence diversity: based on the weight distribution of each interest dimension in the user interest profile vector set, videos with high matching degrees (usually requiring a matching degree of no less than 0.6) are selected from interest dimensions not covered by the initial sequence and inserted into the push sequence according to their weight proportions. During the insertion process, the system uses a greedy algorithm to optimize the insertion position, ensuring that the introduction of new videos does not significantly reduce the overall sequence relevance. The resulting diversity-enhanced push notification sequence retains highly relevant core content while also covering multiple interest dimensions, thus improving the richness of the content and users' exploration opportunities.

[0081] Finally, the system delivers the diversity-enhanced push sequence to user terminals in real time via the Content Delivery Network (CDN) interface. The CDN interface is designed based on a RESTful API, supporting high-concurrency, low-latency data transmission. The push sequence is encapsulated in a JSON-formatted message body, including metadata such as the video list, push identifier, and timestamp. During the push process, the system monitors the transmission status in real time. If a transmission failure or timeout is detected, a retransmission mechanism is automatically triggered to ensure reliable delivery of the push sequence. The entire push process is completed without the user's awareness, with push latency controlled within 100 milliseconds, ensuring the real-time nature of the recommendations and a smooth user experience.

[0082] In addition, the system records the delivery logs of the push sequence, including push time, user ID, and push content, for subsequent feedback collection and strategy optimization. The push logs are stored in a distributed database, supporting rapid querying and analysis, and providing data support for closed-loop optimization.

[0083] S500 collects user feedback on pushed videos in real time, and dynamically adjusts the user interest profile vector set and push strategy based on the feedback data to achieve closed-loop optimization; specifically including the following steps:

[0084] The system collects user feedback data on pushed videos in real time through the user terminal interface and the backend log system. This data includes viewing completion rate, interaction type and frequency, and explicit feedback rating data. The viewing completion rate is calculated as the ratio of the user's actual viewing time to the total video duration. Interaction types include at least one of the following: liking, favorites, comments, and sharing. Explicit feedback rating data comes from user-initiated ratings, typically ranging from 1 to 5. The data collection mechanism combines event triggering with periodic retrieval to ensure the real-time nature and completeness of the feedback data. The collected data undergoes preprocessing, including deduplication, invalid value removal, and normalization, for subsequent quantitative analysis. The preprocessed feedback behavior data is then subjected to multi-dimensional quantitative processing to generate a feedback intensity vector set. The feedback intensity value is calculated using a weighted summation formula.

[0085] ;in, Indicates the feedback strength value; Indicates the completion rate of viewing; This represents the normalized frequency of "like" actions. Normalized frequency value representing the act of collecting; This represents the normalized frequency of commenting behavior. This represents the normalized frequency of sharing behavior. This represents the normalized value of the explicit feedback rating data; , , , , and These represent the weighting coefficients for view completion rate, likes, favorites, comments, shares, and explicit rating data, respectively. In the actual quantification process, the system assigns differentiated weights to different types of feedback behavior, with the preferred weighting being... , , , , , Ultimately, each pushed video corresponds to a feedback intensity value, and the feedback values ​​of all videos constitute a feedback intensity vector set, which is used for dynamic adjustment of the user interest profile in the future.

[0086] The user interest profile vector set is dynamically updated based on the feedback intensity vector set and the user-video matching confidence matrix. The update process first adjusts the weights of the corresponding user interest dimensions according to the feedback intensity values: if a video under a certain interest dimension receives a high feedback intensity, the weight of that dimension is increased; conversely, it is decreased. After weight adjustment, the system recalculates the user interest profile vectors by weighted summation of the vectors of each interest dimension and applying L2 normalization to generate updated user interest profile vectors. This process ensures that the user interest representation can reflect the latest changes in preferences in a timely manner, improving the timeliness and accuracy of the profile. Based on the updated user interest profile vector set, the system re-executes vector space alignment and similarity calculation to generate an updated user-video matching confidence matrix. This matrix, after threshold filtering and sorting optimization, is transformed into a push sequence and distributed to user terminals in real time through the content delivery network interface. During the push process, the system uses an asynchronous message queue and multi-threaded concurrency mechanism to ensure that the push latency is less than 100 milliseconds, guaranteeing a smooth user experience. Simultaneously, the system records a complete log of the push and feedback process for subsequent model iteration and strategy optimization, forming a complete closed-loop learning mechanism.

[0087] The data acquisition, processing, and implementation of the technical solutions in this invention strictly comply with relevant laws and regulations such as the "Cybersecurity Law of the People's Republic of China," the "Data Security Law of the People's Republic of China," and the "Personal Information Protection Law of the People's Republic of China," and conform to socialist public order and good morals. Regarding user data collection and use, the "informed consent" principle is consistently upheld, clearly informing users of the purpose, method, and scope of data use and obtaining their explicit authorization. Strict technical and management measures are adopted to protect user privacy, prevent the leakage, alteration, or loss of user information, and ensure the legal, compliant, and ethical application of the technical solutions.

[0088] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.

Claims

1.A method for intelligent pushing of online video content combined with a learning interest model, characterized in that, The method comprises the following steps: S100, constructing a dynamic interest vector based on multi-source user behavior data, performing multi-modal feature extraction and fusion on user historical viewing records, search keywords and interaction behavior data, generating a user interest portrait vector set, and performing interest dimension clustering and weight distribution on the user interest portrait vector set; S200, according to the clustering result of the user interest portrait vector set, performing deep semantic analysis and label mapping on the video content in the online video library, extracting the theme feature, knowledge structure feature and emotional tendency feature of the video content, and generating a video content feature vector set; S300, establishing a multi-dimensional association relationship between the user interest portrait vector set and the video content feature vector set, performing vector space alignment and matching degree calculation by using a similarity calculation model, and outputting a user-video matching confidence matrix; comprising: performing vector space alignment processing on the user interest portrait vector set and the video content feature vector set; wherein the vector space alignment processing comprises dimension mapping and scale normalization of each user interest portrait vector in the user interest portrait vector set and each video content feature vector in the video content feature vector set, so that the user interest portrait vector set and the video content feature vector set are in the same vector space; based on the aligned user interest portrait vector set and the video content feature vector set, a multi-modal similarity calculation model is used to calculate the matching degree value between each user interest portrait vector and each video content feature vector, and a preliminary matching degree matrix is generated; wherein the multi-modal similarity calculation model fuses cosine similarity and Euclidean distance two measurement methods, and generates a preliminary matching degree matrix through weighted linear combination; calibrating the confidence of the preliminary matching degree matrix to generate a calibrated matching degree matrix; threshold filtering and sorting optimization are performed on the calibrated matching degree matrix to generate a final user-video matching confidence matrix; wherein threshold filtering is used to eliminate user-video matching pairs with a matching confidence score lower than a preset threshold, and then the remaining matching pairs are arranged in descending order according to the matching degree value; S400, based on the multi-objective optimization strategy, the user-video matching confidence matrix is converted into a push sequence, and the push sequence is real-time delivered to the user terminal through the content distribution network interface; comprising: performing push sequence generation processing on the user-video matching confidence matrix to generate an initial push sequence, and then performing content diversity enhancement processing on the initial push sequence based on a diversity control strategy to generate a diversity enhanced push sequence; wherein the diversity control strategy comprises limiting the number of videos belonging to the same interest dimension in the initial push sequence, and introducing cross-interest dimension videos to improve the content coverage of the push sequence; the diversity enhanced push sequence is real-time delivered to the user terminal through the content distribution network interface; S500, real-time collection of user feedback behavior on the pushed video, dynamic adjustment of user interest portrait vector set and push strategy based on feedback data, to realize closed-loop optimization. 2.The online video content intelligent pushing method combined with a learning interest model according to claim 1, characterized in that, The S100 specifically comprises the following steps: Obtaining multi-source user behavior data; wherein the multi-source user behavior data includes user historical viewing records, a search keyword set, and an interactive behavior data set; the user historical viewing records include identification information and viewing duration of at least one video that has been watched; the search keyword set includes a keyword sequence input by a user in one or more search sessions; and the interactive behavior data set includes at least one of a like, a collection, a comment, and a share operation of a user on a video; Multi-modal feature extraction is performed on the multi-source user behavior data to respectively generate a viewing behavior feature vector, a search intent feature vector, and an interactive behavior feature vector; Cross-modal feature fusion is performed on the viewing behavior feature vector, the search intent feature vector, and the interactive behavior feature vector to generate a user interest portrait vector set; wherein the cross-modal feature fusion adopts a combination of a self-attention mechanism and a fully connected layer to align and weight splice different types of feature vectors to generate a unified user interest representation vector; Interest dimension clustering analysis is performed on the user interest portrait vector set, and an unsupervised clustering algorithm is used to divide user interests into multiple interest dimensions, and an initial weight is assigned to each interest dimension; wherein the initial weight is determined based on the frequency and timeliness of the user behavior data contained in the interest dimension. 3.The online video content intelligent pushing method with learning interest model according to claim 2, characterized in that, Multi-modal feature extraction is performed on the multi-source user behavior data to respectively generate a viewing behavior feature vector, a search intent feature vector, and an interactive behavior feature vector, specifically including the steps of: The viewing behavior feature vector is obtained by embedding encoding and time sequence modeling of the video identification and viewing duration in the user historical viewing records; the search intent feature vector is obtained by semantic encoding and attention weighting of the search keyword sequence; and the interactive behavior feature vector is obtained by one-hot encoding and weight distribution of different types of interactive behaviors. 4.The online video content intelligent pushing method combined with a learning interest model according to claim 1, characterized in that, The S200 specifically includes the steps of: Obtaining multi-modal video content data in an online video library; wherein the multi-modal video content data includes a video metadata set, a subtitle text set, and an audio feature set; the video metadata set includes a title, a description, a classification label, uploader information, and a release time of at least one video; the subtitle text set includes transcription text corresponding to the video or manually added subtitle text; and the audio feature set includes acoustic features, speech emotion features, and background music features extracted from a video audio stream; Multi-dimensional feature extraction is performed on the multi-modal video content data to respectively generate a theme feature vector, a knowledge structure feature vector, and an emotional tendency feature vector; Cross-modal feature fusion is performed on the theme feature vector, the knowledge structure feature vector, and the emotional tendency feature vector to generate a video content feature vector set; wherein the cross-modal feature fusion adopts a feature weighting and splicing method based on an attention mechanism to align and integrate different source feature vectors to generate a unified video content representation vector; Based on the clustering result of the user interest portrait vector set, the video content feature vector set is mapped with labels; wherein, the label mapping process matches the video content feature vector with the interest dimension, and assigns one or more interest labels and their matching degree weights to each video. 5.The online video content intelligent pushing method with learning interest model according to claim 4, characterized in that, Multi-dimensional feature extraction is performed on the multi-modal video content data to generate a theme feature vector, a knowledge structure feature vector and an emotional tendency feature vector, specifically including the steps of: The theme feature vector is obtained by theme modeling and keyword extraction on the subtitle text set; the knowledge structure feature vector is obtained by analyzing the knowledge point distribution and logical relationship of the video metadata set; and the emotional tendency feature vector is obtained by emotional analysis on the audio feature set and the subtitle text set. 6.The online video content intelligent pushing method with learning interest model according to claim 1, wherein, Based on the aligned user interest portrait vector set and the video content feature vector set, a multi-modal similarity calculation model is used to calculate the matching degree value between each user interest portrait vector and each video content feature vector, and a preliminary matching degree matrix is generated, specifically including the steps of: The cosine similarity value and the Euclidean distance value between each user interest portrait vector and each video content feature vector are calculated respectively, and the Euclidean distance value is normalized; the cosine similarity value and the normalized Euclidean distance value are fused by weighted linear combination to generate a preliminary matching degree matrix; wherein, the weighted linear combination uses a preset weight coefficient, and the weight coefficient includes a cosine similarity weight coefficient and a Euclidean distance weight coefficient, and the cosine similarity weight coefficient is greater than the Euclidean distance weight coefficient. 7.The online video content intelligent pushing method with learning interest model according to claim 1, wherein, The preliminary matching degree matrix is subjected to confidence calibration to generate a calibrated matching degree matrix, specifically including the steps of: A user historical feedback data set and a video popularity data set are obtained; wherein, the user historical feedback data set is extracted based on the user historical viewing records to obtain the skipping behavior and complete viewing behavior data of the user for the video; and the video popularity data set includes the number of plays, the number of shares, the number of collections and the number of comments of the video; Based on the user historical feedback data set and the video popularity data set, the feedback calibration coefficient and the popularity calibration coefficient of each user-video matching pair are calculated respectively; wherein, the feedback calibration coefficient is obtained by weighted averaging the historical feedback behavior of the corresponding user for similar videos, and the popularity calibration coefficient is obtained by normalizing and weighted fusion of the number of plays, the number of shares, the number of collections and the number of comments of the corresponding video; The feedback calibration coefficient and the popularity calibration coefficient are fused by weighted combination to generate the comprehensive calibration coefficient of each user-video matching pair; Each matching degree value in the preliminary matching degree matrix is multiplied by the corresponding comprehensive calibration coefficient to obtain the calibrated matching degree value, and a calibrated matching degree matrix is generated; The calibrated matching degree matrix is subjected to smoothing processing; wherein, the smoothing processing adopts a sliding average algorithm or a Gaussian filtering algorithm to eliminate abnormal fluctuations of the matching degree value. 8.The online video content intelligent pushing method with learning interest model according to claim 1, wherein, The S500 specifically includes the steps of: Real-time collection of user feedback behavior data on the pushed video; wherein the feedback behavior data includes a watching completion rate, an interaction behavior type and frequency, and explicit feedback score data; the watching completion rate is the ratio of the actual watching time of the user to the total time of the video; the interaction behavior type includes at least one of likes, collections, comments, and shares; the explicit feedback score data is the numerical value of the user's active scoring on the video; Multi-dimensional quantitative processing of the feedback behavior data to generate a feedback intensity vector set; wherein the feedback intensity vector set includes a feedback intensity value corresponding to each pushed video, which is obtained by assigning different weights to different types of feedback behavior and weighted summation; Based on the feedback intensity vector set and the user-video matching confidence matrix, the user interest portrait vector set is dynamically updated; wherein the dynamic updating process includes adjusting the weight of the corresponding user interest dimension according to the feedback intensity value, and recalculating and normalizing the user interest portrait vector; According to the updated user interest portrait vector set, an updated user-video matching confidence matrix and a push sequence are generated, and the updated push sequence is real-time delivered to the user terminal through the content distribution network interface.

Citation Information

Patent Citations

  • Content recommendation method and system based on semantic recognition

    CN119089398A

  • Video feature extraction and multi-dimensional matching-based movie advertisement real-time pushing system

    CN120181922A