Video recommendation method, feature processing model training method, device and equipment
Through the feature processing model, the semantic feature sequence and twin network training are generated, which solves the problems of low recommendation efficiency and insufficient interest recognition in the video recommendation system, and achieves more accurate and diverse video recommendations, improving user experience.
Patent Information
- Application Number
- CN202510353921.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing video recommendation system is inefficient in recommendation when facing a large number of videos to be selected, and it is difficult to accurately identify users' medium- and long-tail and long-term interests, resulting in insufficient recommendation accuracy and stability.
The feature processing model performs feature extraction of the user's historical viewing videos, maps them to the preset codebook collection to generate semantic feature sequences, and matches the videos to be recommended from the video library. The feature processing model is trained using the twin network model to improve the accuracy and diversity of video recommendations.
It improves the accuracy and user experience of video recommendations, ensures that the recommended content covers the main interests and long-tail interests of users, improves the diversity and coverage of recommendations, and meets the complex and dynamic needs of users.
Smart Images

Figure CN120336581A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to intelligent search in the field of artificial intelligence, and particularly to a video recommendation method, a training method of a feature processing model, an apparatus, and a device. Background Art
[0002] In current video recommendation systems, user features can generally be generated for each user, and then the similarity is calculated based on the user features of the user and the video features of the candidate videos, and the recommended videos are selected from the candidate videos according to the similarity.
[0003] However, in the face of a large number of candidate videos, this method has the problem of low recommendation efficiency. Summary of the Invention
[0004] The present disclosure provides a video recommendation method, a training method of a feature processing model, an apparatus, and a device for improving the accuracy of video recommendation and enhancing the user experience.
[0005] According to a first aspect of the present disclosure, there is provided a video recommendation method, including:
[0006] Performing feature extraction on the historical watched videos of a user based on a feature processing model to obtain video features of the historical watched videos;
[0007] Mapping the video features to a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical watched videos; wherein, the preset codebook set includes a plurality of codebook vectors, and the codebook vectors represent semantic features of videos;
[0008] Determining candidate recommended videos matching the semantic feature sequence of the historical watched videos from a video library; and recommending the candidate recommended videos to the user.
[0009] According to a second aspect of the present disclosure, there is provided a training method of a feature processing model, including:
[0010] Inputting a first training set into a first feature processing model to obtain first video information features of samples in the first training set; inputting a second training set into a second feature processing model to obtain second video information features of samples in the second training set; wherein, the samples in the first training set include historical watched videos and historical unwatched videos; the first video information features represent the features after decoding the semantic feature sequences of the samples in the first training set; the samples in the second training set include historical watched videos and historical unwatched videos; the second video information features represent the features after decoding the semantic feature sequences of the samples in the second training set;
[0011] Based on the first video information feature and the second video information feature, train the first feature processing model and the second feature processing model to obtain the trained first feature processing model and second feature processing model;
[0012] Wherein, the trained first feature processing model, or the second feature processing model, is used to process the video in the first aspect and any possible method in the first aspect.
[0013] According to a third aspect of the present disclosure, there is provided a video recommendation device, including:
[0014] An extraction unit, configured to extract features of a user's historical viewed videos based on a feature processing model to obtain video features of the historical viewed videos;
[0015] A mapping unit, configured to map the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical viewed videos; wherein, the preset codebook set includes a plurality of codebook vectors, and the codebook vectors represent semantic features of videos;
[0016] A recommendation unit, configured to determine a video to be recommended that matches the semantic feature sequence of the historical viewed videos from a video library; and recommend the video to be recommended to the user.
[0017] According to a fourth aspect of the present disclosure, there is provided a training device for a feature processing model, including:
[0018] An input module, configured to input a first training set into a first feature processing model to obtain first video information features of samples in the first training set; input a second training set into a second feature processing model to obtain second video information features of samples in the second training set; wherein, the samples in the first training set include historical viewed videos and historical unviewed videos; the first video information features represent features after decoding the semantic feature sequences of the samples in the first training set; the samples in the second training set include historical viewed videos and historical unviewed videos; the second video information features represent features after decoding the semantic feature sequences of the samples in the second training set;
[0019] A training module, configured to train the first feature processing model and the second feature processing model based on the first video information features and the second video information features to obtain the trained first feature processing model and second feature processing model; wherein, the trained first feature processing model, or the second feature processing model, is used to process the video in the first aspect and any possible method in the first aspect.
[0020] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0021] at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect or the second aspect.
[0022] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect or the second aspect.
[0023] According to a seventh aspect of the present disclosure, there is provided a computer program product, which includes: a computer program stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the method described in the first aspect or the second aspect.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0026] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0027] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0028] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0029] Figure 4 shows a histogram of a semantic feature sequence;
[0030] Figure 5 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0031] Figure 6 shows a training illustration of a feature processing model Figure 1 ;
[0032] Figure 7 shows a training illustration of a feature processing modelFigure 2 ;
[0033] Figure 8 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0034] Figure 9 is a schematic diagram according to the sixth embodiment of the present disclosure;
[0035] Figure 10 is a schematic diagram according to the seventh embodiment of the present disclosure;
[0036] Figure 11 is a schematic diagram according to the eighth embodiment of the present disclosure;
[0037] Figure 12 is a schematic diagram according to the ninth embodiment of the present disclosure;
[0038] Figure 13 shows an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0039] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0040] During the process of a user watching a video, the user usually has multiple points of interest. For example, user A may be interested in three fields: sports, food, and technology. These points of interest can reflect the diverse content needs of the user. Currently, in the field of video recommendation, common video recall models such as the dual-tower structure recall model are limited by their representation ability and usually can only recall the TOP1 or TOP2 points of interest of the user, with poor representation ability for the medium and long-tail interests of the user. Moreover, the commonly used online learning method for model training has the problem of catastrophic forgetting and can only focus on the short-term points of interest of the user, resulting in the loss of the user's long-term interests.
[0041] Among them, the medium and long-tail interest representation and long-term interest of users can better reflect the long-term interests of users. Moreover, the long-term interests of users reflect stable behavior patterns such as the fields that users continuously focus on and the types of content they prefer within a period of time, which is of great help in identifying the core needs of users and can avoid the interference of short-term behaviors on video recommendations. In addition, the video recommendation model generated based on long-term interests can also better smooth the fluctuations of short-term behaviors such as accidental clicks, improve the stability of video recommendations, and be more in line with user preferences. Also, the parallel capture of more user interest points can improve the richness of recommended videos based on multiple interest dimensions and avoid the over-simplification of the content of recommended videos.
[0042] Moreover, since users' interests change over time. Therefore, video recall based on a single interest point or a small number of interest points is likely to be unable to meet users' needs when users' interests change. While video recall based on multiple interest points can better ensure the proportion of videos that users are interested in among the recommended videos, thus leaving enough time for the optimization of the video recommendation model for this user and improving the user experience. Therefore, in the construction of video recommendation strategies, the long-term interests and multiple interests of users complement each other in the recall stage. The use of these two features can not only ensure the accuracy and stability of recommendations, but also enhance diversity and coverage, more accurately match the complex and dynamic needs of users, provide a high-quality candidate set for the subsequent ranking stage, and ultimately increase user activity and retention rate.
[0043] Currently, multi-interest modeling methods mainly include multi-interest extraction based on capsule networks (Multi-Interest Network with Dynamic routing, MIND), multi-interest modeling based on multi-head attention mechanisms (Controllable Multi-Interest Framework for Recommendation, ComiRec), and search-based interest modeling (Search-based Interest Modeling, SIM). Among them, MIND clusters the user behavior sequence into multiple interest vectors through a dynamic routing algorithm and is suitable for scenarios with complex user interest distributions. ComiRec uses the multi-head attention mechanism to generate multiple independent interest vectors and is suitable for short behavior sequence scenarios. SIM extracts multi-interests from long behavior sequences through two stages (hard search + soft search) and can model the long-term interests of users. These methods have improved the ability of multi-interest modeling to a certain extent, but there are still significant problems.
[0044] First, the ability to represent multiple interests is insufficient. Methods such as MIND and ComiRec automatically learn a user's multiple interests through specially designed network structures, but there is an interest coupling problem, that is, there will be overlap among the learned multiple interests. At the same time, the number of user interests usually needs to be preset in advance, making it difficult to adapt to the dynamic changes of user interests. Although the SIM method can model the long-term confident interests of users, in the recall model, an interest vector needs to be generated for each interest point for recall. Considering the computational complexity, usually only the TOP 5-10 interest points can be modeled. Second, the computational complexity is high. When the length of the user's historical consumption sequence becomes longer or the number of interests to be modeled increases, the model complexity of MIND, ComiRec, and SIM methods increases significantly. Usually, only a sequence length of thousands or the number of interests within 10 can be modeled. Finally, the ability to model long-term and long-tail interests is weak. The MIND and ComiRec methods mainly model the user's recent interests and have poor ability to model long-term and long-tail interests. Although the SIM method can model long-term interests, it is difficult to model the long-tail interests of users.
[0045] To solve the above problems, the present disclosure proposes a video recommendation method and a training method for a feature processing model. Based on the training method of the feature processing model, the training of the feature processing model can be realized. Furthermore, based on the trained feature processing model, the video recommendation in the video recommendation method can be realized, thereby improving the accuracy of video recommendation and enhancing the user experience. Specifically, the present disclosure can extract the semantic feature sequence of each video through the feature processing model. Based on this semantic feature sequence, the video recommendation method of the present disclosure can count the user's historical watched videos over a long period of time, so as to determine that the semantic feature sequence with a relatively large number of user views is the semantic feature sequence that the user is interested in. Based on the statistics of the number of views, the video recommendation method can select the videos corresponding to the semantic feature sequences that the user is interested in from the video library and recommend the videos to the user. In this recommendation process, according to the user's needs, multiple semantic feature sequences that the user is interested in can be selected in descending order of the number of views, so as to ensure the richness of the recommended content and avoid the decline of the user experience caused by the over-single type of recommended videos.
[0046] The present disclosure provides a video recommendation method, a training method for a feature processing model, an apparatus, and a device, which are applied to intelligent search in the field of artificial intelligence to achieve the effect of improving the accuracy of video recommendation and enhancing the user experience.
[0047] It should be noted that the video recommendation model in this embodiment is not a video recommendation model for a specific user and does not reflect the personal information of a specific user. It should be noted that the user's historical watched videos in this embodiment are data obtained after the user's authorization.
[0048] In the technical solution of the present disclosure, the processing of the user's personal information, such as collection, storage, use, processing, transmission, provision, and disclosure, complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0049] To enable readers to more deeply understand the implementation principle of the present disclosure, the following Figures 2 - 8 will be Figure 1 further refined for the
[0050] Figure 1 schematic diagram shown in accordance with the first embodiment of the present disclosure. As Figure 1 shown, the present disclosure provides a method for video recommendation, and the method includes:
[0051] 101. Extract features of the user's historical watched videos based on a feature processing model to obtain video features of the historical watched videos.
[0052] In this embodiment, the feature processing model is used to extract features from the encoded video information. Optionally, the feature processing model can be a convolutional neural network. For example, the feature processing model can be a transformer, DNN, CNN, etc. Based on this feature processing model, the extracted video features are obtained by extracting the encoded information of the historical watched videos, and are high-dimensional feature vectors that can represent comprehensive information such as visual features, text features, and audio features.
[0053] Optionally, the data input to the feature processing model is encoded video information. Optionally, the encoded video information can be data obtained by encoding the multi-modal information of the historical watched videos through a multi-modal encoder. Optionally, the multi-modal encoder can be provided with encoders for encoding data of each modality. Optionally, the data input to the feature processing model can be obtained by splicing after encoding the multi-modal information.
[0054] For example, when the modality is text information, the text information can be spliced from at least one of the text description information of the historical watched videos, image text recognition information, and audio text recognition information. The multi-modal encoder can encode the text information to obtain the data input to the feature processing model. Optionally, the text description information of the historical watched videos can be text description information edited by technicians for describing the videos. For example, the text description information can include the title, category, tags, introduction, etc. of the historical watched videos.
[0055] 102. Map the video features to a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical watched videos. Among them, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent the semantic features of the videos.
[0056] In this embodiment, after the video features of the user's historical watched videos are extracted through the feature processing model, the extracted video features can be further mapped to a preset codebook set of the feature processing model to obtain the semantic feature sequence of the historical watched videos. Optionally, the preset codebook set is a set containing multiple codebook vectors, and each codebook vector represents a specific video semantic feature. Optionally, the mapping process can be a residual quantization process. Optionally, through residual quantization, the video features are converted into a semantic feature sequence composed of multiple codebook vectors. The process of mapping the video features to this semantic feature sequence can be understood as the process of extracting the semantic features of the historical watched videos. The generation of this semantic feature sequence enables the video content to be represented in a more abstract form, facilitating subsequent matching and recommendation. Optionally, this semantic feature sequence reflects the semantic features of different resources that the user has come into contact with in the past.
[0057] 103. Determine the videos to be recommended that match the semantic feature sequence of the historical watched videos from the video library, and recommend the videos to be recommended to the user.
[0058] In this embodiment, after obtaining the semantic feature sequence of the user's historical watched videos, the semantic feature sequence can be matched with the semantic feature sequences of the videos in the video library. After determining the videos to be recommended that match the semantic feature sequence of the historical watched videos from the video library, these videos can be recommended to the user.
[0059] Among them, the video library is a database containing a large number of videos and their corresponding information. Optionally, the videos in the video library can be processed through the feature processing model shown in the above steps 101 and 102 to obtain the semantic feature sequence of each video. Optionally, since the above feature processing model is a pre-trained feature processing model. Therefore, after obtaining this feature processing model, each video in the video library can be processed first to obtain the semantic feature sequence corresponding to each video. The video library can record the semantic feature sequence of each video. Optionally, when a new video is written into the video library, the video can be processed using this feature processing model and then the video and the semantic feature sequence can be written into the video library together.
[0060] Optionally, the user may have multiple historical watched videos. Optionally, based on the user's multiple historical watched videos, the occurrence frequency of each semantic feature sequence can be counted, and then one or more semantic feature sequences with the highest occurrence frequency can be selected to match the videos in the video library. Optionally, since there are a large number of videos in the video library, each semantic feature sequence may correspond to multiple videos. Therefore, this matching process can be to select a certain number of videos from the videos in the video library that have the same semantic feature sequence as the historical watched video as the matching videos. Optionally, the selection of the videos can be random selection. Or, it can be selected after sorting according to information such as the video release time and release location. Optionally, the certain number can be a preset value.
[0061] In this embodiment, the feature processing model is used to extract the features of the user's historical watched videos to obtain the semantic feature sequences of the historical watched videos. And based on the semantic feature sequences of the user's historical watched videos, after obtaining the videos to be recommended by matching from the video library, the videos to be recommended are then sent to the user, so as to improve the accuracy of video recommendation, enable the recommended videos to cover the user's main interests and long-tail interests, enhance the diversity and coverage of the recommendation, and can more accurately meet the complex and dynamic needs of the user, improving the user experience and platform activity.
[0062] On the basis of the above first embodiment, the specific process of using the feature processing model in step 101 to extract the video features of the historical watched videos may include:
[0063] 1011. Obtain the multimodal information of the historical watched video. And use the multimodal encoder to process the multimodal information to obtain the feature vector of the historical watched video.
[0064] In this embodiment, first, the multimodal information of the historical watched video can be obtained. Optionally, the multimodal information includes at least one of the video's image information, text information, and audio information. Since the feature processing model of this application is a language large model. Therefore, after obtaining the multimodal information, the multimodal information can be processed to obtain the text information of each modality. The text information of each modality can be spliced and then the multimodal encoder can be used to process the multimodal information to obtain the feature vector of the video.
[0065] 1012. Use the feature extraction model in the feature processing model to process the feature vector to obtain the video features of the historical watched video.
[0066] In this embodiment, the feature vector can be input into the feature processing model to process and obtain video features. Optionally, the feature processing model can be a convolutional neural network (CNN), Transformer, deep neural network (DNN), etc., which extracts high-dimensional feature vectors that can represent the comprehensive information of the historical watched videos as the video features of the historical watched videos.
[0067] In this embodiment, by obtaining the multimodal information of the historical watched videos and processing the multimodal information to obtain video features, the video features can fully exploit the multimodal information of the videos, fuse features such as vision, text, and audio into a unified feature representation, so as to more comprehensively and accurately depict the video content, provide high-quality feature inputs for subsequent tasks such as video recommendation, classification, or retrieval, and improve the system's ability to understand user interests and recommendation accuracy.
[0068] Based on the above first embodiment, the specific process of obtaining the video features of the historical watched videos in step 1011 may include:
[0069] 10111. Obtain the image information, text information, and audio information of the historical watched video
[0070] In this embodiment, after obtaining the historical watched video, the text information of the historical watched video can be obtained. Optionally, the text information may include text descriptions such as the title, category, label, and introduction of the video. The image information of the video can be obtained. Optionally, the image information may include non-repeating frame images in the video. The audio information of the video can be obtained.
[0071] 10112. Concatenate the text information, the image text recognition information of the image information, and the audio text recognition information of the audio information to obtain text features.
[0072] In this embodiment, based on the frame image, the text information appearing in the frame image can be recognized by means of image text recognition to obtain the image text recognition information. And / or, based on the frame image, the target object appearing in the frame image can also be recognized by sampling the image target recognition method, so as to obtain the semantic information of the image, and then convert the semantic information into the image text recognition information. Optionally, the image text recognition model can be a Convolutional Recurrent Neural Network (CRNN), an Efficient and Accurate Scene Text Detector (EAST), a Transformer-based Text Recognition Model (Transformer-TRM), etc. Optionally, the image target recognition model can be a Convolutional Neural Network (CNN), a Region-based Convolutional Neural Network (R-CNN), a You Only Look Once (YOLO), etc.
[0073] Based on the audio information, the text information corresponding to the audio can be recognized by means of audio text recognition to obtain the audio text recognition information of the audio information. Optionally, the audio text recognition model can be a Deep Speech (DS), a Connectionist Temporal Classification (CTC), a Speech Recognition Transformer (SR-Transformer), etc.
[0074] Furthermore, the text information, the image text recognition information of the image information, and the audio text recognition information of the audio information can be sequentially concatenated to obtain the text feature.
[0075] 10113. Input the text feature into a multi-modal encoder for encoding to obtain a feature vector.
[0076] , in this embodiment, text information can be input into a preset multimodal encoder, which can encode the text features to achieve the encoding of the text features and obtain an encoded feature vector. Optionally, the encoding method can be implemented by means of a Bag of Words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), Word Embedding (WE), etc.
[0077] In this embodiment, by extracting the image, text, and audio information of the historical watched video, using the image text recognition and audio text recognition models to generate text features, and then encoding them into high-dimensional feature vectors through a multimodal encoder, the video content can be comprehensively characterized, achieving the effects of improving the recommendation accuracy, diversity, and coverage rate, meeting the dynamic needs of users, and enhancing the user experience and platform activity.
[0078] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure. As Figure 2 shown, the present disclosure provides a method for video recommendation. The specific process of mapping the video features to a preset codebook set based on a feature processing model in step 102 to obtain the semantic feature sequence of the historical watched video includes:
[0079] 201. Map the video features to a preset codebook set to obtain multiple target codebook vectors that match the video features.
[0080] In this embodiment, the preset codebook set is a set composed of a group of predefined codebook vectors. Each codebook vector represents a specific semantic feature information. Optionally, the semantic feature information represented by each codebook vector is the semantic feature in the abstract concept.
[0081] The process of mapping the video features to the preset codebook set is the process of residual quantization. By mapping the video features to the preset codebook set, multiple target codebook vectors can be obtained according to the matching order. The target codebook vector is the codebook vector in the preset codebook set.
[0082] Optionally, the mapping process can actually be a process of calculating the similarity or distance between the video features and each codebook vector in the codebook set. Based on the similarity or distance, the multiple target codebook vectors can be selected.
[0083] 202. Arrange the multiple target codebook vectors in order according to the matching order of the target codebook vectors to form a semantic feature sequence.
[0084] In this embodiment, according to the matching order of the target codebook vectors, multiple target codebook vectors are arranged in sequence to form the final semantic feature sequence. This semantic feature sequence is an ordered vector sequence. The combination of this semantic feature sequence can represent the multi-level semantic information of the video content. Among them, the position of each target codebook vector can reflect the degree of relevance of the historical viewed video to this target codebook vector. That is, the higher the degree of relevance between the semantic feature corresponding to this target codebook vector and the historical viewed video, the higher the ranking of this target codebook vector. Different arrangement orders reflect the importance or relevance of different semantic features. This serialized representation method not only retains the semantic information of the video features, but also enhances the richness and structure of the semantic expression through the sequential relationship, providing a more accurate input for subsequent semantic matching and recommendation tasks.
[0085] In this embodiment, by mapping the video features to a preset codebook set, selecting multiple target codebook vectors with the highest matching degree, and arranging them in the matching order to form a semantic feature sequence, the capture of the key semantic information of the video and the multi-level semantic representation are realized, the richness and structure of the semantic expression are enhanced, and a precise input is provided for subsequent semantic matching and recommendation tasks.
[0086] Based on the above second embodiment, the specific process of mapping the video features to the preset codebook set in step 201 to obtain multiple target codebook vectors matching the video features can be a process of cyclic matching. In this cyclic matching process, 2011 can be cyclically executed until the exit condition of 2012 is met. Taking one cycle as an example, this process may include:
[0087] 2011. Match the video features with each codebook vector in the preset codebook set to determine the target codebook vector with the highest matching degree with the video features. Calculate the difference between the video features and the target codebook vector, and use the difference as the updated video features.
[0088] In this embodiment, first, the video features can be matched with each codebook vector in the preset codebook set. This matching process is the process of calculating the similarity or distance between the video features and each codebook vector in the preset codebook set. For example, the calculation of the similarity can be obtained through algorithms such as Cosine Similarity (CS), Euclidean Distance (ED), and Dot Product Similarity (DPS). After calculating the similarity between the video features and each codebook vector, the codebook vector with the highest matching degree with the video features can be determined, and this codebook vector can be used as the target codebook vector determined in this cycle. Subsequently, the difference between the video features and the target codebook vector can be calculated, and this difference can be used as the new video features. This difference can represent the remaining information in the video features that is not captured by the target codebook vector. Using this difference as the updated video features for the next round of matching process can capture the semantic features in the remaining information, thereby realizing the layer-by-layer extraction of multi-level semantic information in the video features.
[0089] 2012. Repeat the above process until the number of target codebook vectors obtained by matching reaches the first preset number.
[0090] In this embodiment, after step 2011 is executed, it can be first determined whether the number of target codebooks that have been matched reaches the first number. If it reaches, output the multiple target codebook vectors that have been matched. If not, return to step 2011 and continue the matching.
[0091] In this embodiment, the target codebook vector with the highest matching degree is determined by calculating the similarity between the video features and the preset codebook vector, and the difference is calculated as the updated video features for the next round of matching; the loop of matching and difference calculation extracts multi-level semantic information from the video features to generate a feature sequence that comprehensively represents the video semantics, providing high-quality input for subsequent tasks.
[0092] Figure 3 It is a schematic diagram according to the third embodiment of the present disclosure. As Figure 3 shown, the present disclosure provides a method for video recommendation. Among them, the specific process of determining the video to be recommended that matches the semantic feature sequence of the historical watched video in step 103 can include:
[0093] 301. Determine the recall sequence of the user according to the semantic feature sequence of all the historical watched videos of the user.
[0094] In this embodiment, after obtaining the semantic feature sequences of all the historical videos watched by the user, the semantic feature sequences of all the historical videos watched by the user can be statistically analyzed to determine multiple semantic feature sequences with the highest occurrence frequencies of the user as the recall sequences. Optionally, all the historical videos watched by the user may include the videos watched by the user within a certain period of time recently. Optionally, the certain period of time may be the last year, the last half year, the last 3 months, the last 1 month, etc. Since the semantic feature sequence is composed of a first preset number of codebook vectors. Therefore, in the case of a large number of historical videos watched by the user and the user's preferences being relatively concentrated, there are a large number of repetitions in the semantic feature sequences. Therefore, the multiple semantic feature sequences with the highest occurrence frequencies can be determined by statistically analyzing the occurrence frequencies of the semantic feature sequences of all the historical videos as the recall sequences.
[0095] 302. Determine the videos matching the recall sequences from the video library as the videos to be recommended.
[0096] In this embodiment, the video library contains a vast number of videos. Usually, the number of videos in the video library is much larger than the historical videos watched by the user. Therefore, after determining the recall sequences, the videos to be recommended can be selected by matching the semantic feature sequences of the videos in the video library with the recall sequences. Optionally, the videos with the semantic feature sequences in the video library that are the same as the recall sequences are the videos that meet the user's preferences. Optionally, by statistically analyzing the semantic feature sequences of the historical videos watched by the user over a long period of time, recall sequences that can cover the user's main interest points and potential long-tail interests can be obtained, so as to ensure that the videos to be recommended can better cover the user's main interest points and potential long-tail interests.
[0097] In this embodiment, by statistically analyzing the semantic feature sequences of the user's recent historical videos, the sequences with the highest occurrence frequencies are selected as the recall sequences, and the semantic feature sequences in the video library that are the same as the recall sequences are matched to screen out the videos that meet the user's interests as the videos to be recommended, so as to achieve that the recommended content covers the user's main interest points and potential long-tail interests and improve the diversity and accuracy of the recommendation.
[0098] Based on the above third embodiment, the specific process of step 301 for determining the user's recall sequences according to the semantic feature sequences of all the historical videos watched by the user may include:
[0099] 3011. Statistically analyze the occurrence frequency of each semantic feature sequence in all the historical videos watched by the user.
[0100] In this embodiment, since the semantic feature sequence is a sequence composed of the first preset number of codebook vectors, there are a large number of repetitions in the semantic feature sequence when there are a large number of historical videos watched by the user and the user's preferences are relatively concentrated. Count and record the number of occurrences of each semantic feature sequence in all the historical videos watched by the user. Based on the ratio of the number of occurrences to the total number of videos in all the historical videos watched by the user, the occurrence frequency of each semantic feature sequence can be obtained. Optionally, different occurrence frequencies can reflect different degrees of interest of the user. Optionally, high-frequency semantic feature sequences usually correspond to the main interest points of the user, while low-frequency sequences may reflect the long-tail interests of the user.
[0101] Optionally, a histogram of the semantic feature sequences can also be generated according to the occurrence frequencies of the semantic feature sequences in all the historical videos watched by the user. The histogram indicates the corresponding relationship between each semantic feature sequence and the occurrence frequency. Optionally, the histogram can also be a frequency distribution table. The histogram or the frequency distribution table reflects the preference degree of the user for different semantic features. Optionally, the histogram can be as Figure 4 shown. Where each position on the horizontal axis represents a semantic feature sequence. For example, as Figure 4 shown, there are a total of 5 semantic feature sequences a_1 to a_5 in all the historical videos watched by this user. Each semantic feature sequence corresponds to an entry. Among them, the vertical axis represents the occurrence frequency. The height of the entry of each semantic feature sequence represents the preference degree of the user for this semantic in historical behavior. This statistical method can intuitively display the diversity and distribution characteristics of the user's interests. By generating a histogram or a frequency distribution table of the semantic feature sequences, a visual display of the preference degrees of the user for different semantic features is realized.
[0102] When counting the long-term historical behavior of the user, the occurrence frequency of each semantic feature sequence can be used as an indication of the user's long-term interests. Therefore, based on this histogram, the preference information of the user in multiple semantic dimensions can be captured, so as to reflect the user's preference for popular topics and retain the fine-grained description of long-tail interests. In practical applications, this representation can also be used to calculate the similarity between the user and candidate resources, so as to preferentially select resources that match the user's long-term interest distribution in the recall stage.
[0103] Optionally, in the specific calculation process of the occurrence frequency of the semantic feature sequence, first, the number of occurrences of each semantic feature sequence in all the historical videos watched by the user can be counted. Secondly, the ratio of the number of occurrences of each semantic feature sequence to the number of videos in all the historical videos watched by the user can be calculated to obtain the occurrence frequency of each semantic feature sequence.
[0104] 3012. Select the second preset number of semantic feature sequences with the largest occurrence frequencies as the recall sequences.
[0105] In this embodiment, based on the calculated multiple occurrence frequencies, the second preset number of semantic feature sequences with the largest occurrence frequencies can be selected. These second preset number of semantic feature sequences can be used as the second preset number of recall sequences. Optionally, the second preset number can be determined according to a number preset according to task requirements. The setting of the second preset number is used to control the number and coverage range of the recall sequences. The second preset number of semantic feature sequences with the largest occurrence frequencies can centrally reflect the user's core interest preferences. Optionally, in order to balance the diversity of recommendations, a certain degree of randomness or weight adjustment can be introduced during the selection process to ensure that low-frequency but potentially interesting semantic feature sequences can also be included in the recall sequences, thereby improving the coverage rate and user satisfaction of the recommendation results.
[0106] In this embodiment, by counting the number of occurrences of each semantic feature sequence in all the historical watched videos of the user and calculating the ratio of it to the total number of videos, the analysis of the occurrence frequency of each semantic feature sequence is realized, and by selecting several semantic feature sequences with the largest occurrence frequencies as the recall sequences, the centralized reflection of the user's core interest preferences is realized.
[0107] Based on the above third embodiment, the specific process of step 302 for determining the videos matching the recall sequences from the video library as the videos to be recommended may include:
[0108] 3021. Extract the features of the videos in the video library based on the feature processing model to obtain the video features of the videos in the video library.
[0109] 3022. Map the video features to the preset codebook set of the feature processing model to obtain the semantic feature sequences of the videos in the video library.
[0110] In this embodiment, the implementation of the above steps 3021 and 3022 is the same as the process of obtaining the semantic feature sequences by processing the user's historical watched videos in the above steps 101 and 102. The difference is that here the videos in the video library are processed to obtain the semantic feature sequences of the videos in the video library. It should be noted that in order to improve the processing efficiency, the semantic feature sequences of each video in the video library are usually pre-processed. That is, during the video recommendation process, the semantic feature sequences of each video can be directly extracted from the video library. Instead of calculating the semantic feature sequences of each video in the video library every time a recommendation is executed. Since the number of videos in the video library is extremely large, the video recommendation efficiency can be greatly improved by the way of pre-extracting the semantic feature sequences.
[0111] 3023. Select the third preset number of videos from the videos in the video library corresponding to the semantic feature sequences matching each recall sequence as the videos to be recommended.
[0112] In this embodiment, after obtaining the recall sequence, videos in the video library whose semantic feature sequences are consistent with the recall sequence can be determined according to the recall sequence. Usually, the number of videos in the video library whose semantic feature sequences are consistent with the recall sequence is much larger than the third preset number. Therefore, a third preset number of videos can be randomly selected from the videos whose semantic feature sequences are consistent with the recall sequence as the videos to be recommended corresponding to the recall sequence. Optionally, when there are multiple recall sequences, multiple third numbers of videos to be recommended can be selected from the video library.
[0113] In this embodiment, by matching the semantic feature sequences of the videos in the video library with the recall sequence, the selection of the videos to be recommended is realized, and the diversity of the videos to be recommended is improved.
[0114] Based on the above third embodiment, after obtaining the videos to be recommended, the videos to be recommended can also be processed to generate a video recommendation sequence, and based on the video recommendation sequence, the videos to be recommended are recommended to the user. This process may include:
[0115] 303. Randomly sort the multiple videos to be recommended corresponding to each recall sequence to generate a video recommendation sequence.
[0116] In this embodiment, the third number of videos to be recommended corresponding to each recall sequence are randomly sorted to generate a video recommendation sequence. Optionally, random sorting is to increase the diversity and exploration of the recommendation and avoid the recommendation result being too single. Optionally, the sorting process can also be sorted according to other strategies. Optionally, when sorting according to other strategies, in order to ensure the exploration of the recommendation, random parameters are usually increased.
[0117] 304. Sequentially send the videos to be recommended in the video recommendation sequence to the user according to the video recommendation sequence.
[0118] In this embodiment, after obtaining the video recommendation sequence, the videos to be recommended in the video recommendation sequence can be sequentially sent to the user according to the video recommendation sequence. Optionally, the number of videos viewed by the user is not determined. Therefore, multiple videos can be sent to the user each time according to the user's needs. After the user finishes watching, the subsequent videos of the video recommendation sequence are continued to be sent to the user, thereby improving the cache utilization rate of the user terminal and ensuring the viewing efficiency of the user. Optionally, when the number of remaining videos in the video recommendation sequence is less than the preset threshold, step 101 can be returned to regenerate the videos to be recommended.
[0119] In this embodiment, by randomly sorting the videos to be recommended in each recall sequence, sending videos to the user according to the sequence order, and adjusting the sending quantity according to requirements, a diversified generation of the video recommendation sequence is achieved, improving the cache utilization rate and the viewing efficiency.
[0120] Figure 5 It is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 5 shown, the present disclosure provides a training method for a feature processing model, including:
[0121] 501. Input the first training set into the first feature processing model to obtain the first video information features of the samples in the first training set; input the second training set into the second feature processing model to obtain the second video information features of the samples in the second training set. Among them, the samples in the first training set include historical viewed videos and historical unviewed videos. The first video information features represent the features after decoding the semantic feature sequence of the samples in the first training set. The samples in the second training set include historical viewed videos and historical unviewed videos. The second video information features represent the features after decoding the semantic feature sequence of the samples in the second training set.
[0122] In this embodiment, during the training process of the feature processing model, a siamese network model will be constructed first. The siamese network includes a first feature processing model and a second feature processing model. The first feature processing model and the second feature processing model have the same network structure and network parameters. That is, the first feature processing model and the second feature processing model are two completely identical network models. Optionally, the siamese network can be as Figure 6 shown. Among them, the left side is the first feature processing model, and the right side is the second feature processing model. Optionally, the first feature processing model can be a user tower. The second feature processing model can be a resource tower.
[0123] Secondly, two training sets will be prepared. Among them, the first training set is used to output the first feature processing model. The second training set is used to input the second feature processing model. Both of these sets include historical viewed videos and historical unviewed videos as samples. And, in these two sets, the number of historical viewed videos is the same as the number of historical unviewed videos.
[0124] After the amounts in the first training set are input into the first feature processing model, the model extracts and maps the features of each sample in the first training set to obtain the semantic feature sequence of the sample, and then decodes the semantic feature sequence of the sample to obtain the first video information feature. Similarly, the second training set is input into the second feature processing model to generate the second video information feature. The first video information feature and the second video information feature are used to reflect the description information of the video. Optionally, in an ideal state, the first video information feature and the second video information feature are consistent with the features of the sample after being encoded using a multi-modal encoder.
[0125] Optionally, as Figure 6 shown, the data input into the multi-modal encoders of the first feature processing model and the second feature processing model can be the titles, categories, and pre-targets of each sample, speech recognition information, text recognition information of images, etc.
[0126] The first feature processing model and the second feature processing model take the user's historical watched videos as input and apply the causal attention mechanism to ensure that the currently processed resources only calculate the attention weights with the historical video content, avoiding the interference of future information, and achieving the generation of the first video information feature at each time step and the second video information feature The first video information feature and the second video information feature respectively capture the user's preferences at a specific moment and the features of the video content.
[0127] 502. Based on the first video information feature and the second video information feature, train the first feature processing model and the second feature processing model to obtain the first feature processing model and the second feature processing model that have completed training. Among them, the first feature processing model that has completed training, or the second feature processing model, is used to process the videos in the Figures 1 - 4 embodiment shown above.
[0128] In this embodiment, based on the first video information feature output by the first feature processing model and the second video information feature output by the second feature processing model, the training of the first feature processing model and the second feature processing model can be realized. The process of this training is actually a process of optimizing the model parameters. This training process can enable the model to extract video information features more accurately and efficiently.
[0129] Optionally, the model training process will calculate the loss function using the first video information feature and the second video information feature. Furthermore, based on this loss function, the optimization of the first feature processing model and the second feature processing model can be realized.
[0130] In this embodiment, by constructing a twin network and inputting the training set composed of historical watched and historical unwatched videos into the first feature processing model and the second feature processing model respectively, the extraction of the first video information feature and the second video information feature is realized, and based on the first video information feature and the second video information feature, the training of the first feature processing model and the second feature processing model is realized, improving the accuracy and efficiency of the feature extraction of the first feature processing model and the second feature processing model.
[0131] Based on the above Figure 5 In the fourth embodiment shown, taking the input of the first training set into the first feature processing model to obtain the first video information feature of the samples in the first training set in step 501 as an example, the processing processes of the first feature processing model and the second feature processing model are described.
[0132] Among them, the first feature processing model may specifically include, for example, the model structure as Figure 7 shown. The first feature processing model is a Residual Quantized Variational Autoencoder (RQ-VAE). Among them, Figure 6 the shown transformer part may correspond to Figure 7 the DNN encoding (Encoder), DNN decoding (Decoder), and residual quantization of Figure 6 shown is Figure 7 the input feature vector (embedding) in Figure 6 shown is Figure 7 the first video information feature (embedding) in Figure 7 shown. As shown in Figure 7 , the encoding (Encoder) process is a process of using DNN for convolutional processing with the output of the multimodal encoder as the input. The residual quantization process is a process of gradually quantizing the video features through a preset codebook set, mapping with discrete codebook vectors, and obtaining a sequence of continuous representation codebook vectors. The sequence of continuous representation codebook vectors is the semantic feature sequence. Optionally, the preset codebook set may include multiple discrete codebook vectors. The decoding (Decoder) process is a process of restoring the input according to the semantic feature sequence. This process is used to ensure that semantic information is retained as much as possible.
[0133] Based on the first feature processing model as Figure 7 shown, the specific process may include:
[0134] 5013. Use the first feature processing model to extract features from the samples in the first training set to obtain the video features of the samples.
[0135] In this embodiment, the samples in the first training set can be input into the first feature processing model for feature extraction. Optionally, the samples input into the first feature processing model are feature vectors (embeddings) encoded by a multi-modal encoder. The first feature processing model can process the feature vectors (embeddings) to obtain the video features of the samples.
[0136] 5014. Map the video features to the preset codebook set of the feature processing model to obtain the semantic feature sequence of the historical watched videos. Wherein, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent the semantic features of the videos.
[0137] In this embodiment, an initial preset codebook set can be preset in the first feature processing model. During the model training process, through parameter optimization, the adjustment of each codebook vector in the preset codebook set can be realized, so as to improve the expression ability of the codebook vectors in the preset codebook set for the semantic features of the videos. Map the video features extracted in step 5013 to the codebook space composed of the codebook vectors of the preset codebook set of the feature processing model to generate the semantic feature sequence. The semantic feature sequence can include multiple sequentially arranged codebook vectors.
[0138] 5015. Use the first feature processing model to decode the semantic feature sequence to obtain the first video information feature of the sample.
[0139] In this embodiment, use the network for decoding (Decoder) in the first feature processing model to decode the semantic feature sequence obtained in step 5014 to obtain the first video information feature (embedding) of the sample. Optionally, the decoding process may be implemented by the backpropagation algorithm. Optionally, as much semantic information in the feature vectors as possible should be retained in the first video information feature.
[0140] In this embodiment, by using the first feature processing model to extract features and map the samples in the first training set to obtain the semantic feature sequence, and then processing through the decoding network to obtain the first video information feature rich in semantic information, the purpose of extracting effective features from video data and converting them into numerical feature vectors is realized.
[0141] In the above Figure 5 shown in the fourth embodiment and Figure 6 、 Figure 7 On the basis of the shown embodiment, the process of inputting the second training set into the second feature processing model to obtain the second video information feature of the samples in the second training set may include:
[0142] 5016. Use the second feature processing model to extract features from the samples in the second training set to obtain the video features of the samples.
[0143] 5017. Map the video features to the preset codebook set of the feature processing model to obtain the semantic feature sequence of the historical watched videos. Among them, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent the semantic features of the videos.
[0144] 5018. Use the second feature processing model to decode the semantic feature sequence to obtain the second video information features of the samples.
[0145] The execution of the above steps 5016, 5017, and 5018 is the same as the execution of the above 5013, 5014, and 5015, and will not be elaborated here.
[0146] Based on the above fourth embodiment, in the above step 501, before inputting the first training set and the second training set into the first feature processing model and the second feature processing model respectively in 5013, the first training set and the second training set can also be generated. This process can include:
[0147] 5011. From the historical watched videos of the user within the preset duration, obtain the fourth preset number of historical watched videos that have been completely played and write them into the first training set. And obtain the fourth preset number of historical watched videos that are not repeated with the historical watched videos written into the first training set and write them into the second training set.
[0148] In this embodiment, first, based on the historical watched videos of the user within the preset duration, select the videos that have been completely played. The videos that have been completely played can be considered as the videos that the user is interested in during this viewing process. Otherwise, if the user is not interested, the user may swipe away before the video is completely played and move on to the next video. Optionally, the preset duration can be selected according to the need for data volume. For example, in order to avoid the video types being too single due to the user's short-term interests, the preset duration can be a relatively long time such as 1 year or half a year.
[0149] After that, 2 times the fourth preset number of videos can be screened from the historical watched videos within the preset duration. And write the fourth preset number of these videos into the first training set, and write the other fourth preset number of videos into the second training set. This method can ensure that the videos written into the first training set and the second training set are not repeated. Optionally, this selection method can be random selection. This method can ensure the richness of the data. Optionally, the historical watched videos of this user can be used as the positive samples in the first training set and the second training set.
[0150] 5012. Obtain the videos that the fourth preset number of users have not watched and write them into the first training set. And obtain the videos that the fourth preset number of users have not watched and are non-repetitive with the videos written into the first training set, and write them into the second training set.
[0151] In this embodiment, it is also possible to select videos from the videos that the user has not watched in the video library, and the number of selected videos is twice the fourth preset number. And write the fourth preset number of these videos into the first training set, and write the other fourth preset number of videos into the second training set. This method can ensure that the videos written into the first training set and the second training set are non-repetitive. Optionally, the selection method can be random selection. This method can ensure the richness of the data. Optionally, the videos that the user has not watched can be used as negative samples in the first training set and the second training set.
[0152] In this embodiment, by randomly screening videos from the historical watched videos and unwatched videos that the user has completed playing within the preset duration, and distributing the screened fourth preset number of videos to the first training set and the second training set respectively, and ensuring that the videos in the two sets are non-repetitive, a data set for subsequent model training is constructed in this way.
[0153] Figure 8 is a schematic diagram according to the fifth embodiment of the present disclosure, as Figure 8 shown, the present disclosure provides a training method for a feature processing model. The specific process of training the first feature processing model and the second feature processing model based on the first video information feature and the second video information feature in step 502 may include:
[0154] 801. Calculate the loss functions of the first feature processing model and the second feature processing model based on the first video information feature and the second video information feature.
[0155] In this embodiment, during the model training process, the loss function is a key indicator for measuring the difference between the model prediction result and the actual label. After calculating the first video information feature through the first feature processing model and calculating the second video information feature based on the second feature processing model, based on the model structure of this siamese network model, the loss value of this iteration can be calculated according to the preset loss function. This loss value will be used to guide the subsequent model optimization process.
[0156] 802. Reverse-optimize the first feature processing model and the second feature processing model according to the loss function. Among them, the first feature processing model and the second feature processing model have the same model structure and model parameters.
[0157] In this embodiment, based on this loss function, the backpropagation algorithm can be used to optimize the first feature processing model and the second feature processing model. Since these two models have the same structure and parameters, after optimizing the first feature processing model, the second processing model can be obtained by copying the first feature processing model. Specifically, based on this loss function, the gradient of the loss function with respect to the model parameters can be calculated by the backpropagation algorithm. Furthermore, according to this gradient, the direction of parameter adjustment can be guided. Specifically speaking, the algorithm will first calculate the gradient of the loss function with respect to the output layer, and then pass the gradient backward layer by layer until the input layer is reached. During the passing process, the gradient of each layer will be calculated according to the parameters and activation functions of this layer. Finally, according to the gradient information, the parameters of the model are updated to gradually reduce the loss value and make the prediction result of the model more accurate. This optimization process will continue during the model iteration process until the preset convergence condition or the number of iterations is reached.
[0158] In this embodiment, by calculating the loss function of the video information features and using the backpropagation algorithm to optimize the first feature processing model and the second feature processing model until the convergence condition is met, the improvement of the model prediction accuracy is realized.
[0159] Based on the above fifth embodiment, step 801 realizes the process of calculating the loss function based on the first video information feature and the second video information feature. This process may specifically include:
[0160] 8011. Calculate the first parameters corresponding to each first video information feature and each second video information feature of each user according to the first video information feature and the second video information feature.
[0161] In this embodiment, the first training set and the second training set of multiple users may be included. There is usually a certain correlation between the first video information feature and the second video information feature extracted from the first training set and the second training set of one user. Therefore, after calculating the multiple first video information features and the second video information features of each user, the first parameter between each first video information feature and the second video information feature can be calculated. This first parameter is used to represent an abstract parameter representation between the two features.
[0162] Optionally, the calculation process of this first parameter may include:
[0163] Step 1: Calculate the similarity parameter between the first video information feature and the second video information feature of the user. Step 2: Take the similarity parameter as the exponent of the base of the natural logarithm to calculate the corresponding exponent parameter of the first video information feature and the second video information feature of the user. Step 3: Calculate the sum of the similarity parameters between the first video information feature of other users and each second video information feature of other users to obtain the sum parameter. Step 4: Calculate the ratio of the similarity parameter to the sum of the similarity parameter and the sum parameter to obtain the ratio parameter. Step 5: Calculate the logarithm of the ratio parameter to obtain the first parameter corresponding to each first video information feature and each second video information feature of each user.
[0164] In this embodiment, the Information Noise Contrastive Estimation (InfoNCE) is used to calculate the loss between the first feature extraction model and the second feature extraction model. The InfoNCE loss aims to optimize the feature extraction model by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. Among them, the positive sample pair is composed of the first video information feature of the positive sample in the first training set and the second video information feature of the positive sample in the second training set of the same user. The negative sample pair is composed of the first video information feature of the negative sample in the first training set and the second video information feature of the negative sample in the second training set of the same user.
[0165] Based on this positive sample pair and negative sample pair, the calculation formula of the first parameter can be:
[0166]
[0167] where A is the first parameter. represents the first video information feature corresponding to the jth sample in the first training set of the ith user. represents the second video information feature corresponding to the kth sample in the second training set of the ith user. s(·) represents a function for calculating the similarity between the user and the resource. Optionally, the s(·) can be a cosine function or an inner product. z represents the zth user other than the ith user.
[0168] 8012. Take the negative of the cumulative sum of the first parameters corresponding to each first video information feature and each second video information feature of each user as the loss function.
[0169] In this embodiment, after obtaining the first parameters between each first video information feature and each second video information feature of each user, a loss function can be constructed to quantify the overall performance of these parameters. Specifically, in this embodiment, the negative of the sum of the first parameters corresponding to each first video information feature and each second video information feature of each user is used as the loss function. The formula of this loss function can be:
[0170]
[0171] where L is the loss function. A is the first parameter. N u is the number of users, i represents the i-th user. H represents the number of samples in the first training set of this user, j represents the j-th sample in this first training set. T represents the number of samples in the second training set, and k represents the k-th sample in this second training set.
[0172] When the first parameter is larger, the loss value is smaller, indicating that the correlation or similarity between features is stronger. By minimizing this loss function, this embodiment can enable the model to learn more effective feature representations, thereby improving the performance of video processing and analysis. Optionally, the summation operation in this loss function is performed for all users and their corresponding feature pairs to ensure that the loss function can comprehensively reflect the feature correlation of the entire data set.
[0173] In this embodiment, by calculating the first parameters between user video information features and using the negative of the sum of these parameters as the loss function, it comprehensively considers user feature pairs, improves the model learning efficiency, enhances the accuracy and efficiency, and provides a new optimization means for video processing and analysis.
[0174] Based on the above fifth embodiment, the specific process of step 802 for reversely optimizing the first feature processing model and the second feature processing model according to the loss function may include:
[0175] 8021. According to the loss function, reversely optimize the feature extraction network in the first feature processing model for extracting features from the samples in the first training set.
[0176] In this embodiment, during the model training process, after calculating the loss function, the model parameters can be reversely optimized based on this loss function. The loss function can quantify the gap between the model prediction result and the true label, thus providing us with the direction to optimize the model. Based on the feedback of the loss function, we use the backpropagation algorithm to optimize the feature extraction network. This algorithm calculates the gradient of the loss function with respect to the network parameters and gradually adjusts the parameters in the opposite direction of the gradient, so that the loss value gradually decreases and the feature extraction ability gradually increases. Multiple optimization algorithms, such as stochastic gradient descent, Adam, etc., can be used in this process to update the network parameters and improve the model performance.
[0177] 8022. Reverse-optimize the preset codebook set of the feature processing model according to the loss function.
[0178] In this embodiment, the preset codebook set plays a crucial role in the feature processing model. As a benchmark in the feature space, it is used to map the extracted video features into codebook vectors with semantic information. To improve the expression ability of the codebook vectors for video semantic features, we also optimize the preset codebook set according to the loss function. During the training process, the loss function measures the difference between the semantic feature sequence output by the model and the true label, and then guides the adjustment of the codebook vectors in the codebook set. Through the backpropagation algorithm, we can calculate the gradient of the loss function with respect to the codebook vectors and gradually adjust the vectors in the codebook set accordingly, so that the mapped semantic feature sequence is closer to the true label, thereby improving the overall performance of the model.
[0179] 8023. Reverse-optimize the feature decoding network in the first feature processing model for decoding the semantic feature sequence according to the loss function.
[0180] In this embodiment, the feature decoding network is another key part of the first feature processing model. It is responsible for decoding the semantic feature sequence into video information features with practical application value. To improve the decoding accuracy and efficiency, we also use the loss function as a guide to optimize the feature decoding network. During the training process, the loss function measures the difference between the decoded video information features and the true label, and then guides the adjustment of the decoding network parameters. Through the backpropagation algorithm, we can calculate the gradient of the loss function with respect to the decoding network parameters and gradually adjust the network parameters accordingly, so that the decoded video information features are closer to the true label, thereby improving the performance of the decoding network.
[0181] 8024. Use the optimized first feature processing model to update the second feature processing model.
[0182] In this embodiment, after optimizing the first feature processing model, the parameters of the first feature processing model can be updated to the parameters of the second feature processing model.
[0183] In this embodiment, by calculating the loss function and based on its feedback, the feature extraction network, the preset codebook set, and the feature decoding network in the first feature processing model are optimized using the backpropagation algorithm, thereby updating the model parameters. Finally, the parameters of the optimized first feature processing model are used to update the second feature processing model to improve the overall video processing performance.
[0184] Figure 9 is a schematic diagram according to the sixth embodiment of the present disclosure, as Figure 9 shown, the present disclosure provides a video recommendation device 900, including:
[0185] An extraction unit 901, configured to extract features of the user's historical watched videos based on a feature processing model to obtain video features of the historical watched videos.
[0186] A mapping unit 902, configured to map the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical watched videos. Wherein, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent semantic features of the videos.
[0187] A recommendation unit 903, configured to determine a video to be recommended that matches the semantic feature sequence of the historical watched videos from a video library, and recommend the video to be recommended to the user.
[0188] The device of this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0189] Figure 10 is a schematic diagram according to the seventh embodiment of the present disclosure, as Figure 10 shown, the present disclosure provides a video recommendation device 1000, including:
[0190] An extraction unit 1001, configured to extract features of the user's historical watched videos based on a feature processing model to obtain video features of the historical watched videos.
[0191] A mapping unit 1002, configured to map the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical watched videos. Wherein, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent semantic features of the videos.
[0192] A recommendation unit 1003, configured to determine a video to be recommended that matches the semantic feature sequence of the historical watched videos from a video library, and recommend the video to be recommended to the user.
[0193] Optionally, the mapping unit 1002 includes:
[0194] A mapping module 10021, configured to map video features to a preset codebook set to obtain a plurality of target codebook vectors that match the video features.
[0195] A combination module 10022, configured to sequentially arrange a plurality of target codebook vectors according to the matching order of the target codebook vectors to form a semantic feature sequence.
[0196] Optionally, the mapping module 10021 includes:
[0197] A first sub-module 100211, configured to match video features with each codebook vector in the preset codebook set to determine a target codebook vector with the highest matching degree with the video features. Calculate the difference between the video features and the target codebook vector, and use the difference as the updated video features.
[0198] A second sub-module 100212, configured to loop the above process until the number of target codebook vectors obtained by matching reaches a first preset number.
[0199] Optionally, the recommendation unit 1003 includes:
[0200] A determination module 10031, configured to determine a recall sequence of a user according to the semantic feature sequences of all historical watched videos of the user.
[0201] A matching module 10032, configured to determine, from a video library, videos that match the recall sequence as videos to be recommended.
[0202] Optionally, the determination module 10031 includes:
[0203] A statistics sub-module, configured to count the occurrence frequency of each semantic feature sequence in all historical watched videos of the user.
[0204] A selection sub-module 100312, configured to select a second preset number of semantic feature sequences with the largest occurrence frequencies as the recall sequence.
[0205] Optionally, the statistics sub-module 100311 is configured to:
[0206] Count the occurrence times of each semantic feature sequence in all historical watched videos of the user.
[0207] Calculate the ratio of the occurrence times of each semantic feature sequence to the number of videos in all historical watched videos of the user to obtain the occurrence frequency of each semantic feature sequence.
[0208] Optionally, the determination module 10031 further includes:
[0209] A display sub-module 100313 is configured to generate a histogram of semantic feature sequences based on the occurrence frequencies of each semantic feature sequence in all historical watched videos of a user. The histogram indicates the corresponding relationship between each semantic feature sequence and the occurrence frequency.
[0210] Optionally, the matching module 10032 includes:
[0211] An extraction sub-module 100321 is configured to perform feature extraction on videos in a video library based on a feature processing model to obtain video features of the videos in the video library.
[0212] A mapping sub-module 100322 is configured to map the video features into a preset codebook set of the feature processing model to obtain semantic feature sequences of the videos in the video library.
[0213] A matching sub-module 100323 is configured to select a third preset number of videos from the videos in the video library corresponding to the semantic feature sequences matching each recall sequence as the videos to be recommended.
[0214] Optionally, the recommendation unit 1003 includes:
[0215] A generation module 10033 is configured to randomly sort the multiple videos to be recommended corresponding to each recall sequence to generate a video recommendation sequence.
[0216] A recommendation module 10034 is configured to sequentially send the videos to be recommended in the video recommendation sequence to a user according to the video recommendation sequence.
[0217] Optionally, the processing unit 1001 includes:
[0218] A first processing module 10011 is configured to obtain multi-modal information of historical watched videos, and process the multi-modal information using a multi-modal encoder to obtain a feature vector of the historical watched videos.
[0219] A second processing module 10012 is configured to process the feature vector using a feature extraction model in the feature processing model to obtain video features of the historical watched videos.
[0220] Optionally, the first processing module 10011 includes:
[0221] An acquisition sub-module 100111 is configured to acquire image information, text information, and audio information of the historical watched videos;
[0222] A splicing sub-module 100112 is configured to splice the text information, the image text recognition information of the image information, and the audio text recognition information of the audio information to obtain a text feature;
[0223] The encoding sub-module 100113 is configured to input the text features into the multi-modal encoder for encoding to obtain the feature vectors.
[0224] The device in this embodiment can execute the technical solutions in the above method. The specific implementation process and technical principle are the same and will not be elaborated here.
[0225] Figure 11 It is a schematic diagram according to the eighth embodiment of the present disclosure. As Figure 11 shown, the present disclosure provides a training device 1100 for a feature processing model, including:
[0226] An input unit 1101 is configured to input a first training set into a first feature processing model to obtain first video information features of samples in the first training set. The second training set is input into a second feature processing model to obtain second video information features of samples in the second training set. Among them, the samples in the first training set include historical watched videos and historical unwatched videos. The first video information features represent the features after decoding the semantic feature sequences of the samples in the first training set. The samples in the second training set include historical watched videos and historical unwatched videos. The second video information features represent the features after decoding the semantic feature sequences of the samples in the second training set.
[0227] A training unit 1102 is configured to train the first feature processing model and the second feature processing model based on the first video information features and the second video information features to obtain the first feature processing model and the second feature processing model that have completed training. Among them, the first feature processing model or the second feature processing model after completing training is used to process Figures 1 - 3 the videos in the method shown.
[0228] The device in this embodiment can execute the technical solutions in the above method. The specific implementation process and technical principle are the same and will not be elaborated here.
[0229] Figure 12 It is a schematic diagram according to the ninth embodiment of the present disclosure. As Figure 12 shown, the present disclosure provides a training device 1200 for a feature processing model, including:
[0230] An input unit 1201 is configured to input a first training set into a first feature processing model to obtain first video information features of samples in the first training set. The second training set is input into a second feature processing model to obtain second video information features of samples in the second training set. Wherein, the samples in the first training set include historical watched videos and historical unwatched videos. The first video information features represent the features after decoding the semantic feature sequence of the samples in the first training set. The samples in the second training set include historical watched videos and historical unwatched videos. The second video information features represent the features after decoding the semantic feature sequence of the samples in the second training set.
[0231] A training unit 1202 is configured to train the first feature processing model and the second feature processing model based on the first video information features and the second video information features to obtain the trained first feature processing model and the second feature processing model.
[0232] Wherein, the trained first feature processing model, or the second feature processing model, is used to process the Figures 1 - 3 videos in the method shown.
[0233] Optionally, the input unit 1201 includes:
[0234] A first extraction module 12011 is configured to use the first feature processing model to extract features of samples in the first training set to obtain video features of the samples.
[0235] A first mapping module 12012 is configured to map the video features to a preset codebook set of the feature processing model to obtain a semantic feature sequence of historical watched videos. Wherein, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent the semantic features of the videos.
[0236] A first decoding module 12013 is configured to use the first feature processing model to decode the semantic feature sequence to obtain the first video information features of the samples.
[0237] Optionally, the input unit 1201 includes:
[0238] A second extraction module 12014 is configured to use the second feature processing model to extract features of samples in the second training set to obtain video features of the samples.
[0239] A second mapping module 12015 is configured to map the video features to a preset codebook set of the feature processing model to obtain a semantic feature sequence of historical watched videos. Wherein, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent the semantic features of the videos.
[0240] The second decoding module 12016 is configured to decode the semantic feature sequence using the second feature processing model to obtain the second video information feature of the sample.
[0241] Optionally, the training unit 1202 includes:
[0242] The loss module 12021 is configured to calculate the loss functions of the first feature processing model and the second feature processing model based on the first video information feature and the second video information feature.
[0243] The optimization module 12022 is configured to reversely optimize the first feature processing model and the second feature processing model according to the loss functions.
[0244] Wherein, the first feature processing model and the second feature processing model have the same model structure and model parameters.
[0245] Optionally, the loss module 12021 includes:
[0246] The first calculation sub-module 120211 is configured to calculate the first parameters corresponding to the first video information features and the second video information features of each user according to the first video information feature and the second video information feature.
[0247] The second calculation sub-module 120212 is configured to use the negative value of the sum of the first parameters corresponding to the first video information features and the second video information features of each user as the loss function.
[0248] Optionally, the first calculation sub-module 120211 is configured to:
[0249] Calculate the similarity parameter between the first video information feature and the second video information feature of the user.
[0250] Use the similarity parameter as the exponent of the base of the natural logarithm to calculate the exponential parameter corresponding to the first video information feature and the second video information feature of the user.
[0251] Calculate the sum of the similarity parameters between the first video information features of other users and the second video information features of other users to obtain the summation parameter.
[0252] Calculate the ratio of the similarity parameter to the sum of the similarity parameter and the summation parameter to obtain the ratio parameter.
[0253] Calculate the logarithm of the ratio parameter to obtain the first parameters corresponding to the first video information features and the second video information features of each user.
[0254] Optionally, the optimization module 12022 includes:
[0255] The first optimization sub-module 120221 is used to reversely optimize the feature extraction network in the first feature processing model for extracting features from the samples in the first training set according to the loss function.
[0256] The second optimization sub-module 120222 is used to reversely optimize the preset codebook set of the feature processing model according to the loss function.
[0257] The third optimization sub-module 120223 is used to reversely optimize the feature decoding network in the first feature processing model for decoding the semantic feature sequence according to the loss function.
[0258] The synchronization sub-module 120224 is used to update the second feature processing model using the optimized first feature processing model.
[0259] Optionally, the input unit 1201 further includes:
[0260] The training data generation module 12010 is used to obtain the fourth preset number of historical watched videos that have been completed for playback from the user's historical watched videos within a preset duration and write them into the first training set. And obtain the fourth preset number of historical watched videos that are not repeated with the historical watched videos written into the first training set and write them into the second training set. Obtain the fourth preset number of videos that the user has not watched and write them into the first training set. And obtain the fourth preset number of videos that the user has not watched and are not repeated with the videos that the user has not watched written into the first training set and write them into the second training set.
[0261] The device in this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0262] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0263] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, and at least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program to enable the electronic device to execute the solution provided in any one of the above embodiments.
[0264] Figure 13FIG. shows a schematic block diagram of an exemplary electronic device 1300 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0265] As Figure 13 shown, the device 1300 includes a computing unit 1301 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0266] Multiple components in the device 1300 are connected to the I / O interface 1305, including: an input unit 1306, such as a keyboard, a mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, an optical disk, etc.; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1309 allows the device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0267] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 executes the various methods and processes described above, such as the video recommendation method or the training method of the feature processing model. For example, in some embodiments, the video recommendation method or the training method of the feature processing model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the video recommendation method or the training method of the feature processing model described above can be executed. Alternatively, in other embodiments, the computing unit 1301 can be configured to execute the video recommendation method or the training method of the feature processing model in any other suitable way (e.g., by means of firmware).
[0268] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0269] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0270] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0271] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0272] The systems and techniques described herein can be implemented in a computing system including a back-end component (e.g., as a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0273] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0274] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0275] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for video recommendation, comprising: Performing feature extraction on the user's historical viewed videos based on a feature processing model to obtain video features of the historical viewed videos; Mapping the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical viewed videos; wherein, the preset codebook set includes a plurality of codebook vectors, and the codebook vectors represent semantic features of videos; Determining videos to be recommended that match the semantic feature sequence of the historical viewed videos from a video library; and recommending the videos to be recommended to the user.
2. The method according to claim 1, wherein The mapping the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical viewed videos includes: Mapping the video features into a preset codebook set to obtain a plurality of target codebook vectors that match the video features; Arranging the plurality of target codebook vectors in sequence according to the matching order of the target codebook vectors to form the semantic feature sequence.
3. The method according to claim 2, wherein, The mapping the video features into a preset codebook set to obtain a plurality of target codebook vectors that match the video features includes: Matching the video features with each of the codebook vectors in the preset codebook set to determine the target codebook vector with the highest matching degree with the video features; calculating the difference between the video features and the target codebook vector, and using the difference as the updated video features; Repeating the above process until the number of the obtained target codebook vectors reaches a first preset number.
4. The method according to any one of claims 1-3, wherein The determining videos to be recommended that match the semantic feature sequence of the historical viewed videos from a video library includes: Determining a recall sequence of the user according to the semantic feature sequences of all the historical viewed videos of the user; Determining videos in the video library that match the recall sequence as the videos to be recommended.
5. The method according to claim 4, wherein, The determining a recall sequence of the user according to the semantic feature sequences of all the historical viewed videos of the user includes: Counting the occurrence frequencies of each semantic feature sequence in all the historical viewed videos of the user; Selecting a second preset number of the semantic feature sequences with the largest occurrence frequencies as the recall sequence.
6. The method according to claim 4 or 5, wherein, The determining videos in the video library that match the recall sequence as the videos to be recommended includes: Performing feature extraction on the videos in the video library based on a feature processing model to obtain video features of the videos in the video library; Mapping the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the videos in the video library; Selecting a third preset number of videos from the videos in the video library corresponding to each semantic feature sequence that matches the recall sequence as the videos to be recommended.
7. The method according to any one of claims 1-6, wherein, The recommending the videos to be recommended to the user includes: Randomly sorting the multiple videos to be recommended corresponding to each recall sequence to generate a video recommendation sequence; Sequentially sending the videos to be recommended in the video recommendation sequence to the user according to the video recommendation sequence.
8. The method according to any one of claims 1-7, wherein, Performing feature extraction on the historical viewing videos of the user based on the feature processing model to obtain video features of the historical viewing videos, including: Obtaining multi-modal information of the historical viewing videos; and processing the multi-modal information using a multi-modal encoder to obtain a feature vector of the historical viewing videos; Processing the feature vector using a feature extraction model in the feature processing model to obtain video features of the historical viewing videos.
9. A method for training a feature processing model, including: Inputting a first training set into a first feature processing model to obtain first video information features of samples in the first training set; Inputting a second training set into a second feature processing model to obtain second video information features of samples in the second training set; wherein, the samples in the first training set include historical viewing videos and historical non-viewing videos; the first video information features represent features after decoding the semantic feature sequence of samples in the first training set; the samples in the second training set include historical viewing videos and historical non-viewing videos; the second video information features represent features after decoding the semantic feature sequence of samples in the second training set; Training the first feature processing model and the second feature processing model based on the first video information features and the second video information features to obtain a first feature processing model and a second feature processing model that are completed in training; Wherein, the first feature processing model that is completed in training, or, the second feature processing model, is used to process videos in the method according to any one of claims 1-8.
10. The method according to claim 9, wherein, The inputting the first training set into the first feature processing model to obtain first video information features of samples in the first training set includes: Performing feature extraction on samples in the first training set using the first feature processing model to obtain video features of the samples; Mapping the video features to a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historical viewing videos; wherein, the preset codebook set includes multiple codebook vectors, and the codebook vectors represent semantic features of videos; Decoding the semantic feature sequence using the first feature processing model to obtain first video information features of the samples.
11. The method according to claim 9 or 10, wherein The training the first feature processing model and the second feature processing model based on the first video information features and the second video information features includes: Calculating loss functions of the first feature processing model and the second feature processing model based on the first video information features and the second video information features; Backward optimizing the first feature processing model and the second feature processing model according to the loss functions; Wherein, the first feature processing model and the second feature processing model have the same model structure and model parameters.
12. The method according to claim 11, wherein, The calculating loss functions of the first feature processing model and the second feature processing model based on the first video information features and the second video information features includes: Calculate a first parameter corresponding to each of the first video information features and each of the second video information features of each user according to the first video information features and the second video information features; Use the negative of the sum of the first parameters corresponding to each of the first video information features and each of the second video information features of each user as the loss function.
13. The method according to claim 12, wherein, The calculating, according to the first video information features and the second video information features, a first parameter corresponding to each of the first video information features and each of the second video information features of each user includes: Calculate a similarity parameter between the first video information feature and the second video information feature of the user; Use the similarity parameter as the exponent of the base of the natural logarithm to calculate an exponent parameter corresponding to the first video information feature and the second video information feature of the user; Calculate the sum of the similarity parameters between the first video information features of other users and each of the second video information features of the other users to obtain a sum parameter; Calculate the ratio of the similarity parameter to the sum of the similarity parameter and the sum parameter to obtain a ratio parameter; Calculate the logarithm of the ratio parameter to obtain a first parameter corresponding to each of the first video information features and each of the second video information features of each user.
14. The method according to any one of claims 11 - 13, wherein, The reversely optimizing the first feature processing model and the second feature processing model according to the loss function includes: According to the loss function, reversely optimize the feature extraction network in the first feature processing model for extracting features from samples in the first training set; According to the loss function, reversely optimize the preset codebook set of the feature processing model; According to the loss function, reversely optimize the feature decoding network in the first feature processing model for decoding the semantic feature sequence; Use the optimized first feature processing model to update the second feature processing model.
15. The method according to any one of claims 9-14, wherein, It further includes: Obtain a fourth preset number of historically viewed videos that have been played to completion from the historically viewed videos of the user within a preset duration and write them into the first training set; And obtain a fourth preset number of historically viewed videos that do not repeat the historically viewed videos written into the first training set and write them into the second training set; Obtain a fourth preset number of videos that the user has not viewed and write them into the first training set; And obtain a fourth preset number of videos that the user has not viewed and do not repeat the videos written into the first training set and write them into the second training set.
16. A video recommendation device, comprising: An extraction unit, configured to extract features of a user's historically viewed videos based on a feature processing model to obtain video features of the historically viewed videos; A mapping unit, configured to map the video features into a preset codebook set of the feature processing model to obtain a semantic feature sequence of the historically viewed videos; wherein, the preset codebook set includes a plurality of codebook vectors, and the codebook vectors represent semantic features of videos; A recommendation unit, configured to determine a video to be recommended that matches the semantic feature sequence of the historical watched video from a video library; and recommend the video to be recommended to the user.
17. A training device for a feature processing model, comprising: An input unit, configured to input a first training set into a first feature processing model to obtain first video information features of samples in the first training set; Input a second training set into a second feature processing model to obtain second video information features of samples in the second training set; wherein, the samples in the first training set include historical watched videos and historical unwatched videos; the first video information features represent the features after decoding the semantic feature sequences of the samples in the first training set; the samples in the second training set include historical watched videos and historical unwatched videos; the second video information features represent the features after decoding the semantic feature sequences of the samples in the second training set; A training unit, configured to train the first feature processing model and the second feature processing model based on the first video information features and the second video information features to obtain a first feature processing model and a second feature processing model that are completed in training; wherein, the first feature processing model after the completion of training, or, the second feature processing model, is configured to process the videos in the method according to any one of claims 1-8.
18. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8; or, the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 9-15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8; or, the computer instructions are used to cause the computer to execute the method according to any one of claims 9-15.
20. A computer program product, comprising a computer program, where when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8; or, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 9-15.