A cross-modal short video recommendation method, system, terminal and storage medium

By acquiring the multimodal preferences of target users and constructing a video recommendation knowledge graph, the problem of insufficient cross-modal recommendation capabilities is solved, and high-accuracy short video recommendation is achieved.

CN117290544BActive Publication Date: 2026-04-28NOTTINGHAM (NINGBO FREE TRADE ZONE) BLOCKCHAIN CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NOTTINGHAM (NINGBO FREE TRADE ZONE) BLOCKCHAIN CO LTD
Filing Date
2023-10-13
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing short video recommendation methods lack the ability to recommend across different modalities, resulting in low recommendation accuracy and an inability to effectively utilize the rich domain knowledge and user history data of traditional print media and cultural promotion departments.

Method used

By acquiring the multimodal preferences of target users, a video recommendation knowledge graph is constructed. Combining user preferences and historical video analysis results, relevant resources are selected from a pre-set resource library to achieve cross-modal video recommendation.

Benefits of technology

It improves the accuracy and precision of short video recommendations, and can better utilize user multimodal preference data for video recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117290544B_ABST
    Figure CN117290544B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of short videos, in particular to a cross-modal short video recommendation method and system, a terminal and a storage medium, which comprises the following steps: acquiring the user preference of a target user; performing multi-modal analysis on historical videos based on the user preference and generating an analysis result; constructing a video recommendation knowledge graph based on a preset resource library and a preset construction method; obtaining a video to be recommended based on the analysis result and the video recommendation knowledge graph; and recommending the target video to be recommended to the target user. The application helps to realize video recommendation of high-accuracy content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of short video technology, and in particular to a cross-modal short video recommendation method, system, terminal and storage medium. Background Technology

[0002] Short video recommendation methods primarily rely on video content and related text, images, and audio information for recommendations. Text-based recommendations extract textual information such as video titles, descriptions, and tags, and use techniques like bag-of-words models, TF-IDF, and word vectors (such as Word2Vec and GloVe) to convert the text into feature vectors. Then, they calculate the similarity of the same modal data with candidate short videos for recommendations. Image-based recommendations extract keyframes from the video, use pre-trained deep learning models (such as VGG, ResNet, and Inception) to extract image feature vectors, and calculate the similarity of the same modal data with candidate short videos for recommendations. Audio-based recommendations extract audio signals from the video, use audio feature extraction techniques to generate feature vectors, and calculate the similarity of the same modal data with candidate short videos for recommendations.

[0003] With the increasing penetration of short videos, traditional print media and cultural promotion departments (such as cultural heritage sites and research institutions) that primarily rely on text and images are also transforming into "converged media" centered on short videos. The inadequacy of recommendation methods across different modalities has become a bottleneck in this transformation process. At the same time, traditional print media and cultural promotion departments possess a wealth of systematic domain knowledge and user history data in text-based formats. How to better utilize this information to improve the accuracy of short video recommendations to users is crucial—that is, how to conduct cross-modal video recommendations. Current recommendation methods, however, suffer from insufficient support and low accuracy in this area. Summary of the Invention

[0004] To facilitate high-accuracy video recommendation, this application provides a cross-modal short video recommendation method, system, terminal, and storage medium.

[0005] Firstly, this application provides a cross-modal short video recommendation method, which adopts the following technical solution:

[0006] A cross-modal short video recommendation method includes:

[0007] Obtain the user preferences of the target users;

[0008] Based on the user preferences, multimodal analysis is performed on historical videos and analysis results are generated;

[0009] Based on a pre-defined resource library and a pre-defined construction method, a video recommendation knowledge graph is constructed.

[0010] Based on the analysis results and the video recommendation knowledge graph, obtain the videos to be recommended;

[0011] The target video to be recommended is recommended to the target user.

[0012] By adopting the above technical solution, the user preferences of the target user are first obtained, and multimodal analysis of the historical videos watched is performed based on these preferences to generate analysis results. Relevant resources are then selected from a preset resource library, and a short video recommendation knowledge graph is constructed according to a preset framework method. Finally, by combining the generated analysis results and the constructed video recommendation knowledge graph, videos to be recommended are obtained and recommended to the target user. Performing multimodal analysis on historical videos based on user preferences and combining data from different modalities based on the video recommendation knowledge graph helps improve the accuracy of short video recommendations to users, thereby helping to achieve high-accuracy video recommendations.

[0013] Optionally, the specific steps for obtaining the target user's user preferences include:

[0014] Obtain the target user's browsing history;

[0015] Based on the historical browsing records, obtain text reading records, image viewing records, and audio listening records;

[0016] Based on the text reading records, the user preferences based on text modality are obtained as text preferences;

[0017] Based on the image viewing history, the user preferences based on image modality are obtained as image preferences;

[0018] Based on the audio listening record, the user preference based on the audio modality is obtained as the audio preference;

[0019] User preferences are obtained based on the text preferences, image preferences, and audio preferences.

[0020] By adopting the above technical solution, text reading records, image viewing records, and audio listening records are obtained based on the target user's historical browsing history. Then, user preferences are obtained based on the text reading records, image viewing records, and audio listening records respectively. Finally, the user preferences under all different modalities are combined to obtain the final user preferences. Combining user preferences under multiple modalities to obtain the final user preferences results in more complete data. Using this as the basis for video recommendation helps to improve the accuracy of video content.

[0021] Optionally, the specific steps for obtaining the user preferences based on the text reading record as text preferences include:

[0022] Based on the text record, obtain the reading text;

[0023] Based on preset word segmentation rules, the reading text is segmented into words, and segmented word groups are obtained;

[0024] Obtain the word frequency and weight of the segmented word groups;

[0025] Based on the word frequency and the weight, keywords are obtained, and text keyword vectors are generated based on the keywords;

[0026] Based on the text records, obtain the target quantity and reading time for different reading content;

[0027] Based on the reading content, obtain the number of keywords for different keywords;

[0028] Based on the target number, the reading time, the number of keywords, and a preset calculation function, a weighted reading time is obtained; based on the weighted reading time, text preferences are obtained.

[0029] By employing the above technical solution, the reading text, the target number of different reading contents, and the reading duration are obtained based on the text records. The reading text is then segmented into words, and word frequencies and weights are obtained. Based on these word frequencies and weights, a text keyword vector is generated. According to the reading content, the number of different keywords is obtained. Finally, combining the target number, reading duration, and number of keywords, a weighted reading time is calculated using a preset calculation function. Text preferences are then obtained based on this weighted reading time. Based on the weighted reading time, it is helpful to identify the reading content that the target user is more interested in. Video recommendations are then made based on the reading content that the user is more interested in, thereby helping to improve the accuracy of video content.

[0030] Optionally, the specific steps for obtaining the user preference based on the image modality as image preference based on the image viewing history include:

[0031] Based on the image viewing history, obtain the target image;

[0032] Obtain the image style of the target image;

[0033] Based on the target image and the preset recognition method, identify the target in the image;

[0034] Based on the image style and the image target, generate an image keyword vector;

[0035] Based on the image viewing history, obtain the total number of images and the viewing duration of different target images;

[0036] Based on the total number of images, the viewing duration, the image keyword vector, and the preset calculation function, the viewing time of key images is obtained;

[0037] Image preferences are obtained based on the viewing time of the key images.

[0038] By adopting the above technical solution, based on image viewing records, the system obtains the target image, the total number of images, and the viewing duration of different target images. It also acquires the image style of the target images, identifies the image targets within them, and generates image keyword vectors based on the image style and image targets. Finally, combining the total number of images, viewing duration, and image keyword vectors, a preset calculation function is used to calculate the key image viewing time, and image preferences are obtained based on this key image viewing time. This key image viewing time helps identify images that target users are more interested in, allowing for video recommendations based on these images, thereby improving the accuracy of video content.

[0039] Optionally, the specific steps for obtaining the user preference based on the audio modality as audio preference based on the audio listening record include:

[0040] Based on the audio listening record, the target audio is obtained;

[0041] Feature extraction is performed on the target audio, and the extracted features are generated;

[0042] Based on the extracted features, an audio keyword vector is generated;

[0043] Based on the audio listening records, the listening time of different audio and the total listening time of all audio are obtained;

[0044] Based on the audio keyword vector, the listening time, the total time, and the preset calculation function, the listening time of key audio is obtained;

[0045] Based on the key audio listening time, audio preferences are obtained.

[0046] By adopting the above technical solution, based on audio listening records, the listening time of the target audio, different audios, and the total listening time of all audio are obtained. Features are extracted from the target audio, and an audio keyword vector is generated. Finally, by combining the audio keyword vector, listening time, and total time, the key audio listening time is calculated using a preset calculation function, and audio preferences are obtained based on the key audio listening time. Based on the key audio listening time, it is helpful to identify the audio that the target user is more interested in, and video recommendations are made based on the audio that the user is more interested in, thereby helping to improve the accuracy of video content.

[0047] Optionally, the specific steps for performing multimodal analysis on historical videos and generating analysis results based on the user preferences include:

[0048] Based on the historical video, characteristic targets are obtained;

[0049] The feature targets are analyzed, and feature labels corresponding to the feature targets are generated;

[0050] Based on the feature labels, generate a label keyword vector;

[0051] Based on the historical videos, the viewing duration of different videos is obtained;

[0052] Based on all the aforementioned viewing durations, obtain the total video viewing duration;

[0053] Based on the tag keyword vector, the viewing duration, the total video viewing duration, and the preset calculation function, the viewing time of key tags is obtained;

[0054] Based on the obtained text keyword vectors, image keyword vectors, audio keyword vectors, tag keyword vectors, and the preset calculation function, a target vector is obtained, and the target vector is used as the analysis result.

[0055] By adopting the above technical solution, feature targets are obtained from historical videos, and feature tag vectors corresponding to the feature targets are generated. Then, the viewing duration and total viewing duration of different videos are obtained. Finally, the viewing duration of key tags is calculated by combining the tag keyword vector, viewing duration, and total viewing duration through a preset calculation function. Based on the viewing duration of key tags, text keyword vectors, image keyword vectors, audio keyword vectors, tag keyword vectors, and the preset calculation function, a target vector is obtained, and the target vector is used as the analysis result. Based on the text keyword vector, image keyword vector, audio keyword vector, and tag keyword vector, a high-dimensional vector composed of relevant data tags is finally obtained. Video recommendation is performed based on the high-dimensional vector, which helps to improve the accuracy of video content.

[0056] Optionally, the specific steps for obtaining the video to be recommended based on the analysis results and the video recommendation knowledge graph include:

[0057] Based on the aforementioned tag keyword vector, obtain the keyword weight vector;

[0058] Based on the target vector, the tag keyword vector, the keyword weight vector, and the preset calculation function, the target mapping relationship is obtained;

[0059] Based on the target mapping relationship and the video recommendation knowledge graph, the video to be recommended is obtained.

[0060] By adopting the above technical solution, a target mapping relationship can be obtained based on the target vector, tag keyword vector, keyword weight vector, and preset calculation function. This allows user preferences in one modality to be mapped to user preferences in other different modalities. Then, based on the target mapping relationship and the video recommendation knowledge graph, videos to be recommended can be obtained, thereby achieving cross-modal video recommendation. Video recommendation based on the mapping relationship and the video recommendation knowledge graph helps improve the accuracy of video content recommendation.

[0061] Secondly, this application also discloses a cross-modal short video recommendation system, which adopts the following technical solution:

[0062] A cross-modal short video recommendation system, comprising:

[0063] The first acquisition module is used to acquire the user preferences of the target user;

[0064] The generation module is used to perform multimodal analysis on historical videos based on the user preferences and generate analysis results;

[0065] The construction module is used to build a video recommendation knowledge graph based on a preset resource library and preset construction methods;

[0066] The second acquisition module is used to acquire videos to be recommended based on the analysis results and the video recommendation knowledge graph;

[0067] The recommendation module is used to recommend the target video to the target user.

[0068] By adopting the above technical solution, the user preferences of the target user are first obtained, and multimodal analysis of the historical videos watched is performed based on these preferences to generate analysis results. Relevant resources are then selected from a preset resource library, and a short video recommendation knowledge graph is constructed according to a preset framework method. Finally, by combining the generated analysis results and the constructed video recommendation knowledge graph, videos to be recommended are obtained and recommended to the target user. Performing multimodal analysis on historical videos based on user preferences and combining data from different modalities based on the video recommendation knowledge graph helps improve the accuracy of short video recommendations to users, thereby helping to achieve high-accuracy video recommendations.

[0069] Thirdly, the computer device provided in this application adopts the following technical solution:

[0070] A smart terminal includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the processor loads the computer program, it executes the method of the first aspect.

[0071] By adopting the above technical solution, a computer program is generated based on the method of the first aspect and stored in a memory for loading and execution by a processor. Thus, a smart terminal is made based on the memory and the processor, making it convenient for users to use.

[0072] Fourthly, the computer-readable storage medium provided in this application adopts the following technical solution:

[0073] A computer-readable storage medium storing a computer program that, when loaded by a processor, executes the method of the first aspect.

[0074] By adopting the above technical solution, a computer program is generated based on the method of the first aspect and stored in a computer-readable storage medium for loading and execution by a processor. The computer-readable storage medium facilitates the reading and storage of the computer program.

[0075] In summary, this application includes the following beneficial technical effects:

[0076] First, the user preferences of the target user are obtained, and then multimodal analysis is performed on the historical videos watched based on these preferences to generate analysis results. Relevant resources are then selected from a pre-set resource library, and a short video recommendation knowledge graph is constructed according to a pre-set framework. Finally, by combining the generated analysis results and the constructed video recommendation knowledge graph, videos to be recommended are obtained and recommended to the target user. Performing multimodal analysis on historical videos based on user preferences and combining data from different modalities based on the video recommendation knowledge graph helps improve the accuracy of short video recommendations to users, thereby contributing to the achievement of high-accuracy video recommendation. Attached Figure Description

[0077] Figure 1 This is a flowchart of the main process of a cross-modal short video recommendation method according to an embodiment of this application;

[0078] Figure 2 This is a flowchart of steps S201 to S206;

[0079] Figure 3 This is a flowchart of steps S301 to S308;

[0080] Figure 4 This is a flowchart of steps S401 to S407;

[0081] Figure 5 This is a flowchart of steps S501 to S506;

[0082] Figure 6 This is a flowchart of steps S601 to S607;

[0083] Figure 7 This is a flowchart of steps S701 to S703;

[0084] Figure 8 This is a block diagram of a cross-modal short video recommendation system according to an embodiment of this application.

[0085] Explanation of reference numerals in the attached figures:

[0086] 1. First acquisition module; 2. Generation module; 3. Construction module; 4. Second acquisition module; 5. Recommendation module. Detailed Implementation

[0087] Firstly, this application discloses a cross-modal short video recommendation method.

[0088] Reference Figure 1 A cross-modal short video recommendation method, comprising steps S101 to S108:

[0089] Step S101: Obtain the user preferences of the target user.

[0090] Specifically, the target user is the user who needs video recommendations, and the user preference is the target user's interest and liking when watching videos. In this embodiment, the user preference can be inferred from the text read by the target user, the pictures viewed, and the audio listened to.

[0091] Step S102: Based on user preferences, perform multimodal analysis on historical videos and generate analysis results.

[0092] Specifically, in this embodiment, the analysis result is the result formed after performing multimodal analysis on historical videos, and the analysis result includes user preference vectors. and its adjoint vector

[0093] Step S103: Construct a video recommendation knowledge graph based on a preset resource library and a preset construction method.

[0094] Specifically, in this embodiment, the preset resource library is a pre-set database that stores a large amount of data resources, including user historical data and related documents; the preset construction scheme is a pre-set method for constructing a video recommendation knowledge graph.

[0095] In this embodiment, constructing a video recommendation knowledge graph includes extracting and filtering relevant literature content from the knowledge graph, then constructing a subject-verb-object structure using word segmentation to build corresponding knowledge graph triples, and finally merging these triples to construct a general knowledge graph. The triples in the knowledge graph are defined as <entity 1, relation, entity 2>. Based on the constructed knowledge graph, all entities and relations are extracted and represented by triples to form a master knowledge graph table, i.e., the video recommendation knowledge graph.

[0096] Select the node types in the knowledge graph used as video tags, such as "Year of Origin," "Person," "Place of Origin," "Alias," "Address," "Object," "Sect," "Cultural Gene," "Concept," and "Level." Extract all nodes under these types and their corresponding indices in the knowledge graph to form a knowledge graph entity table corresponding to the video tags. It is worth noting that entity 1 and entity 2 in a certain triple have a first-degree relationship. If there are <entity 1, relation, entity 2> and <entity 2, relation, entity 3>, then entity 1 and entity 3 have a second-degree relationship, and so on.

[0097] Step S104: Based on the analysis results and the video recommendation knowledge graph, obtain the videos to be recommended.

[0098] Specifically, in this embodiment, the video to be recommended is the video that corresponds to the user preferences of the target user after comprehensive analysis based on the analysis results and the video recommendation knowledge graph.

[0099] Step S105: Recommend the target video to the target user.

[0100] The cross-modal short video recommendation method provided in this embodiment first obtains the target user's user preferences, and then performs multimodal analysis on the historical videos watched based on these preferences to generate analysis results. Relevant resources are then selected from a preset resource library, and a short video recommendation knowledge graph is constructed according to a preset framework method. Finally, by combining the generated analysis results and the constructed video recommendation knowledge graph, videos to be recommended are obtained and recommended to the target user. Performing multimodal analysis on historical videos based on user preferences and combining data from different modalities using the video recommendation knowledge graph helps improve the accuracy of short video recommendations to users, thereby contributing to the achievement of high-accuracy content video recommendations.

[0101] Reference Figure 2 In one embodiment of this example, the specific steps of obtaining the target user's user preferences in step S101 include steps S201 to S206:

[0102] Step S201: Obtain the target user's browsing history.

[0103] Specifically, in this embodiment, historical browsing records refer to the video and data records that the target user browsed on the specified APP within a certain historical time threshold range (such as one year or other time periods).

[0104] Step S202: Based on historical browsing records, obtain text reading records, image viewing records, and audio listening records.

[0105] Specifically, in this embodiment, text reading record is the record of reading text, image viewing record is the record of viewing images, and audio listening record is the record of listening to audio.

[0106] It is worth noting that in this embodiment, after obtaining the historical browsing records, the data needs to be processed, that is, by processing the data based on rules, abnormal data is removed and the data format is unified, specifically including data cleaning, deduplication and data format conversion.

[0107] Data cleaning involves cleaning up invalid or abnormal data, such as removing obviously illogical data like age and occupation from user registration information, or data that is incomplete due to user misoperation or abnormal interruption; deduplication involves removing duplicate data, such as different operations at the same time and identical duplicate records; and data format conversion involves unifying the data record format and content to the same data format.

[0108] Step S203: Based on the text reading record, obtain user preferences based on text modality as text preferences.

[0109] Specifically, in this embodiment, text preferences are user preferences based on image modalities.

[0110] Step S204: Based on the image viewing history, obtain user preferences based on image modality as image preferences.

[0111] Specifically, in this embodiment, image preference refers to user preference based on image modality.

[0112] Step S205: Based on the audio listening record, obtain user preferences based on audio modality as audio preferences.

[0113] Specifically, in this embodiment, audio preference refers to user preference based on audio modality.

[0114] Step S206: Obtain user preferences based on text preferences, image preferences, and audio preferences.

[0115] Specifically, in this embodiment, user preferences include text preferences, image preferences, and audio preferences. Combining text preferences, image preferences, and audio preferences can be considered as user preferences.

[0116] The cross-modal short video recommendation method provided in this embodiment obtains text reading records, image viewing records, and audio listening records based on the target user's historical browsing history. Then, it obtains the corresponding user preferences based on the text reading records, image viewing records, and audio listening records respectively. Finally, it combines the user preferences from all different modalities as the final user preference. Combining user preferences from multiple modalities as the final user preference results in more complete data. Using this as the basis for video recommendation helps improve the accuracy of video content.

[0117] Reference Figure 3 In one embodiment of this example, step S203, which involves obtaining user preferences based on text modality from text reading records, specifically includes steps S301 to S308:

[0118] Step S301: Obtain the reading text based on the text record.

[0119] Specifically, in this embodiment, the text being read refers to the text that the target user has already read.

[0120] Step S302: Based on the preset word segmentation rules, segment the reading text and obtain the segmented word groups.

[0121] Specifically, in this embodiment, the preset word segmentation rule is a pre-set rule for segmenting the reading text into words; according to the preset word segmentation rule, the word groups formed after segmenting the reading text into words are the segmented word groups.

[0122] Step S303: Obtain the word frequency and weight of the segmented word groups.

[0123] Specifically, in this embodiment, word frequency refers to the number of times different word segments appear; weight refers to the weight of different word segments. In this embodiment, the weight can be preset or calculated by word frequency.

[0124] Step S304: Based on word frequency and weight, obtain keywords and generate text keyword vectors based on the keywords.

[0125] Specifically, in this embodiment, keywords are word segments corresponding to content that the user is interested in, and text keyword vectors are vectors composed of keywords in the reading text.

[0126] Step S305: Based on the text records, obtain the target number and reading time for different reading content.

[0127] Specifically, in this embodiment, the target quantity is the count c of each part of the content (the content presented in the reading window); the reading duration is the length of time t1 spent in different windows.

[0128] Step S306: Based on the reading content, obtain the number of keywords for different keywords.

[0129] Specifically, in this embodiment, the number of keywords is the number of times different keywords appear, f.

[0130] Step S307: Obtain the weighted reading time based on the target number, reading time, number of keywords, and a preset calculation function.

[0131] Specifically, in this embodiment, the preset calculation function is a pre-set function, which is stored in a function library containing a large number of various functions used for calculation; the weighted reading time is the weighted dwell time of keywords in the reading window, and the specific calculation formula is as follows: The Count() function is defined as a function to calculate the frequency of keywords.

[0132] Step S308: Obtain text preferences based on weighted reading time.

[0133] Specifically, in this embodiment, the data from all reading windows is accumulated to obtain the weighted reading time of keywords and sort them. In this embodiment, user preferences are derived from the dwell time from long to short, that is, the longer the dwell time, the more the user likes it.

[0134] The cross-modal short video recommendation method provided in this embodiment obtains the reading text, the target number of different reading contents, and the reading duration based on the text record. It then segments the reading text into words and obtains word frequencies and weights. Based on these word frequencies and weights, it generates a text keyword vector. According to the reading content, it obtains the number of different keywords. Finally, combining the target number, reading duration, and keyword number, it calculates a weighted reading time using a preset calculation function and obtains text preferences based on this weighted reading time. The weighted reading time helps identify the reading content that the target user is more interested in, and video recommendations are made based on this content, thereby improving the accuracy of video content recommendations.

[0135] Reference Figure 4 In one embodiment of this example, step S204, which obtains user preferences based on image modality as image preferences based on image viewing records, includes steps S401 to S407:

[0136] Step S401: Obtain the target image based on the image viewing history.

[0137] Specifically, the target image refers to the image viewed by the target user. In this embodiment, it can be obtained from existing apps and related web pages, or it can be supplemented from the background database.

[0138] Step S402: Obtain the image style of the target image.

[0139] Specifically, in this embodiment, the image style includes a wide range of content, such as the tags for sunset, fast, and blooming flowers. The relevant content of the image can be obtained through the image's tags and title, and the image style can be obtained through an artificial intelligence training model.

[0140] Step S403: Identify the target in the image based on the target image and the preset recognition method.

[0141] Specifically, in this embodiment, the image target refers to the target in the image, which can be an animal, a person, or a scenic spot, etc., while also incorporating text descriptions and text information within the image. This embodiment can employ a pre-trained model and further train and recognize it using labeled specific image features.

[0142] Step S404: Generate image keyword vectors based on image style and image target.

[0143] Specifically, in this embodiment, the image keyword vector is a vector composed of image style and image target.

[0144] Step S405: Based on the image viewing history, obtain the total number of images and the viewing duration of different target images.

[0145] Specifically, in this embodiment, the total number of images is the total number of target images viewed by the target user, h1; the viewing time is the duration t2 spent viewing different target images.

[0146] Step S406: Based on the total number of images, viewing duration, image keyword vectors, and a preset calculation function, obtain the viewing time of key images.

[0147] Specifically, in this embodiment, the viewing time of key images, i.e., the dwell time of image keywords, is calculated using the following formula:

[0148] Step S407: Obtain image preferences based on key image viewing time.

[0149] Specifically, in this embodiment, the keyword data of all images is accumulated to obtain the dwell time of the keywords and sort them. The user's preference is determined by the dwell time from longest to shortest, that is, the longer the dwell time, the more the user likes it.

[0150] The cross-modal short video recommendation method provided in this embodiment obtains the target image, the total number of images, and the viewing duration of different target images based on image viewing records. It also obtains the image style of the target image, identifies the image target in the target image, and generates image keyword vectors based on the image style and image target. Finally, it combines the total number of images, viewing duration, and image keyword vectors to calculate the key image viewing time using a preset calculation function, and obtains image preferences based on the key image viewing time. Based on the key image viewing time, it helps to identify images that target users are more interested in, and recommends videos based on these images, thereby improving the accuracy of video content.

[0151] Reference Figure 5 In one embodiment of this example, step S205, which involves obtaining user preferences based on audio modality from audio listening records, includes steps S501 to S506:

[0152] Step S501: Obtain the target audio based on the audio listening record.

[0153] Specifically, in this embodiment, the target audio is the audio that the target user has listened to.

[0154] Step S502: Extract features from the target audio and generate the extracted features.

[0155] Specifically, feature extraction involves extracting features from the target audio. In this embodiment, the extracted features can be obtained by using Mel-frequency cepstral coefficients (MFCC) for speech classification and then using deep learning models (such as CNN, RNN, etc.) for feature extraction.

[0156] Step S503: Generate audio keyword vectors based on extracted features.

[0157] Specifically, in this embodiment, the audio keyword vector is a vector generated based on the extracted features.

[0158] Step S504: Based on the audio listening records, obtain the listening time of different audio and the total listening time of all audio.

[0159] Specifically, in this embodiment, the listening time is the time t3 for listening to different audios; the total time is the total listening time h2 for all audios.

[0160] Step S505: Based on the audio keyword vector, listening time, total time, and preset calculation function, obtain the listening time of key audio.

[0161] Specifically, in this embodiment, the key audio listening time, i.e. the dwell time of audio keywords, is calculated using the following formula:

[0162] Step S506: Obtain audio preferences based on key audio listening times.

[0163] Specifically, in this embodiment, the keyword data of all audios is accumulated to obtain the dwell time of the keywords and sort them. The user's preference is determined by the dwell time from long to short, that is, the longer the dwell time, the more the user likes it.

[0164] The cross-modal short video recommendation method provided in this embodiment obtains the listening time of the target audio, different audios, and the total listening time of all audios based on audio listening records. It then extracts features from the target audio and generates an audio keyword vector. Finally, it combines the audio keyword vector, listening time, and total time to calculate the key audio listening time using a preset calculation function, and obtains audio preferences based on this key audio listening time. Based on the key audio listening time, it helps to identify the audio that the target user is more interested in, and recommends videos based on the audio that the user is more interested in, thereby helping to improve the accuracy of video content.

[0165] ReferenceFigure 6 In one embodiment of this example, step S102, which involves performing multimodal analysis on historical videos based on user preferences and generating analysis results, includes steps S601 to S607:

[0166] Step S601: Obtain feature targets based on historical videos.

[0167] Specifically, the feature target is the target in the video content.

[0168] Step S602: Analyze the feature targets and generate feature labels corresponding to the feature targets.

[0169] Specifically, in this embodiment, object detection and recognition algorithms (such as YOLO, Faster R-CNN, etc.) are used to identify objects in the video, and the categories of these objects are used as text labels; scene recognition algorithms (such as Places-CNN) are used to identify scenes in the video, and the scene type is used as a text label; action recognition algorithms (such as I3D, C3D, etc.) are used to identify actions in the video, and the action category is used as a text label; optical character recognition (OCR) technology (such as Tesseract, Google OCR, etc.) is used to identify text content in the video and use it as a label.

[0170] Step S603: Generate a tag keyword vector based on the feature labels.

[0171] Specifically, in this embodiment, the tag keyword vector is a vector generated based on the feature tags.

[0172] Step S604: Based on historical videos, obtain the viewing duration of different videos.

[0173] Specifically, in this embodiment, the viewing duration is the viewing duration t4 of each video.

[0174] Step S605: Obtain the total video viewing time based on all viewing durations.

[0175] Specifically, in this embodiment, the total video viewing time is the total viewing time h3 of all videos.

[0176] Step S606: Based on the tag keyword vector, viewing duration, total video viewing duration, and preset calculation function, obtain the viewing time of key tags.

[0177] Specifically, in this embodiment, the viewing time of the key tag, i.e., the dwell time of the tag keyword, is calculated using the following formula:

[0178] Step S607: Based on the obtained text keyword vectors, image keyword vectors, audio keyword vectors, tag keyword vectors, and preset calculation functions, obtain the target vector and use the target vector as the analysis result.

[0179] Specifically, in this embodiment, based on the text keyword vector Image keyword vector Audio keyword vectors and tag keyword vector The user's preference vector, a high-dimensional vector composed of relevant data labels, is obtained through a preset calculation function. The calculation formula is as follows: Then based on the preference vector Weighted reading time Key image viewing time Key audio listening time and viewing time of key tags The adjoint vector is calculated. The calculation formula is as follows:

[0180] Wherein, ∪ indicates the join operation. Let Wf, Wp, Wa, and Wv represent the vector inner product, and Wf, Wp, Wa, and Wv represent the weights of each mode, respectively.

[0181] In this embodiment, the target vector is... and

[0182] The cross-modal short video recommendation method provided in this embodiment obtains feature targets from historical videos and generates feature tag vectors corresponding to the feature targets. It then obtains the viewing duration and total viewing duration of different videos. Finally, combining the tag keyword vectors, viewing durations, and total viewing durations, it calculates the viewing duration of key tags using a preset calculation function. Based on this key tag viewing duration, text keyword vectors, image keyword vectors, audio keyword vectors, tag keyword vectors, and the preset calculation function, it obtains a target vector, which is used as the analysis result. Finally, based on the text keyword vectors, image keyword vectors, audio keyword vectors, and tag keyword vectors, it obtains a high-dimensional vector composed of relevant data tags. Video recommendations are then performed based on this high-dimensional vector, thereby helping to improve the accuracy of video content.

[0183] Reference Figure 7 In one embodiment of this example, step S104, based on the analysis results and the video recommendation knowledge graph, specifically involves steps S701 to S703 to obtain the video to be recommended:

[0184] Step S701: Obtain the keyword weight vector based on the tag keyword vector.

[0185] Specifically, in this embodiment, the specific formula for calculating the keyword weight vector is as follows:

[0186] Where t5 is the duration of each modality in the video watched by the target user; h4 is the total duration of a single watched video.

[0187] Step S702: Obtain the target mapping relationship based on the target vector, tag keyword vector, keyword weight vector, and preset calculation function.

[0188] Specifically, in this embodiment, the target mapping relationship is determined by... and arrive and The mapping relationship is calculated using the following formula: Where ∩ represents the intersection operation.

[0189] Step S703: Obtain the videos to be recommended based on the target mapping relationship and the video recommendation knowledge graph.

[0190] Specifically, in this embodiment, for Augmentation based on knowledge graphs involves introducing relationships of one degree or higher to add attributes to two vectors, thereby augmenting their weight vectors accordingly. The augmented data in one dimension of the vectors are then assigned the same original weights. It is a vector inclusion operation, C n Here, C2 is an n-degree association operator, and Kg is a knowledge graph; for example, C2 is defined as... right Each dimension item is input into the knowledge graph one by one. If the corresponding entity exists in the knowledge graph, then entities with first-degree or higher relations to that entity are included in the user preference vector. It is then assigned a weight equivalent to that of the initial entity.

[0191]

[0192] right Augmentation based on the knowledge graph involves introducing relationships of one degree or higher to add attributes to two vectors, thereby augmenting the weight vector accordingly. The augmented data from one dimension of the vector is then assigned the same original weight. Each dimension is input into the knowledge graph one by one. If a corresponding entity exists in the knowledge graph, then entities with first-degree or higher relations to that entity are included in the video vector. It then assigns it the same rights as the initial entity.

[0193]

[0194] The augmented vector is matched, and the matched vector is...

[0195] Recommend association strength to users based on association strength. Videos sorted from highest to lowest.

[0196] The cross-modal short video recommendation method provided in this embodiment obtains a target mapping relationship based on the target vector, tag keyword vector, keyword weight vector, and preset calculation function. This allows user preferences in one modality to be mapped to user preferences in other different modalities. Then, based on the target mapping relationship and the video recommendation knowledge graph, the video to be recommended is obtained, thereby realizing cross-modal video recommendation. Video recommendation based on the mapping relationship and the video recommendation knowledge graph helps to improve the accuracy of video content recommendation.

[0197] Secondly, this application also discloses a cross-modal short video recommendation system.

[0198] Reference Figure 8 A cross-modal short video recommendation system, comprising:

[0199] The first acquisition module 1 is used to acquire the user preferences of the target user;

[0200] Module 2 is used to perform multimodal analysis on historical videos based on user preferences and generate analysis results;

[0201] Module 3 is used to build a video recommendation knowledge graph based on a preset resource library and preset construction methods.

[0202] The second acquisition module 4 is used to acquire videos to be recommended based on the analysis results and the video recommendation knowledge graph;

[0203] Recommendation module 5 is used to recommend target videos to target users.

[0204] Thirdly, this application discloses a smart terminal, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor loads the computer program, it executes a cross-modal short video recommendation method according to the above embodiment.

[0205] Fourthly, embodiments of this application disclose a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is loaded by a processor, it executes a cross-modal short video recommendation method of the above embodiments.

[0206] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A cross-modal short video recommendation method, characterized in that, include: Obtain the user preferences of the target users; Based on the user preferences, multimodal analysis is performed on historical videos and analysis results are generated; The specific steps for performing multimodal analysis on historical videos based on the user preferences and generating analysis results include: obtaining feature targets based on the historical videos; analyzing the feature targets and generating feature tags corresponding to the feature targets; generating tag keyword vectors based on the feature tags; obtaining the viewing duration of different videos based on the historical videos; obtaining the total video viewing duration based on all the viewing durations; obtaining the key tag viewing duration based on the tag keyword vectors, the viewing durations, the total video viewing durations, and a preset calculation function; obtaining a target vector based on the obtained text keyword vectors, image keyword vectors, audio keyword vectors, tag keyword vectors, and the preset calculation function, and using the target vector as the analysis result. Based on a pre-defined resource library and a pre-defined construction method, a video recommendation knowledge graph is constructed. Based on the analysis results and the video recommendation knowledge graph, obtain the videos to be recommended; The specific steps for obtaining the video to be recommended based on the analysis results and the video recommendation knowledge graph include: obtaining a keyword weight vector based on the tag keyword vector; obtaining a target mapping relationship based on the target vector, the tag keyword vector, the keyword weight vector, and the preset calculation function; and obtaining the video to be recommended based on the target mapping relationship and the video recommendation knowledge graph. The video to be recommended is recommended to the target user.

2. The cross-modal short video recommendation method according to claim 1, characterized in that, The specific steps for obtaining the target user's user preferences include: Obtain the target user's browsing history; Based on the historical browsing records, obtain text reading records, image viewing records, and audio listening records; Based on the text reading records, the user preferences based on text modality are obtained as text preferences; Based on the image viewing history, the user preferences based on image modality are obtained as image preferences; Based on the audio listening record, the user preference based on the audio modality is obtained as the audio preference; User preferences are obtained based on the text preferences, image preferences, and audio preferences.

3. The cross-modal short video recommendation method according to claim 2, characterized in that, The specific steps for obtaining user preferences based on text modality as text preferences based on the text reading records include: Based on the text reading records, obtain the read text; Based on preset word segmentation rules, the reading text is segmented into words, and segmented word groups are obtained; Obtain the word frequency and weight of the segmented word groups; Based on the word frequency and the weight, keywords are obtained, and text keyword vectors are generated based on the keywords; Based on the text reading records, obtain the target quantity and reading time for different reading contents; Based on the reading content, obtain the number of keywords for different keywords; Based on the target number, the reading duration, the number of keywords, and a preset calculation function, obtain the weighted reading time; Based on the weighted reading time, text preferences are obtained.

4. The cross-modal short video recommendation method according to claim 2, characterized in that, The specific steps for obtaining user preferences based on image modality from the image viewing history include: Based on the image viewing history, obtain the target image; Obtain the image style of the target image; Based on the target image and the preset recognition method, identify the target in the image; Based on the image style and the image target, generate an image keyword vector; Based on the image viewing history, obtain the total number of images and the viewing duration of different target images; Based on the total number of images, the viewing duration, the image keyword vector, and the preset calculation function, the viewing time of key images is obtained; Image preferences are obtained based on the viewing time of the key images.

5. The cross-modal short video recommendation method according to claim 2, characterized in that, The specific steps for obtaining the user preference based on the audio modality as audio preference based on the audio listening record include: Based on the audio listening record, the target audio is obtained; Feature extraction is performed on the target audio, and the extracted features are generated; Based on the extracted features, an audio keyword vector is generated; Based on the audio listening records, the listening time of different audio and the total listening time of all audio are obtained; Based on the audio keyword vector, the listening time, the total time, and the preset calculation function, the listening time of key audio is obtained; Based on the key audio listening time, audio preferences are obtained.

6. A cross-modal short video recommendation system, characterized in that, The system is used to execute a cross-modal short video recommendation method as described in claim 1, including: The first acquisition module (1) is used to acquire the user preferences of the target user; The generation module (2) is used to perform multimodal analysis on historical videos based on the user preferences and generate analysis results; Module (3) is used to build a video recommendation knowledge graph based on a preset resource library and a preset construction method. The second acquisition module (4) is used to acquire the video to be recommended based on the analysis results and the video recommendation knowledge graph; The recommendation module (5) is used to recommend the video to be recommended to the target user.

7. A smart terminal, comprising a memory and a processor, characterized in that, The memory is used to store computer programs that can run on the processor, and when the processor loads the computer program, it executes the method of any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is loaded by the processor, it executes the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Label-based cross-modal resource recommendation method and system

    CN114722275A

  • Resource determination method and device, electronic equipment and storage medium

    CN114741502A