Recommendation method for multimedia data, recommendation model training method and related devices

By extracting and fusion of multimedia data and user browsing records, combined with attention mechanism, deep fusion features are generated, and the existing recommendation algorithms are solved in the deep feature expression, sparsity, and cold start problems, and accurate personalized recommendations are achieved.

CN114610913BActive Publication Date: 2025-07-04ASIAINFO TECH CHINA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111681672.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-04
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The existing recommendation algorithms rely on the construction of user characteristics and recommended content features, and it is difficult to fully express the deep-level characteristics of recommended content, and it is powerless in terms of sparseness and cold start issues.

Method used

By obtaining multimedia data and user browsing records, feature extraction is performed separately, feature fusion is performed using graph embedding models and representation learning models, and combining time and visual attention mechanisms to generate deep fusion features for recommendation.

Benefits of technology

It realizes a deep exploration of precise capture of user interests and hobbies and multimedia data correlation. The recommended results are well interpreted and solves the problems of sparseness and cold start.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114610913B_ABST
    Figure CN114610913B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for recommending multimedia data, a method for training a recommendation model, and related devices; relating to the field of artificial intelligence technology. The method includes: obtaining at least one piece of multimedia data and at least one browsing record related to a target user, then respectively performing feature extraction on each browsing record to obtain corresponding first vector representations, respectively performing feature extraction on each piece of multimedia data to obtain corresponding second vector representations, and performing feature fusion on the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fused feature, and then determining the multimedia data for recommendation from at least one piece of multimedia data based on the fused feature. The implementation of the present application not only helps to discover the deep interest points of users, but also helps to solve the problems of sparsity and cold start.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence and recommendation technology. Specifically, the present application relates to a method for recommending multimedia data, a method for training a recommendation model, and related devices. Background Art

[0002] With the development of technology, the wide application of recommendation systems in fields such as Internet e-commerce, search engines, and video websites has brought huge economic benefits, enabling enterprises to start paying attention to the research and application of recommendation algorithms and focusing on the improvement of user experience.

[0003] Traditional recommendation algorithms mostly adopt collaborative filtering or content-based recommendation algorithms. For example, algorithms that recommend based on user personality characteristics (including interests, hobbies, and behavior habits), users similar to him, the similarity between recommended contents, and representative characteristics of the recommended contents themselves.

[0004] However, the above-mentioned recommendation algorithms rely on the construction and similarity characterization of user features and recommended content features, and only consider the superficial connections between recommended contents (or users) more. It is difficult to provide explanations for the recommended contents. Obviously, how to extract the deep features of recommended contents has become a major technical problem in the field of recommendation technology, and the above-mentioned recommendation algorithms also rely on the habitual data of users and are difficult to solve the problems of sparsity and cold start. Summary of the Invention

[0005] Embodiments of the present application provide a method for recommending multimedia data, a method for training a recommendation model, and related devices, which are used to solve the technical problem of how to recommend based on the deep content features of recommended contents.

[0006] According to one aspect of the embodiments of the present application, a method for recommending multimedia data is provided. The method includes:

[0007] Obtain at least one piece of multimedia data and at least one browsing record related to a target user;

[0008] Extract features from each browsing record respectively to obtain corresponding first vector representations;

[0009] Extract features from each piece of multimedia data respectively to obtain corresponding second vector representations;

[0010] Perform a fusion step for the first vector representation corresponding to each browsing record: fuse the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fusion feature;

[0011] Determine the multimedia data for recommendation from at least one piece of multimedia data based on the fusion feature.

[0012] In a possible implementation, the fusion step further includes:

[0013] Adjusting the fused features by using the feature weights determined based on the visual attention mechanism and / or the temporal weights determined based on the temporal attention mechanism.

[0014] In a possible implementation, feature extraction is respectively performed on each browsing record to obtain corresponding first vector representations, including:

[0015] Generating a node sequence for each node corresponding to a browsing record; the distance information and similarity information between the nodes in the node sequence respectively have corresponding random walk weights;

[0016] Generating a corresponding first vector representation based on the node sequence.

[0017] In a possible implementation, feature extraction is respectively performed on each multimedia data to obtain corresponding second vector representations, including:

[0018] Generating a vector matrix for each multimedia data;

[0019] Extracting the implicit features of the multimedia data from the vector matrix;

[0020] Generating a corresponding second vector representation based on the implicit features.

[0021] In a possible implementation, determining the multimedia data for recommendation from at least one multimedia data based on the fused features includes:

[0022] Calculating the similarity between the fused features and the second vectors corresponding to the multimedia data not in the target user's browsing record to determine at least one similar multimedia data for recommendation;

[0023] And / or calculating the similarity between the fused features and the fused features of other users to determine at least one similar user, and determining at least one similar multimedia data not in the target user's browsing record based on the browsing records of the similar users.

[0024] According to another aspect of the embodiments of the present application, a training method for a recommendation model is provided. The recommendation model includes a first feature extraction module for extracting features of user browsing records and a second feature extraction module for extracting features of multimedia data; the training method includes:

[0025] Obtaining a training data set; the training data set includes at least one multimedia data and at least one browsing record related to a user;

[0026] Through the first feature extraction module, based on the browsing record, obtaining a predicted first vector representation;

[0027] Through a second feature extraction module, based on the multimedia data, a predicted second vector representation is obtained;

[0028] Update the recommendation model based on the predicted first vector and the predicted second vector;

[0029] Wherein, the trained recommendation model is applied to the above-mentioned multimedia data recommendation method.

[0030] According to another aspect of the embodiments of the present application, a multimedia recommendation device is provided, including:

[0031] A data acquisition module, configured to acquire at least one piece of multimedia data and at least one browsing record related to a target user;

[0032] A first feature extraction module, configured to perform feature extraction on each browsing record respectively to obtain a corresponding first vector representation;

[0033] A second feature extraction module, configured to perform feature extraction on each piece of multimedia data respectively to obtain a corresponding second vector representation;

[0034] A fusion module, configured to perform a fusion step for the first vector representation corresponding to each browsing record: perform feature fusion on the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fusion feature;

[0035] A determination module, configured to determine the multimedia data for recommendation from at least one piece of multimedia data based on the fusion feature.

[0036] According to another aspect of the embodiments of the present application, a training device for a recommendation model is provided, where the recommendation model includes a first feature extraction module for extracting the features of multimedia data in a user's browsing record and a second feature extraction module for extracting the features of multimedia data in a database; the training device includes:

[0037] An acquisition module, configured to acquire a training data set; the training data set includes at least one piece of multimedia data and at least one browsing record related to a user;

[0038] A training module, configured to, through the first feature extraction module, based on the browsing record, obtain a predicted first vector representation; through the second feature extraction module, based on the multimedia data, obtain a predicted second vector representation; update the recommendation model based on the predicted first vector representation and the predicted second vector representation;

[0039] Wherein, the trained recommendation model is applied to the above-mentioned multimedia data recommendation method.

[0040] According to another aspect of the embodiments of the present application, a computer device is provided, including:

[0041] One or more memories;

[0042] A processor and a computer program stored in the memory,

[0043] The processor executes the computer program to implement the steps of the above method.

[0044] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, including:

[0045] A computer program stored thereon,

[0046] When the computer program is executed by a processor, the steps of the above method are implemented.

[0047] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0048] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:

[0049] The embodiments of the present application provide a method for recommending multimedia data. Specifically, the present application can respectively extract features from at least one browsing record of a target user to obtain a corresponding first vector representation, so as to explore the features related to the multimedia data by the browsing record of the target user; at the same time, features can be respectively extracted from at least one multimedia data to obtain a corresponding second vector representation; then, the first vector representation corresponding to the browsing record and the second vector representation corresponding to the multimedia data corresponding to the browsing record are feature-fused to obtain a fused feature, so as to perform multimedia data recommendation based on the fused feature, making the recommendation method of the present application not only can accurately capture the user's interests and hobbies, but also can deeply explore the correlation between the user and the multimedia data, such as the user's interest points in the multimedia data, etc.; in addition, the present application performs multimedia data recommendation based on the extracted first vector representation and second vector representation, making the recommendation result have better interpretability, and no longer relying on the user's habitual data to depict the features of the multimedia data or the user, which helps to solve the problems of sparsity and cold start. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application.

[0051] Figure 1 It is a schematic flowchart of a method for recommending multimedia data provided by an embodiment of the present application;

[0052] Figure 2Schematic flowchart of a training method for a recommendation model provided by an embodiment of the present application;

[0053] Figure 3 Schematic diagram of a random walk process provided by an embodiment of the present application;

[0054] Figure 4 Schematic diagram of the system architecture of a recommendation system provided by an embodiment of the present application;

[0055] Figure 5 Schematic diagram of an application scenario of a recommendation method for multimedia data provided by an embodiment of the present application;

[0056] Figure 6 Schematic diagram of the structure of a training device for a recommendation model provided by an embodiment of the present application;

[0057] Figure 7 Schematic diagram of the structure of a multimedia data recommendation device provided by an embodiment of the present application;

[0058] Figure 8 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0059] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0060] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0061] For a better understanding and illustration of the solutions provided by the embodiments of the present application, the related technologies involved in the present application will be described below.

[0062] Regarding the user-based collaborative filtering recommendation algorithm, it relies on professionals to construct user features. Even if the construction of features is very sophisticated, there are still problems in capturing some user features. That is to say, the recall rate of user preferences is relatively low. At the same time, this recommendation algorithm also relies on user behavior data, and there are problems of sparsity and cold start.

[0063] Regarding the content-based collaborative filtering recommendation algorithm, it relies on professionals to construct the features of the recommended content. If recommendations are made only based on the features of the recommended content, it is possible to keep recommending content that the user does not like although the features are closely related, thus losing the diversity of recommendations.

[0064] Regarding the content-based recommendation algorithm, it is easily limited by the detailed degree of description of the recommended content, and there are still problems in capturing some specific features of the content (people usually have difficulty explaining what they like about the things they like).

[0065] According to the above content, it can be seen that the existing recommendation algorithms rely heavily on the artificial construction of the features of the recommended content. The constructed features not only have the technical problem of being difficult to comprehensively express the recommended content, but also have the problem of being difficult to capture the deep interests of users. At the same time, the recommendation algorithms based on the recommended features require a large amount of user data and are also powerless in the face of the technical problems of sparsity and cold start. Therefore, the embodiments of this application provide a recommendation method, a recommendation model training method and related devices for multimedia data to solve at least one of the above problems.

[0066] To make the purpose, technical solution and advantages of this application clearer, the following introduces and explains several terms involved in this application:

[0067] Graph embedding model: Maps the original graph data (usually a sparse high-dimensional adjacency matrix) into a low-dimensional dense vector. There are generally two types of graph embedding models: Node embedding is to obtain a vector representation for each node on the original graph through embedding. This vector representation can be used in simple tasks, such as judging the similarity of two nodes. In a social network, it can be used to judge whether two users are similar and then recommend friends. It can also be used as an input representation in more upstream complex tasks, such as using the representation of nodes (users) in a social network in a commodity recommendation system; Whole graph embedding is to obtain a vector for the whole graph through embedding, which is generally used to compare the similarity of two graphs, such as judging the similarity of protein molecules and whether two communities are similar.

[0068] Representation learning model: In the field of deep learning, representation refers to the form and method by which the input observation samples X of the model are represented through the parameters of the model. Representation learning means learning an effective representation for the observation samples X. There are many forms of representation learning. For example, the supervised training of the parameters of CNN (Convolutional Neural Networks) is a form of supervised representation learning, the unsupervised pre-training of the parameters of autoencoders and restricted Boltzmann machines is a form of unsupervised representation learning, and the semi-supervised shared representation learning form is to first perform unsupervised pre-training on the parameters of DBN (Deep neural network) and then perform supervised fine-tuning.

[0069] Temporal attention mechanism: Imitating the human visual attention mechanism, it learns a weight distribution for the features of multimedia data, and then applies these weight distributions to the original features, providing different feature influences for recommendations based on multimedia data and user-based recommendations, etc., so that the task mainly focuses on some important features, ignores unimportant features, and improves the task efficiency. Among them, the visual attention mechanism is a unique signal processing mechanism of the human brain. Human vision observes the global image, selects some local key attention areas, and then invests more attention in these areas to obtain more detailed information and suppress other useless information.

[0070] Homogeneity of the network: The embeddings of nodes that are close in distance should be as similar as possible. In practical applications, programs with homogeneity are likely to be of the same category, the same attribute, or programs that are often purchased or clicked together.

[0071] Structure of the network: The embeddings of nodes that are similar in structure should be as close as possible. In a recommendation system, items with similar structures are generally hit shows in various categories and other programs with similar trends or structural attributes.

[0072] Next, through the description of several exemplary embodiments, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0073] The recommendation method involved in the embodiments of the present application is a recommendation method based on multi-model fusion. Its special feature is that the involved graph embedding model and representation learning model complement each other's advantages, not only solving the technical problem of the difficulty in extracting the content features of multimedia data, but also helping to alleviate problems such as sparsity and cold start. The ultimate goal of the recommendation based on the multi-model fusion recommendation method in the embodiments of the present application is to make the recommendation result have good interpretability.

[0074] Please refer to Figure 1 , and the following will describe in detail the recommended method of the embodiments of the present application in conjunction with Figure 1 . Figure 1 FIG. shows a schematic flow chart of a recommended method for multimedia data provided by an embodiment of the present application. This method can be executed by any electronic device. For example, it can be executed by a terminal device. The terminal device can execute this method to perform feature fusion based on the first vector representation of the browsing record of the target user and the second vector representation of the multimedia data corresponding to the browsing record, so as to obtain a fusion feature. Subsequently, multimedia data for recommendation can be determined based on this fusion feature. This method can also be executed by a server. Optionally, the server can be a cloud server. This method can be implemented as an application program or as a plugin or functional module of an existing application program with a recommendation function. For example, it can be a new functional module of a multimedia application program. By executing the method of the embodiments of the present application, for different multimedia data or users, deeper multimedia data features or user features can be characterized. Subsequently, multimedia data for recommendation can also be determined based on these features and pushed to the user's terminal device for display to the user; more multimedia data features can be learned based on this method, which is beneficial to deeply discover the user's interest points, is beneficial to solving problems such as sparsity and cold start, and provides more accurate recommended content for users. Among them, the above-mentioned terminal device includes a user terminal, and the user terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, wearable electronic devices, AR / VR devices, etc.

[0075] As Figure 1 shown in, the recommended method for the multimedia data provided by the embodiments of the present application may include the following steps S100 to step S500.

[0076] Step S100: Obtain at least one piece of multimedia data and at least one browsing record related to the target user.

[0077] Specifically, the browsing record related to the target user refers to the browsing traces left by the target user when using an application program or a small program. Among them, the browsing record of the target user can be used to characterize the browsing behavior of the target user for multimedia data. It can be understood that the objects browsed by the target user can include TV programs, short videos, tweets, advertisements, etc. In addition, the browsing record not only includes the objects browsed by the user, but also can include information such as the browsing time and the number of browsing times of the user for the object.

[0078] Among them, the multimedia data refers to the browsable objects provided on the content providing platform, including but not limited to the media asset information of the browsable objects (including: introduction, classic lines, keywords of classic scenes), comment data, types of browsable objects, etc. It should be noted that the multimedia data in the above step S100 includes the browsed objects in the target user's browsing record and the browsable objects not in the user's browsing record (that is, including the multimedia data that has been browsed and not browsed by the target user).

[0079] Optionally, the obtained browsing record of the target user includes: the browsing record that meets the preset requirements obtained by filtering the browsing record based on at least one of a preset time period and the behavior data of the target user.

[0080] Specifically, the preset time period refers to a preset time cycle. Among them, the preset time cycle includes but is not limited to one week, one month, one quarter, etc. It can be understood that the duration of the hot content can also be used as the preset time cycle.

[0081] Specifically, the behavior data of the target user refers to user click data, search data, comment data, browsing duration, etc.

[0082] In a possible case, after obtaining the browsing record of the target user, the browsing record of the target user can be filtered based on the preset time period. For example: set the preset time period to one month, and filter out the total number of browsing records of the target user within one month from all the browsing records of the target user with a month (or 30 days) as the time cycle.

[0083] In another possible case, after obtaining the browsing record of the target user, the browsing record is filtered based on the behavior data of the target user during browsing. For example: personalized feature tags are assigned to the browsing objects based on the browsing duration of the target user. Among them, a browsing duration threshold is set for the browsing duration. When the browsing duration is greater than the browsing duration threshold, a positive tag is assigned to the browsing object, indicating that the browsing object is a positive sample (that is, this browsing record is a positive sample). When the browsing duration is less than or equal to the browsing duration threshold, a negative tag is assigned to the browsing object, indicating that the browsing object is a negative sample (that is, this browsing record is a negative sample). Among them, the positive samples of the target user's browsing record can be selected as the input data of the embodiments of the present application.

[0084] Another possible scenario is that after obtaining the browsing records of the target user, the browsing records of the target user are filtered based on a preset time period and the user's behavior data. For example, the preset time period is set to one month, and the total number of the target user's browsing records within one month is filtered from all the browsing records of the target user in a monthly (or 30-day) time cycle. Then, personalized feature tags are assigned to the browsing objects according to the browsing duration of the target user. Next, the total number of positive samples of the target user's browsing records within one month is filtered out as the input data of the embodiment of the present application.

[0085] Step S200: Feature extraction is performed on each browsing record to obtain a corresponding first vector representation.

[0086] Among them, the first vector representation refers to the vector representation based on the multimedia data in the target user's browsing record, which characterizes the context information of the multimedia data in the user's browsing record.

[0087] Optionally, a node sequence is generated for each node corresponding to a browsing record; the distance information and similarity information between the nodes in the node sequence respectively have corresponding random walk weights; based on the node sequence, a corresponding first vector representation is generated.

[0088] Specifically, the random walk weight refers to the weight on the node network path when randomly walking in the node network in the graph embedding model. Among them, the weight on the path is controlled by the hyperparameters of the graph embedding model, and the specific value of this hyperparameter can be set manually or learned through a semi-supervised method.

[0089] Specifically, the node network refers to a complex network generated according to the browsing records of the user. Among them, this complex network can be a directed graph or an undirected graph, and the present application does not make any restrictions.

[0090] Specifically, nodes that are close in distance in the node network should be similar, that is, the homogeneity of the network. Nodes with similar structures in the node network should also be similar, that is, the structure of the network. In order to balance the homogeneity and structure of the node network, the random walk weight can be adjusted according to the distance between different nodes (or the structure between different nodes), and then a node sequence related to the nodes is obtained. Based on this node sequence, it is input into the recommendation model, so that the recommendation model can learn more context information about the nodes, and then generate a first vector representation based on the context information.

[0091] Specifically, the context information of a node refers to the information closely related to the node in the node network (i.e., other nodes that are close to the node or have a similar structure). Each node represents a browsing record, i.e., a browsing object of a target user. It should be noted that the embodiment of the present application introduces a graph embedding model to extract the features of user browsing objects, so that the recommendation model can learn more features of user browsing objects based on user features based on user browsing records.

[0092] In one possible scenario, the browsing records and multimedia data filtered out based on the preset time period of one month are used as input data and input into an embodiment of the present application. After feature extraction of the browsing records, a first vector representation of the user browsing object based on user features (i.e., a vector representation of the browsing records) is obtained.

[0093] In another possible scenario, the browsing records and multimedia data selected based on the browsing time are used as input data and input into an embodiment of the present application. After feature extraction of the browsing records, a first vector representation of the user browsing object based on user features (i.e., a vector representation of the browsing records) is obtained.

[0094] In another possible scenario, the browsing history of the target user is filtered based on the preset time period and browsing duration, and the browsing history of the target user obtained after filtering is input into an embodiment of the present application as input data, and features are extracted from the browsing history to obtain a first vector representation of the user browsing object based on user features (i.e., a vector representation of the browsing history).

[0095] In other words, the first vector representation obtained by extracting features based on the browsing history of the target user can also be used to characterize the features of the target user. This is because the browsing history is the browsing traces left by the user on the content provider platform based on his or her own interests and hobbies, so the user's features are essentially implied in the user's browsing history. At the same time, the embodiment of the present application can filter the browsing history of the target user based on the preset time period and the user's behavior data, remove the dirty data, and make the recommendation list obtained later in the embodiment of the present application more interpretable.

[0096] Step S300: extracting features from each multimedia data to obtain a corresponding second vector representation.

[0097] Specifically, the multimedia data includes the multimedia data in the target user's browsing record and the multimedia data not in the user's browsing record. Therefore, the second vector representation includes the second vector representation of the multimedia data in the target user's browsing record and the second vector representation of the multimedia data not in the user's browsing record. Among them, the second vector representation refers to the vector representation of the browsable object on the content providing platform, that is, it represents the two-way context information of the browsing object.

[0098] Optionally, convert the multimedia data into a vector matrix; extract the implicit features of the multimedia data from the vector matrix through a weight coefficient matrix; and obtain the corresponding second vector representation based on the implicit features.

[0099] Specifically, the weight coefficient matrix refers to the weight coefficient matrix learned by the recommendation model according to the relationship between words in the multimedia data. The implicit feature refers to the deep feature of the multimedia data.

[0100] In the embodiment of the present application, by converting the multimedia data into a vector matrix and then adjusting the weight coefficient matrix, the implicit features of the multimedia data can be extracted from the vector matrix according to the relationship between words. Based on these implicit features, the information of some relatively important words in the multimedia data can be obtained, and the information of these words is the two-way context information of the vector matrix. Specifically, the two-way context information of the vector matrix (i.e., the vector form of the multimedia data) refers to the context information obtained from the forward or reverse direction of the multimedia data. For example, when people read a text, they usually read from left to right (from top to bottom). However, when there is a doubt about a certain place in the article, they will infer the content of the doubtful place from the above or below of the article, and then obtain the content information of that place. The introduced representation learning model exactly imitates this behavior of humans to learn the multimedia data. Thus, for a single word in the multimedia data, the information of the single word obtained by reading in order is different from the information of the single word obtained by reading in reverse order. In order to obtain the accurate meaning contained in this word, the two-way context information of this word needs to be obtained. It should be noted that by introducing the method of extracting the features of the browsing object by the representation learning model in the embodiment of the present application, the recommendation model can not only effectively extract and capture the features of the potential concerns of the user for the browsing object, but also help to solve the technical problems of sparsity and cold start faced in the existing recommendation algorithms.

[0101] In a possible embodiment, when a certain application lacks user data (where the user data can be the user's browsing record), the deep features of the browsing object browsed by the current user can be extracted, and recommendations can be made based on the deep features of the browsing object.

[0102] Step S400: performing a fusion step for the first vector representation corresponding to each browsing record: performing feature fusion on the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fused feature.

[0103] Specifically, the multimedia data obtained in the embodiment of the present application includes multimedia data of browsing objects in the browsing history of the target user and multimedia data of browsing objects not in the browsing history of the user, so the second vector representation includes the second vector representation of the multimedia data corresponding to the browsing objects in the browsing history of the target user and the second vector representation of the multimedia data corresponding to the browsing objects not in the browsing history of the user. Therefore, there is a corresponding second vector representation for the browsing objects in the browsing history of the target user, so feature fusion of the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing history refers to fusion of two different vector representations of the same browsing object.

[0104] According to the above content, it can be seen that the first vector representation obtained by extracting features based on the browsing history of the target user can also be used to characterize the characteristics of the target user, because the browsing history is the browsing traces left by the user on the content providing platform based on his own interests and hobbies, so the user characteristics are implicit in the user's browsing history. It can be seen that the fusion feature obtained by combining the first vector representation and the second vector representation is a vector representation that can simultaneously characterize the characteristics of the target user and the characteristics of the program. Specifically, if a single first vector representation is fused with its corresponding second vector representation, the fusion feature obtained can be used to characterize the information of the multimedia data characteristics (i.e., the browsing object vector representation), and then fused with the function of the time attention mechanism, which can be summarized as formulas (1) and (2); if multiple first vector representations of a preset time period are fused and summed with their corresponding second vector representations (i.e., multiple fusion features are summed), the fusion feature obtained can be used to characterize the information of the user characteristics (i.e., the user vector representation), and then fused with the function of the time attention mechanism, which can be summarized as formula (3).

[0105] C=[AB]……(1)

[0106] Among them, A is the first vector representation obtained by extracting features from the target user's browsing history based on the graph embedding model, B is the second vector representation obtained by extracting features from multimedia data based on the representation learning model, and C is the fusion feature that characterizes the features of the multimedia data.

[0107] V=C*f(x)……(2)

[0108] Among them, f(x) is the function of the temporal attention mechanism.

[0109]

[0110] Among them, U i is the fused feature of the i-th user, m is the total number of browsing records of user i in a period, and Σc j is the total fused feature of the first vector representation of the browsing records of the target user and the second vector representation of the multimedia data corresponding to the browsing records in a period.

[0111] Optionally, the fused feature is adjusted by using the feature weights determined based on the visual attention mechanism and / or the time weights determined based on the time attention mechanism.

[0112] Specifically, the time attention mechanism refers to the mechanism in which the user's attention object changes over time. When the inventor studies the user preference phenomenon, considering that the user's preference for the object to be recommended has a significant time decay (for example, some of the user's preferences are only due to the pursuit of the object to be recommended), therefore, in order to simulate the phenomenon that the user's interest in the object to be recommended decays over time, the time attention mechanism is introduced.

[0113] Specifically, the essence of the time attention mechanism is to imitate the human visual attention mechanism, learn a weight distribution for the features of the browsing object, and then apply these weight distributions to the features of the browsing object, providing different feature influences for recommendations based on the features of the browsing object and recommendations based on the features of the user, etc., so that the task mainly focuses on some important features, ignores unimportant features, and improves the task efficiency.

[0114] Specifically, based on the visual attention mechanism, a weight distribution for the features of the browsing object can also be learned, and then these weight distributions are applied to the features of the browsing object, providing different feature influences for recommendations based on the features of the browsing object and recommendations based on the features of the user, etc., so that the task mainly focuses on some important features, ignores unimportant features, and improves the task efficiency.

[0115] Among them, the weight calculation process is to design a scoring function, calculate a score for each Attention vector, and the basis for scoring is the degree of correlation with the object that Attention focuses on (essentially a vector). The more relevant, the larger the obtained value, and the score is mapped to a value in (0, 1).

[0116] In a possible scenario, since the user's interests change over time, the recommendation of the object to be recommended should consider the time effect. For example, the user may not be interested in the content of the object that they liked in the past week. Compared with recommending the multimedia data that the user liked in the past, it is more valuable to recommend the multimedia data that the user has recently liked. Therefore, a function constructed for the time attention mechanism needs to be set, and a time parameter is added. Among them, the time parameter is a time period, so that when the recommendation model simulates the change of the user's interests, it is symmetric about the midpoint α of the time period and monotonically decreasing. Optionally, the time parameter can be set to one week, one month, one quarter, etc. For example, to fit the work system cycle, the time parameter can be set to one week, and the output of the time attention mechanism is mapped between (0, 1), symmetric about α = 4 (the fourth day) and monotonically decreasing.

[0117] Specifically, for the function of the time attention mechanism, reference can be made to formula (4):

[0118]

[0119] In another possible scenario, since the user's focus is different, the recommendation of the object to be recommended should consider the user's visual focus. For example, some users focus on the people in the multimedia data or some users focus on the scenes in the multimedia data. It is more valuable to recommend to the user based on the user's focus. Therefore, a visual attention mechanism can be set for the fusion feature, so that the recommendation model simulates the user's focus and makes recommendations based on the user's focus.

[0120] Step S500: Determine the multimedia data for recommendation from at least one piece of multimedia data based on the fusion feature.

[0121] Optionally, calculate the similarity between the fusion feature and the second vector corresponding to the multimedia data not in the target user's browsing record to determine at least one piece of similar multimedia data for recommendation; and / or, calculate the similarity between the fusion feature and the fusion features of other users to determine at least one similar user, and based on the browsing records of the similar users, determine at least one piece of similar multimedia data for recommendation that is not in the target user's browsing record.

[0122] Specifically, similarity is used to measure the similarity between two vectors. It can be understood that the calculation methods of similarity include but are not limited to: cosine similarity, Manhattan distance, Pearson correlation coefficient, Spearman rank correlation coefficient and other similarity calculation methods.

[0123] In a possible embodiment, calculate the similarity between the feature representations through cosine similarity, and screen out one or more relatively similar objects to be recommended, so that subsequent personalized recommendations can be made to the user. Among them, the cosine similarity calculation formula is as follows, referring to formula (5)

[0124]

[0125] Specifically, the object to be recommended refers to the browsing object not in the target user's browsing record. The object to be recommended can also be other similar users. Among them, the vector representation of the browsing object not in the user's browsing record corresponds to the second vector representation of the corresponding multimedia data of the browsing object, which can refer to the following formula (6).

[0126] V = B……(6)

[0127] In a possible embodiment, the fusion feature obtained based on the target user's browsing record and the second vector corresponding to the multimedia data of the browsing object not in the target user's browsing record are passed through the cosine similarity calculation formula to screen out the top several most similar browsing objects as the recommended objects (i.e., recommend based on the deep features of the browsing object).

[0128] In another possible embodiment, the fusion features corresponding to the browsing records of the target user based on a preset time period (which can be one week) are summed up. The obtained fusion feature (i.e., the fusion feature of the target user) and the fusion features of other users under the same conditions are passed through the cosine similarity calculation formula to screen out the top several most similar users as the similar users, and in the browsing records of the similar users, the browsing objects that the target user has not browsed are used as the objects to be recommended (i.e., recommend based on the deep features of the similar users).

[0129] In still another possible embodiment, the objects to be recommended obtained by the above recommendation method can be recommended to the user at the same time, that is, recommend based on the deep features of the browsing object and the similar user at the same time.

[0130] Optionally, if the number of the target user's browsing records is lower than the preset value, recommend based on the popular content.

[0131] Specifically, the existing recommendation algorithms not only rely on feature engineering for the composition of user features or program features, but also rely on the user's behavior data or the attachment information of the items to be recommended. In addition, when the user uses a certain application software less, the traditional recommendation model lacks training data, there is a risk of overfitting (sparsity problem), or the newly added items to be recommended lack corresponding historical information, and it is also difficult for the existing recommendation algorithms to accurately model the recommendation (cold start problem).

[0132] To solve the above problems, in one possible scenario, for users without browsing records or with browsing records less than a preset value, recommendations can be made based on popular content. Among them, the total number of program views is divided by the total number of unique users who have viewed the program to obtain the average click-through rate of browsable objects. The top several browsable objects with a higher average click-through rate are used as recommended objects. The following formula (7) can be referred to:

[0133]

[0134] The embodiment of the present application provides a method for recommending multimedia data. By respectively extracting features from each browsing record of the target user, a corresponding first vector representation is obtained, so that the multimedia data features based on the target user's features can be discovered through the target user's browsing records; by respectively extracting features from the multimedia data, a corresponding second vector representation is obtained, so that the recommendation method of the present application is no longer limited to the surface features of the multimedia data, but further extracts the deep features of the multimedia data corresponding to the multimedia data; the first vector representation corresponding to the browsing record and the second vector representation corresponding to the multimedia data corresponding to the browsing record are feature-fused to obtain a fused feature, so that the recommendation method of the present application can not only accurately capture the user's interests and hobbies, but also deeply discover the user's interest points, making the recommendation results have better interpretability, and at the same time no longer relying on the user's habitual data, characterizing the features of the multimedia data or the user, which helps to solve the problems of sparsity and cold start.

[0135] Based on the same inventive concept, the embodiment of the present application also provides a method for training a recommendation model, referring to Figure 2 As shown, the embodiment of the present application will detail the following steps S101 to S104:

[0136] The recommendation model includes a first feature extraction module for extracting the multimedia data features in the user's browsing records and a second feature extraction module for extracting the multimedia data features in the database; the training method includes

[0137] Step S101: Obtain a training data set; the training data set includes at least one piece of multimedia data and at least one browsing record related to the user.

[0138] Specifically, the user's browsing record refers to the browsing traces left by the user when using the application program. Among them, the user's browsing record can be used to represent the objects browsed by the user, such as movie programs, short videos, tweets, advertisements, etc. Multimedia data refers to the detailed information of all browsable objects provided on the content provider platform, including but not limited to the media asset information of the browsable objects (including: introduction, classic lines, keywords of classic scenes), comment data, types of browsable objects, etc.

[0139] Step S102: Through the first feature extraction module, based on the browsing record, obtain the predicted first vector representation.

[0140] Optionally, convert the browsing record into a node network; according to adjustable hyperparameters, determine the random walk weights, determine the traversal method in the node network based on the random walk weights, and generate the corresponding node sequence; generate the corresponding at least one predicted first vector representation based on at least one node in the node sequence.

[0141] Specifically, the node network can be a directed graph or an undirected graph without limitation in this application. Among them, the nodes in the node network are a user browsing record, that is, a node of an object browsed by a user. The node network can be generated based on the browsing records in a preset time period (such as one week, one month, one quarter, etc.), or can be generated by screening the browsing records based on the user's behavior data. The adjustable hyperparameters are the random walk hyperparameters p and q of the Node2vec graph embedding model (Node to vector model). By adjusting the random walk hyperparameters p and q in the Node2vec model, the random walk weights on the node network path are obtained, so that the Node2vec model can determine the traversal method, enabling the recommendation model to balance between the homogeneity and structure of the node network, generate the corresponding node sequence, and further obtain the context information of each node. Based on the context information of the node, the predicted first vector representation corresponding to the node (i.e., the browsing record) can be generated. The traversal methods include: BFS (Breadth First Search) and DFS (Depth First Search). The values of the random walk hyperparameters p and q can be set manually or learned through a semi-supervised method. It should be noted that p represents the probability of returning to the previous node, and q represents the probability of moving away from the previous node. The weights on the node network path can be controlled through the hyperparameters p and q, such as Figure 3As shown in the figure, if the t node is the next node that the v node wanders to, the probability of wandering to t is 1 / p (i.e., the weight on the node network path, the same below), which represents that the v node wanders back to the previous node (the t node) in the next step; if the x1 node is the next node that the v node wanders to (the t node is connected to the x1 node), the probability of wandering to x1 is 1, which represents that the v node wanders to the adjacent x1 node of the previous node (the t node) in the next step, that is, it shows BFS; if the x2 or x3 node is the next node that the v node wanders to (the t node is not connected to the x2 and x3 nodes), the probability of wandering to x2 or x3 is 1 / q, which represents that the v node wanders to farther nodes (the x2 and x3 nodes) in the next step, that is, it shows DFS; where the v node is the current node, the t node is the previous node that the v wanders to, and α represents the probability of random walk.

[0142] Specifically, the node sequence refers to obtaining the weights on the node network path by adjusting the above-mentioned random walk hyperparameters p and q, and then determining the traversal method in the node network. The node sequence of the random walk is obtained by wandering in the node network. For example, the node sequences that can be obtained from Figure 2 are: t->v->t, t->v->x1, t->v->x2, t->v->x3, etc.

[0143] In order to extract the features of browsable objects based on the user's browsing records and explore more connections between different browsable objects, the obtained node sequences will be input into the recommendation model later. By learning the features between the target nodes and other nodes in each node sequence, the context information of the target nodes is predicted, and the predicted first vector representation of each target node is obtained, that is, the predicted first vector representation that represents the context information of the browsable object.

[0144] In the embodiment of the present application, a node network can be generated by the browsing records of all different users. In a possible case, the hyperparameters of Node2vec are set to explore the features of browsable object nodes in the community based on the network homogeneity, and the predicted first vector representation of the browsable object is obtained. Where the community refers to the node tribe formed by the aggregation of nodes in the node network. For example: war programs and anime programs will form two different communities. The model can predict the next program that the user will browse in the same community after the user browses a certain program.

[0145] In another possible scenario, the hyperparameters of Node2vec are set based on the structural exploration of the network to browse the characteristics of object nodes in different communities, and a predicted first vector representation of the browsable object is obtained. For example, in two different communities, there may be two nodes with similar structures in their respective communities. The model can jump out of the current community and go to other communities to find nodes similar to the target node in the current community, and recommend nodes with similar structures or associated nodes of nodes with similar structures to the user.

[0146] Step S103: Through the second feature extraction module, based on the multimedia data, obtain a predicted second vector representation.

[0147] Optionally, convert the multimedia data into a vector matrix; where the vector matrix is composed of word vectors, segmentation vectors, and position vectors; adjust the weight coefficient matrix to perform word feature extraction on the vector matrix to obtain word-level bidirectional context information; perform sentence feature extraction on the vector matrix to obtain sentence-level bidirectional context information; generate at least one predicted second vector representation based on at least one bidirectional context information.

[0148] Specifically, the word (sentence)-level bidirectional context information refers to the context information obtained from the forward or reverse direction of the word (character) or sentence. For example, when people read a text, they usually read from left to right (from top to bottom). However, when there is a doubt about a certain part of the article, they will deliberate on the content of this doubtful part from the above or below of the article to obtain the content information of this part. The introduced representation learning model exactly imitates this behavior of humans to learn the multimedia data. Thus, it can be seen that for a single character in the multimedia data, the information of the single character obtained by reading in order is different from the information of the single character obtained by reading in reverse order. In order to obtain the accurate meaning contained in this character, it is necessary to obtain the bidirectional context information of this character.

[0149] Specifically, it is necessary to perform character segmentation on the multimedia information (optionally, word segmentation can also be performed), and convert each character (word) into a character vector representing the semantics of each character (word). Since each character (word) should express different meanings at different positions in a sentence (or in different sentences), it is necessary to distinguish different sentences and introduce the position information of the character (word). Therefore, it is necessary to set a sentence index (i.e., a segmentation vector) for each character (word), and set a position index (i.e., a position vector) for different positions of each character in different sentences. Finally, the character vector, the segmentation vector, and the position vector are summed to obtain a vector matrix, which is input into the recommendation model. Among them, the additional information is converted into a character vector, and when setting the segmentation vector for the character vector, the [CLS] symbol is inserted at the beginning of the sentence (on the one hand, it is used to aggregate the information of the entire sequence, and on the other hand, it is used to indicate that this is the beginning of the sentence), and the [SEP] symbol is inserted at the end of the sentence or at the segmentation point.

[0150] Specifically, by performing a linear transformation on the vector matrix, three character vector matrices Q (query, query vector), K (key, vector to be queried), and V (value, content vector) are obtained. In order to obtain the deep meaning of each character (word) in the multimedia data, further, it is necessary to obtain the degree of association between each character (word) and other characters (words). Therefore, the Q and K matrices of each character are multiplied pointwise, and the softmax function is used for normalization processing to obtain the relationship matrix between the current character (word) and other characters (words). Further, in order to obtain the result of Attention, the normalized relationship matrix is multiplied pointwise with V to obtain the result of Attention. Among them, in order to ensure that the dimensions of the matrices obtained before and after the pointwise multiplication of Q and K are consistent, it is necessary to transpose K and then multiply it with Q. In order to ensure the stability of the gradient before and after the pointwise multiplication, it is necessary to reduce the dimension of the matrix after the pointwise multiplication (i.e., divide by the dimension d k ).

[0151] In other words, the input based on the vector matrix can be summarized as shown in the following formula (8):

[0152]

[0153] Specifically, in order to improve the performance of the recommendation model and obtain the deep meaning of more additional information, optionally, the multi-head attention mechanism is used to set more hidden layers, and the results of multiple Attentions are concatenated through formula (9), and then multiplied by a weight coefficient matrix W O . Among them, before concatenation, it is necessary to project the Q, K, and V of different hidden layers through multiple linear transformations by formula (10) (i.e., multiply by different weight coefficient matrices W). The above weight coefficient matrices are all obtained by the recommendation model during the learning process. Refer to formulas (9) and (10) as follows:

[0154] MultiHead(Q, K, V) = Concat(head1,..., head n )W O ……(9)

[0155] head i = Attention(QW i Q , KW i K , VW i V )……(10)

[0156] Specifically, in order to obtain a deeper meaning of the additional information, it is necessary to let the recommendation model remember more context information of words (terms) and the context relationship of sentences. During the training process of the above representation learning model, there are also two pre-training tasks, namely the predicted word training task and the predicted next sentence training task.

[0157] Among them, the predicted word training task means randomly masking 15% of the word vectors during the training process. For 15% of the word vectors, the following processing is performed:

[0158] (1) With a probability of 80%, replace it with the [MASK] symbol.

[0159] (2) With a probability of 10%, replace it with a random word.

[0160] (3) With a probability of 10%, keep the word unchanged.

[0161] Through the above method, the recommendation model cannot know which word vectors are replaced, so that the recommendation model must remember the context expressions of all word vectors, thereby predicting the masked word vectors, and then learning the prediction vector representation at the word level.

[0162] Among them, in the above processing of the additional information, sentence vectors have been introduced to split the multimedia data into sentences. Therefore, the predicted next sentence training task means processing the sentence pairs in the multimedia data (that is, the sentences with context relationships are called sentence pairs) as follows:

[0163] (1) Shuffle the order of 50% of the sentence pairs.

[0164] (2) Learn the context relationship of sentences from the other 50% of the sentence pairs that are not shuffled.

[0165] Through the above method, the recommendation model is made to predict the context relationship of the sentence vectors, and then learn the prediction vector representation at the sentence level.

[0166] Through the above training method, the recommendation model learns the deep meaning in the multimedia data and generates a predicted second vector representation of the multimedia data, that is, represents the deep meaning of the browsable object.

[0167] Step S104: Update the recommendation model based on the predicted first vector and the predicted second vector.

[0168] Specifically, a loss function can be constructed according to the predicted first vector representation, the predicted second vector representation and the fusion feature, and then the value of the loss row number is calculated to update the parameters of the recommendation model.

[0169] Figure 4 Figure 1 shows a schematic diagram of the system architecture of an optional recommendation method, as Figure 4 shown in Figure 1, the system includes the user's terminal device 10, the server side of the first application, that is, Figure 4 the application server 20 shown in Figure 1 and the recommendation model training server 30. The terminal device 10 and the server side of the first application communicate through the network. Among them, an application program APP that requires a recommendation function can be installed in the terminal device 10. By opening the client of the application program, content browsing can be performed. For example, if the application program APP is a video viewing software, opening the video software can watch multimedia data (such as programs).

[0170] Among them, the recommendation model training server 30 can obtain the user's browsing records and multimedia data through the network to train the recommendation model and obtain a trained recommendation model. The trained recommendation model can be deployed in the application server 20 shown in Figure 1. Figure 4 The application server 20 shown in Figure 1 can be used to execute the recommendation method provided in the embodiments of the present application. Based on the browsing records and multimedia data of the target user, the graph embedding representation model is used to extract features from the browsing records and the representation learning model is used to extract features from the multimedia data, and a first vector representation and a second vector representation are obtained respectively. Thus, the first vector representation of the browsing records of the target user and the second vector representation of the multimedia data corresponding to the browsing records are feature-fused to obtain a fusion feature, and then the multimedia data for recommendation can be determined based on the fusion feature, and subsequent personalized recommendations are made to the user, so that the recommendation results are interpretable.

[0171] Next, in combination with Figure 4 the recommendation system shown in Figure 2, the recommendation method process in the recommendation scenario of programs (one type of multimedia data) will be described in detail. As Figure 5 shown in Figure 2, the method includes steps S201 to S207.

[0172] Step S201: Obtain business scenario data, including: user browsing and multimedia data.

[0173] Specifically, the user logs into the client of the video viewing software through the terminal device 10 and leaves a browsing record during the use of the software. The application server 20 obtains the user's browsing record on the client and the multimedia data of the program in the database.

[0174] Step S202: Use the graph embedding model to represent the user browsing record as a vector.

[0175] Specifically, the application server 20 extracts features from the user browsing record through the graph embedding model to obtain program features based on the user browsing record.

[0176] Step S203: Use the representation learning model to represent the multimedia data as a vector.

[0177] Specifically, the application server 20 extracts features from the multimedia data through the representation learning model to obtain the deep information of the program, which can enable the recommendation model to discover the user's true interests.

[0178] Step S204: Fuse features.

[0179] Specifically, the application server 20 fuses the features of two different vector representations of the same program to obtain the fused features, which can enable the vector representation of the program to have more representation methods in the feature space, that is, to obtain more semantic information.

[0180] Step S205: Incorporate the temporal attention mechanism or the visual attention mechanism.

[0181] Specifically, by incorporating the temporal attention mechanism into the fused features, the model can simulate the timeliness of the user's interests changing over time, or by incorporating the visual attention mechanism into the fused features, the model can simulate the user's focus on the program content, making the content recommended by the model to the user more interpretable.

[0182] Step S206: Feature screening.

[0183] Specifically, in one possible scenario, the fusion features corresponding to the programs in the target user's browsing record are calculated for similarity with the second vector representations corresponding to other programs not in the target user's browsing record, and the top several programs not in the target user's browsing record that are relatively similar are selected for recommendation to the target user (i.e., recommendation to the user based on the features of the programs). In another possible scenario, at least one fusion feature within a preset time period is filtered out from the target user's browsing record, and all the fusion features are summed to obtain the fusion feature based on the user, which is then calculated for similarity with the fusion features of other users under the same conditions. At least one similar user is selected, and programs that the target user has not watched yet are filtered out from the browsing records of the similar users for recommendation to the target user (i.e., recommendation based on user features).

[0184] Step S207: Determine the object to be recommended.

[0185] Optionally, the obtained programs are sent to the terminal device 10 (i.e., the user client) to generate a personalized recommendation list for recommendation to the user.

[0186] In the embodiments of the present application, at least one browsing record of the target user is respectively subjected to feature extraction to obtain the corresponding first vector representation, so as to discover the features related to the target user and the multimedia data through the browsing record of the target user; at the same time, at least one multimedia data can be respectively subjected to feature extraction to obtain the corresponding second vector representation; then, the first vector representation corresponding to the browsing record and the second vector representation corresponding to the multimedia data corresponding to the browsing record are subjected to feature fusion to obtain the fusion feature, so as to perform multimedia data recommendation based on the fusion feature, making the recommendation method of the present application not only able to accurately capture the user's interests and hobbies, but also deeply discover the correlation between the user and the multimedia data, such as the user's interest points in the multimedia data, etc.; in addition, the present application performs multimedia data recommendation based on the extracted first vector representation and second vector representation, making the recommendation result have better interpretability, and no longer relying on the user's habitual data to characterize the multimedia data or the user's features, which helps to solve the problems of sparsity and cold start.

[0187] Corresponding to the training method of the recommendation model provided by the present application, the embodiments of the present application also provide a training device for the recommendation model, as Figure 6 shown. The training device 60 for the recommendation model may include:

[0188] An acquisition module 601, configured to acquire a training data set; the training data set includes at least one multimedia data and at least one browsing record related to the user.

[0189] A training module 602, configured to obtain a predicted first vector representation based on browsing records through a first feature extraction module, and obtain a predicted second vector representation based on multimedia data through a second feature extraction module. Update the recommendation model based on the predicted first vector and the predicted second vector.

[0190] Corresponding to the recommendation method provided in this application, an embodiment of this application further provides a recommendation device, as Figure 7 shown. The recommendation device 70 may include:

[0191] A data acquisition module 701, configured to acquire at least one piece of multimedia data and at least one browsing record related to a target user.

[0192] A first feature extraction module 702, configured to perform feature extraction on each browsing record respectively to obtain a corresponding first vector representation;

[0193] A second feature extraction module 703, configured to perform feature extraction on each piece of multimedia data respectively to obtain a corresponding second vector representation;

[0194] A fusion module 704, configured to perform a fusion step for the first vector representation corresponding to each browsing record: perform feature fusion on the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fusion feature.

[0195] A determination module 705, configured to determine the multimedia data for recommendation from at least one piece of multimedia data based on the fusion feature.

[0196] Optionally, the fusion module 704 is further configured to adjust the fusion feature by using a feature weight determined based on a visual attention mechanism and / or a time weight determined based on a time attention mechanism.

[0197] Optionally, the first feature extraction module 702 is further configured to generate a node sequence for each node corresponding to a browsing record; the distance information and similarity information between the nodes in the node sequence respectively have corresponding random walk weights; based on the node sequence, generate a corresponding first vector representation.

[0198] Optionally, the second feature extraction module 703 is further configured to generate a vector matrix for each piece of multimedia data; extract the implicit features of the multimedia data from the vector matrix; based on the implicit features, generate a corresponding second vector representation.

[0199] Optionally, the determining module 705 is further configured to calculate the similarity between the fused feature and a second vector corresponding to multimedia data not in the target user's browsing record, to determine at least one similar multimedia data for recommendation, and / or calculate the similarity between the fused feature and the fused features of other users, to determine at least one similar user, and based on the browsing record of the similar user, determine at least one similar multimedia data for recommendation not in the target user's browsing record.

[0200] An embodiment of the present application provides a method for recommending multimedia data. The first feature extraction module 702 respectively extracts features from each browsing record of the target user to obtain corresponding first vector representations, so that program features based on the target user's features can be discovered through the target user's browsing record; the second feature extraction module 703 can also respectively extract features from multimedia data to obtain corresponding second vector representations, so that the recommendation method of the present application is no longer limited to the surface features of the program, but further extracts the deep features of the multimedia data corresponding to the program; the fusion module 704 can also perform feature fusion on the first vector representation corresponding to the browsing record and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fused feature, so that the recommendation method of the present application can not only accurately capture the user's interests and hobbies, but also deeply discover the user's interest points, making the recommendation result have better interpretability, and at the same time no longer relying on the user's habitual data, which helps to solve the problems of sparsity and cold start.

[0201] In addition, in the embodiment of the present application, the determining module 705 can also perform recommendation based on the deep features of the recommended content extracted by the above method, making the recommendation result more personalized and interpretable.

[0202] In an alternative embodiment, a computer device is provided, such as Figure 8 shown Figure 8 The electronic device 4000 shown in the figure includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiment of the present application.

[0203] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0204] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is used to represent it herein, but it does not mean that there is only one bus or one type of bus.

[0205] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0206] The memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0207] Among them, the electronic device includes but is not limited to: mobile phones, tablet computers, PDAs (Personal Digital Assistants), POS (Point of Sales), in-vehicle computers, servers, and any other electronic devices.

[0208] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.

[0209] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.

[0210] Terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.

[0211] It should be understood that although the flowchart of the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless clearly stated in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0212] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.

Claims

1. A method for recommending multimedia data, characterized in that, Comprising: Obtaining at least one piece of multimedia data and at least one browsing record related to a target user; Generating a corresponding node network according to each browsing record, where each node in the node network represents a browsing record; In the node network, adjusting the random walk weights according to the distances or structures between different nodes to obtain a node sequence. The distance information and similarity information between the nodes in the node sequence respectively have corresponding random walk weights. Based on the node sequence, a corresponding first vector representation is generated through a graph embedding model; For each piece of multimedia data, converting the multimedia data into a vector matrix; extracting the implicit features of the multimedia data from the vector matrix through a representation learning model, and based on the implicit features, obtaining a corresponding second vector representation; Performing a fusion step on the first vector representation corresponding to each browsing record: fusing the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fused feature; Determining the multimedia data for recommendation from the at least one piece of multimedia data based on the fused feature.

2. The method according to claim 1, characterized in that, The fusion step further includes: Adjusting the fused feature by using the feature weights determined based on the visual attention mechanism and / or the time weights determined based on the time attention mechanism.

3. The method according to claim 1, characterized in that, The determining the multimedia data for recommendation from the at least one piece of multimedia data based on the fused feature includes: Calculating the similarity between the fused feature and the second vector corresponding to the multimedia data not in the browsing record of the target user to determine at least one similar multimedia data for recommendation; And / or calculating the similarity between the fused feature and the fused features of other users to determine at least one similar user, and based on the browsing records of the similar user, determining at least one similar multimedia data not in the browsing record of the target user.

4. A training method for a recommendation model, characterized in that, The recommendation model includes a graph embedding model for extracting the features of user browsing records and a representation learning model for extracting the features of multimedia data; the training method includes: Obtaining a training data set; the training data set includes at least one piece of multimedia data and at least one browsing record related to a user; Generating a corresponding node network according to each browsing record, adjusting the random walk weights according to the distances or structures between different nodes to obtain a node sequence, and based on the node sequence, obtaining a predicted first vector representation through the graph embedding model; For each piece of multimedia data, converting the multimedia data into a vector matrix, extracting the implicit features of the multimedia data from the vector matrix through the representation learning model, and based on the implicit features, obtaining a predicted second vector representation; Updating the recommendation model based on the predicted first vector and the predicted second vector; Wherein, the trained recommendation model is applied to the recommendation method for multimedia data described in claims 1-3.

5. A multimedia data recommendation device, characterized in that, Comprising: A data acquisition module for obtaining at least one piece of multimedia data and at least one browsing record related to a target user; The first feature extraction module is used to generate a corresponding node network according to each browsing record, and each node in the node network represents a browsing record; In the node network, according to the distance or structure between different nodes, adjust the random walk weights to obtain a node sequence. The distance information and similarity information between the nodes in the node sequence respectively have corresponding random walk weights. Based on the node sequence, through a graph embedding model, generate a corresponding first vector representation; The second feature extraction module is used to convert each multimedia data into a vector matrix for each multimedia data; through a representation learning model, extract the implicit features of the multimedia data from the vector matrix, and based on the implicit features, obtain a corresponding second vector representation; The fusion module is used to perform a fusion step for the first vector representation corresponding to each browsing record: fuse the first vector representation and the second vector representation corresponding to the multimedia data corresponding to the browsing record to obtain a fusion feature; The determination module is used to determine the multimedia data for recommendation from the at least one multimedia data based on the fusion feature.

6. A training device for a recommendation model, characterized in that, The recommendation model includes a graph embedding model for extracting the features of multimedia data in the user's browsing records and a representation learning model for extracting the features of multimedia data in the database; The training device includes: An acquisition module for acquiring a training data set; the training data set includes at least one multimedia data and at least one browsing record related to the user; A training module for generating a corresponding node network according to each browsing record, adjusting the random walk weights according to the distance or structure between different nodes to obtain a node sequence, and based on the node sequence, through the graph embedding model, obtaining a predicted first vector representation; for each multimedia data, converting the multimedia data into a vector matrix, and through the representation learning model, extracting the implicit features of the multimedia data from the vector matrix, and based on the implicit features, obtaining a predicted second vector representation; updating the recommendation model based on the predicted first vector representation and the predicted second vector representation; Among them, the trained recommendation model is applied to the method for recommending the multimedia data described in claims 1-3.

7. A computer device, characterized in that, Including: One or more memories; A processor and a computer program stored on the memory, The processor executes the computer program to implement the steps of the method according to any one of claims 1-4.

8. A computer-readable storage medium, characterized in that, Including: A computer program is stored thereon, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-4.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Video recommendation method, server and readable storage medium

    CN112822526A