Application method and system of hybrid recommendation model based on knowledge graph in movie recommendation system

By applying a hybrid recommendation model based on knowledge graph in the movie recommendation system, combining collaborative filtering and content recommendation, the problem of lack of transparent recommendation process and personalized recommendation problems in the existing system is solved, and higher recommendation accuracy and user experience are achieved.

CN120216766AActive Publication Date: 2025-06-27HUBEI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510285689.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing movie recommendation system lacks a transparent recommendation process, affects the user experience, and has problems with sparse data, cold start and personalized recommendation.

Method used

Using a hybrid recommendation model based on knowledge graphs, we use the knowledge graph of the film field, combined with collaborative filtering and content recommendation, and use the improved TransH model to perform data fitting and least squares method to reduce calculation errors to achieve personalized recommendations.

Benefits of technology

It improves the accuracy and recall of the recommendation system, enhances the interpretability of the recommendation results, solves the cold start problem, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216766A_ABST
    Figure CN120216766A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of knowledge maps, discloses an application method of a hybrid recommendation model based on a knowledge map in a movie recommendation system, and provides the hybrid recommendation model based on the knowledge map, which shows better accuracy and recall rate than a traditional recommendation model to a certain extent compared with the traditional recommendation model. The effectiveness of the model is verified through experiments, and finally the movie recommendation system based on the knowledge graph mixed recommendation model is realized. The invention relates to design and implementation of a movie recommendation system based on a knowledge graph. And the improved mixed recommendation model is applied to the system. Meanwhile, system function requirements are analyzed, system functions and architecture are designed, E-R graph object analysis and database table design are included, functions of all modules are achieved, finally, the functions of all the modules are tested, and it is guaranteed that the system can stably run while the function requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge graphs, and particularly relates to a method and system for applying a hybrid recommendation model based on a knowledge graph in a movie recommendation system. Background Art

[0002] In recent years, with the rapid development of Internet technology, the massive amount of fragmented information generated every day has caused problems of information overload and data explosion, and it takes a long time for users to retrieve the required content from the massive information. As a key technology to solve this problem, a recommendation system provides personalized suggestions for users based on their historical behaviors and content features through filtering and intelligent decision-making. The recommendation system was initially proposed by Resnick et al. and has been quickly applied to fields such as movies, music, e-commerce, and news. For example, Tencent Video, QQ Music, Taobao, Douyin, Weibo, etc. all adopt a variety of recommendation algorithms.

[0003] Traditional collaborative filtering mainly calculates similarity using user ratings and click data, and there are problems of data sparsity and cold start; while content-based recommendation recommends through item attributes and can only play a role in the new user scenario. To make up for their respective deficiencies, a hybrid recommendation method has emerged, combining the advantages of collaborative filtering and content-based recommendation.

[0004] As a technology proposed by Google, a knowledge graph constructs a knowledge network using entities, attributes, and multi-sided semantic relationships and has been widely applied in intelligent question answering, search, and recommendation systems. A knowledge graph can not only enhance the data association of a recommendation system but also improve the interpretability of recommendation results. However, existing movie recommendation systems often only display results and lack a transparent recommendation process, affecting the user experience. Therefore, how to use a knowledge graph and a recommendation algorithm in a deep integration to provide users with personalized and interpretable recommendation services has become a current research hotspot.

[0005] Certain achievements have been made at home and abroad in aspects such as knowledge graph construction and recommendation algorithm improvement, but problems such as data sparsity, cold start, and personalized recommendation still exist. Generally speaking, using the rich semantic relationships of a knowledge graph to assist in recommendation and combining a hybrid recommendation strategy provide new ideas and methods for solving the problem of information overload, improving retrieval efficiency, and enhancing the user experience. Summary of the Invention

[0006] In view of the problems existing in the prior art, the present invention provides a method for applying a hybrid recommendation model based on a knowledge graph in a movie recommendation system.

[0007] The present invention is implemented as follows. A method for applying a hybrid recommendation model based on a knowledge graph in a movie recommendation system includes:

[0008] Step 1, constructing a knowledge graph in the movie field;

[0009] Step 2: Improved knowledge graph-based hybrid recommendation model in the movie field;

[0010] Step 3, implementation of movie recommendation system.

[0011] Furthermore, the film field knowledge graph is constructed as follows:

[0012] The selected ml-latest-small dataset contains movie names, movie types, user ratings, and external website link information; according to the movie data information in ml-latest-small;

[0013] First, we designed the model layer of the knowledge graph in the film field. Based on the MovieID of each movie in the dataset, we crawled the corresponding semi-structured information of movie type, director, starring actor, and screenwriter on the IMDB website through web crawler technology and stored it in the local MongoDB database.

[0014] The crawled data is further processed to filter out very little invalid and duplicate information. The few non-existent information is filled with null or manually filled. The processed data is stored locally in csv format to obtain a structured movie dataset. Entities and relationships are extracted from the organized structured movie dataset and organized into triples. The load csv command is used to import the data into Neo4j to complete knowledge storage, thus completing the construction of the knowledge graph.

[0015] Furthermore, the improved knowledge graph-based hybrid recommendation model in the movie field:

[0016] Combined with the constructed film field knowledge graph, which contains multiple feature items such as movie actors, directors, and genres, these feature items are used to comprehensively calculate the similarity of movie content features;

[0017] The semantic similarity of movies based on the knowledge graph is calculated and fused with the results of content-based recommendation technology and collaborative filtering recommendation technology to calculate the scoring results of a single recommendation model based on the knowledge graph;

[0018] Then, the two single model recommendation results based on the knowledge graph are fitted to the data by adjusting the weights and eliminating the relative errors caused by the fusion of multiple models as much as possible;

[0019] Finally, the least squares method is used to reduce the calculation error of the hybrid recommendation model;

[0020] Improved Knowledge Representation Learning:

[0021] For the TransE model, for a fact triple (h, r, t), the relationship r is regarded as the translation distance between the head entity (head) and the tail entity (tail) after mapping in the vector space. Then, the scoring function shown in the formula is used to continuously adjust the quality of the model, making h + r in the fact triple as close as possible to t, that is, h + r ≈ t;

[0022] Compared with TransE, for each type of relationship r in the triple, the TransH model gives a hyperplane Wr, defines the relationship vector dr on the hyperplane Wr, and maps the original head entity h and tail entity t onto the hyperplane, denoted as hr and tr; it is required that the correct triple satisfies the following formula:

[0023] h r +d r =t r

[0024] In TransH, if there is a triple (h1, r, t) and (h2, r, t) for both the h1 and h2 vectors; through the mapping of the hyperplane of relationship r in TranH, there is:

[0025] h 1r +d r =t r

[0026] h 2r +d r =t r

[0027] That is to say, the mappings of h1 and h2 on the hyperplane are the same or approximate; however, h1 and h2 themselves can be not close, that is, distinguishable:

[0028] Compared with TransE when updating the vector representation, for each relationship, one more mapping vector Wr is learned, and the projected head entity and tail entity can be expressed as:

[0029]

[0030] And the learned scoring function and loss function are different, and the scoring function and calculation function are as shown in the formula:

[0031]

[0032] L=Σ (h,r,t)∈△ Σ (h′,r′,t′)∈△′ [γ + f r (h, t) - f r (h′, t′)] + 。

[0033] Although the TransH model solves the problems of one-to-many and many-to-many complex relationships, as can be seen from the Figure 10 shown process, the model may introduce the problem of incorrect labels during the negative sampling process. The negative sampling strategy of the original TransH model is as follows: First, a certain probability is set to replace the head entity or the tail entity. Its purpose is to make the head entity more likely to be replaced when the relationship is one-to-many, and to make the tail entity more likely to be replaced when the relationship is many-to-one. This replacement strategy can improve the accuracy of negative sample generation compared to the random replacement method of TransE.

[0034] Furthermore, the implementation of the movie recommendation system:

[0035] By analyzing the requirements of the recommendation system, designing the system architecture and functions, including database design and E-R diagram design, and combining the sorted requirements with the proposed knowledge graph-based hybrid recommendation model, a movie recommendation system including movie management, graph display, movie search, movie relationship query, user evaluation, and movie recommendation functions is designed and implemented; and a visual recommendation display is realized.

[0036] Another object of the present invention is to provide an application system of a knowledge graph-based hybrid recommendation model in a movie recommendation system, including:

[0037] A construction module for constructing a knowledge graph in the movie field;

[0038] An improvement module for an improved knowledge graph-based hybrid recommendation model in the movie field;

[0039] An implementation module for implementing a movie recommendation system.

[0040] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for applying the knowledge graph-based hybrid recommendation model in a movie recommendation system.

[0041] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method for applying the knowledge graph-based hybrid recommendation model in a movie recommendation system.

[0042] Another object of the present invention is to provide an information data processing terminal for implementing the application system of the knowledge graph-based hybrid recommendation model in a movie recommendation system.

[0043] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0044] First, the present invention proposes a knowledge graph-based hybrid recommendation model. Compared with traditional recommendation models, it shows better accuracy and recall rates to a certain extent, and the effectiveness of the model is verified through experiments. Finally, a movie recommendation system based on the knowledge graph-based hybrid recommendation model is realized. The main research results are summarized as follows:

[0045] (1) Construct a movie domain knowledge graph. The present invention uses the movie link information in the ml-latest-small dataset to construct the knowledge graph. Using Python crawler technology, it crawls movie data on the IMDB website, processes the crawled data and stores it in a MySQL database, converts the data into csv format and stores it in a MongoDB database, designs the schema layer of the movie knowledge graph, determines the ontology classes, converts the data into the structured data required by the knowledge graph, and uses the load csv instruction to import the data into a Neo4j graph database to facilitate the persistent storage and visual display of the structured movie knowledge.

[0046] (2) Propose a knowledge graph-based hybrid recommendation model. Considering the deficiencies of traditional recommendation models, it combines the semantic relationships contained in the knowledge graph as supplementary information to assist in recommendation, and optimizes the negative sampling process of the TransH model during the vectorization of entity relationships. The model first calculates the semantic similarity of movies through the constructed movie domain knowledge graph, fuses the movie similarity based on the knowledge graph with the user rating similarity obtained by collaborative filtering recommendation and the movie content similarity based on content recommendation, determines the fusion ratio through experiments, fits the two sets of recommended results, sorts the final fitting results according to the rating size, and finally performs TOP-N recommendation. The effectiveness of the model is verified through comparative experiments, and it has higher precision, recall rate and F1 value compared with traditional recommendation algorithms.

[0047] (3) Design and implementation of a movie recommendation system based on the knowledge graph. And apply the improved hybrid recommendation model to the system. At the same time, analyze the system functional requirements, design the system functions and architecture, including E-R diagram object analysis and database table design, implement the functions of each module, and finally test the functions of each module to ensure that the system can run stably while meeting the functional requirements.

[0048] Second, when the existing movie recommendation system makes recommendations, for user recommendations, only the recommendation results are displayed, and the specific recommendation process and method are not shown. For users, it is not interpretable, resulting in a poor personalized recommendation experience for users who are more interested in the recommendation process. Using a knowledge graph to display can more intuitively show the relevance between results; an improved TransH model is proposed. As can be seen from the experimental results, it can effectively improve the accuracy of recommendation results; using a public dataset for graph construction can, to a certain extent, solve the cold start problem. Brief Description of the Drawings

[0049] Figure 1 is a flowchart of the application method of the hybrid recommendation model based on the knowledge graph in the movie recommendation system provided by an embodiment of the present invention.

[0050] Figure 2 is a block diagram of the system structure of the application of the hybrid recommendation model based on the knowledge graph in the movie recommendation system provided by an embodiment of the present invention.

[0051] Figure 3 is a flowchart of content-based recommendation provided by an embodiment of the present invention.

[0052] Figure 4 is a flowchart of item-based recommendation provided by an embodiment of the present invention.

[0053] Figure 5 is a flowchart of user-based recommendation provided by an embodiment of the present invention.

[0054] Figure 6 is a schematic diagram of the knowledge graph provided by an embodiment of the present invention.

[0055] Figure 7 is a technical architecture diagram of the knowledge graph provided by an embodiment of the present invention.

[0056] Figure 8 is a classification diagram of knowledge extraction tasks provided by an embodiment of the present invention.

[0057] Figure 9 is a spatial diagram of entity relationships provided by an embodiment of the present invention.

[0058] Figure 10 is a flowchart of TransE provided by an embodiment of the present invention.

[0059] Figure 11 is a diagram of movie semantic elements provided by an embodiment of the present invention.

[0060] Figure 12 is a flowchart of knowledge graph construction provided by an embodiment of the present invention.

[0061] Figure 13It is the information graph crawled from the movie Mission: Impossible - Fallout provided by the embodiments of the present invention.

[0062] Figure 14 It is the information graph of the director information crawled and stored in mongodb provided by the embodiments of the present invention.

[0063] Figure 15 It is the graph of importing data into the neo4j database provided by the embodiments of the present invention.

[0064] Figure 16 It is the command graph of storing data into Neo4j provided by the embodiments of the present invention.

[0065] Figure 17 It is the partial movie knowledge graph data graph stored in Neo4j provided by the embodiments of the present invention.

[0066] Figure 18 It is the knowledge graph data graph of the movie Mission: Impossible - Fallout provided by the embodiments of the present invention.

[0067] Figure 19 It is the comparison graph of TransE and TransH models provided by the embodiments of the present invention.

[0068] Figure 20 It is the schematic diagram of vector mapping to the hyperplane provided by the embodiments of the present invention.

[0069] Figure 21 It is the improved TransH flow chart provided by the embodiments of the present invention.

[0070] Figure 22 It is the framework graph of the hybrid recommendation model based on the knowledge graph provided by the embodiments of the present invention.

[0071] Figure 23 It is the data graph of partial content of the data set provided by the embodiments of the present invention.

[0072] Figure 24 It is the Hit@10 graph under different embedding dimensions provided by the embodiments of the present invention.

[0073] Figure 25 It is the Meanrank graph under different embedding dimensions provided by the embodiments of the present invention.

[0074] Figure 26 It is the graph of precision rate, recall rate and F1 value corresponding to different fusion ratios provided by the embodiments of the present invention.

[0075] Figure 27 It is the graph of precision rate, recall rate and F1 value corresponding to different fusion ratios provided by the embodiments of the present invention.

[0076] Figure 28 It is the precision rate graph corresponding to different K values when α1 = 0.8 and α2 = 0.5 provided by the embodiments of the present invention.

[0077] Figure 29 It is the recall rate graph corresponding to different K values when α1 = 0.8 and α2 = 0.5 provided by the embodiments of the present invention.

[0078] Figure 30 It is the F1 value graph corresponding to different K values when α1 = 0.8 and α2 = 0.5 provided by the embodiments of the present invention.

[0079] Figure 31 It is the MAE graph corresponding to different K values when α1 = 0.8 and α2 = 0.5 provided by the embodiments of the present invention.

[0080] Figure 32 It is the user use case diagram provided by the embodiments of the present invention.

[0081] Figure 33 It is the administrator use case diagram provided by the embodiments of the present invention.

[0082] Figure 34 It is the system architecture diagram provided by the embodiments of the present invention.

[0083] Figure 35 It is the system function diagram provided by the embodiments of the present invention.

[0084] Figure 36 It is the E-R diagram provided by the embodiments of the present invention.

[0085] Figure 37 It is the user login page diagram provided by the embodiments of the present invention.

[0086] Figure 38 It is the home page diagram provided by the embodiments of the present invention.

[0087] Figure 39 It is the historical movie viewing record page diagram provided by the embodiments of the present invention.

[0088] Figure 40 It is the user historical movie viewing spectrum viewing diagram provided by the embodiments of the present invention.

[0089] Figure 41 It is the movie evaluation page diagram provided by the embodiments of the present invention.

[0090] Figure 42 It is the movie search interface diagram provided by the embodiments of the present invention.

[0091] Figure 43 It is the movie relationship query interface diagram provided by the embodiments of the present invention.

[0092] Figure 44 It is a movie recommendation interface diagram provided by an embodiment of the present invention.

[0093] Figure 45 It is an interface for viewing the relationship of the movie graph provided by an embodiment of the present invention

[0094] Figure 46 It is a user information addition diagram provided by an embodiment of the present invention.

[0095] Figure 47 It is a user information maintenance interface diagram provided by an embodiment of the present invention.

[0096] Figure 48 It is a movie information maintenance interface diagram provided by an embodiment of the present invention.

[0097] Figure 49 It is a user login page diagram provided by an embodiment of the present invention. Detailed implementation manners

[0098] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0099] As Figure 1 shown, a method for applying a knowledge graph-based hybrid recommendation model in a movie recommendation system provided by an embodiment of the present invention includes the following steps:

[0100] S101, constructing a movie domain knowledge graph;

[0101] S102, an improved movie domain hybrid recommendation model based on a knowledge graph;

[0102] S103, implementing a movie recommendation system.

[0103] As Figure 2 shown, a system for applying a knowledge graph-based hybrid recommendation model in a movie recommendation system provided by an embodiment of the present invention includes:

[0104] A construction module for constructing a movie domain knowledge graph;

[0105] An improvement module for an improved movie domain hybrid recommendation model based on a knowledge graph;

[0106] An implementation module for implementing a movie recommendation system.

[0107] The specific implementation of the present invention:

[0108] 1. For movie recommendation, the main implementation steps can be divided into the following four steps:

[0109] 1) Extract the features and attributes of movies in the dataset, and represent the extracted movie feature information or attribute labels in the form of vectors;

[0110] 2) Perform representation learning on the feature vectors of each movie to obtain a similarity calculation model that can represent user interests;

[0111] 3) According to the similarity calculation model in the second step, calculate the similarity between the movies to be recommended and the movies that the user is interested in, and obtain a recommendation list;

[0112] 4) Recommend the top N movies with the highest similarity to the user in the way of TOP-N recommendation; its main flowchart is as Figure 3 shown:

[0113] Table 2.1 User Movie-Watching Table

[0114]

[0115] Taking Table 2.1 as an example, User 1 has watched movies A, B, and C; User 2 has watched movie B; User 3 has watched movies A and C; According to the idea of item-based collaborative filtering, it is considered that User 1 and User 3 have the same interest preferences, so movie B that User 3 has not watched will be recommended to User 3.

[0116] This algorithm first constructs a user-movie rating matrix by calculating movie similarities based on the user's historical rating data, and recommends after calculating the predicted ratings of the movies to be recommended for other similar movies by the target user. The algorithm implementation steps can be divided into the following steps:

[0117] Construct a user-movie rating matrix based on the user's historical rating data;

[0118] Calculate the similarity between movies in the matrix (the specific calculation method will be introduced later) to obtain an n×n order movie similarity matrix;

[0119] Find the top k movies to be recommended through the movie similarity matrix to form a set;

[0120] Sort the set according to the similarity size to obtain the final recommendation list. The algorithm flowchart is as Figure 4 shown:

[0121] The premise idea of this algorithm is that if two users have similar ratings for the same item, then it is considered that the two users have similar interest preferences. Under this premise condition, according to the historical rating data of similar users, the rating prediction formula is used to predict the rating of the target user for this type of item. Taking the example of users watching movies, as shown in Table 2.2.

[0122] Table 2.2 User Movie Viewing Table

[0123]

[0124]

[0125] The movies watched by User 1 are A and B; the movies watched by User 2 are B and D; the movies watched by User 3 are A, B, and C; User 1 and User 3 both watched Movie A and Movie B; however, User 1 did not watch C. According to the algorithm idea of user-based collaborative filtering, User 1 and User 3 have extremely similar interests, so Movie C is recommended for User 1 to watch.

[0126] For movie recommendations, user-based collaborative filtering recommendation first calculates the set of other users similar to the user, then calculates the predicted ratings of the movies to be recommended through the historical rating data of these similar users on the movies using the rating prediction formula, and finally sorts the predicted ratings from high to low to obtain the list of movies to be recommended. Taking user-based collaborative filtering as an example, its basic implementation steps are as follows:

[0127] 1) Establish a rating matrix of users for movies;

[0128] 2) Establish a similarity calculation model between movies based on the rating information of users for movies;

[0129] 3) Calculate the similarity between every two movies according to the obtained calculation model, then calculate the calculated ratings of all unwatched movies and sort them to obtain the recommendation list;

[0130] 4) Recommend the movies ranked in the top N in terms of predicted ratings in the recommendation list to the target user. Its flowchart is as Figure 5 shown:

[0131] The essence of collaborative filtering recommendation is actually to borrow "collective wisdom". By constructing a similarity model between user sets or item sets, and then using the user's historical rating data to calculate the user's rating for unviewed movies for neighbor-based recommendation. Here, the neighbors refer to user sets or item sets, because the premise of the collaborative filtering algorithm is the assumption that similar users may be interested in the same items, or similar items may be liked by the same users. The collaborative filtering recommendation algorithm can combine collective wisdom and effectively utilize explicit data such as user ratings. Through ratings, information that is not completely similar in content can be recommended, which can effectively mine the potential interests of users compared to content-based recommendation. At the same time, this algorithm also has the following disadvantages: 1) The problem of data sparsity. When there is less rating data of users for movies in the recommendation system or dataset, it will lead to the problem of sparse rating matrix data, and the recommended results are not representative; 2) The cold start problem. New movies have no user ratings, and new users have no historical data, so the similarity cannot be calculated and no recommendation can be made; 3) The scalability problem. When the amount of data in the dataset continues to increase, the recommendation effect and accuracy will become worse and worse.

[0132] As Figure 6 shown, Movie 1 and science fiction and their relationship constitute the triple (Movie 1, genre, science fiction).

[0133] The technical architecture of the knowledge graph is as Figure 7 shown. This section mainly introduces knowledge extraction, knowledge representation, and knowledge storage.

[0134] The task classification technology of knowledge extraction is as Figure 8 shown:

[0135] (1) Entity extraction

[0136] Entity extraction can automatically identify the "tagged" entities from text data. Entity extraction is extremely important for constructing a knowledge graph, and the quality of the extracted entities directly affects the accuracy of the subsequent knowledge base and recommendation system. The main methods of entity extraction are as follows: 1) Rule- and dictionary-based methods. 2) Deep learning-based methods 3) Machine learning-based methods 4) Extraction based on web sites. The method mainly adopted in this invention is extraction through vertical sites.

[0137] (2) Relationship extraction

[0138] Relationship extraction generally extracts the semantic relationships between entities from text information. The main methods of relationship extraction are the following three methods: 1) Template-based relationship extraction; 2) Supervised learning-based relationship extraction; 3) Semi-supervised learning-based relationship extraction.

[0139] (3) Attribute extraction

[0140] Attribute extraction is to extract entity attribute information from data such as text. In the movie field, the director, lead actors, and movie genres are all attributes. For example, for the movie "Movie 1", the attribute value of the director is "Director 1", the attribute values of the actors are "Actor 3" and "Actor 1", and the attribute values of the movie genres are "Science Fiction", "Romance", and "Comedy". The data for constructing the knowledge graph in the movie field in the present invention comes from the movie website IMDB and belongs to semi-structured data.

[0141] 2.2.2 Knowledge Representation

[0142] Knowledge representation is actually to transform the things in our real world into a language that machines can understand, and through computer calculations, simulate human cognition and reasoning about these things, and solve complex computational tasks in the field of artificial intelligence in this way.

[0143] The TransE model was proposed by Antoine et al. in 2013. In this model, the entities and relationships in the triple (h, r, t) are represented in vector form in the same vector space, and then the relationship vector r is regarded as the translation between the head entity vector h and the tail entity vector t, that is, let h + r ≈ t. For example, Actor 3 + Lead Actor ≈ Movie 1. This model is also called the translation model. As Figure 9 shown.

[0144] The scoring function of the model is:

[0145]

[0146] If the fact (h, r, t) exists, the higher the score; on the contrary, if the fact does not exist or the rationality is low, the score is lower. The model has the advantages of fast training and easy implementation. At the same time, it also has good performance even when facing a large-scale sparse knowledge base. However, its effect is not good when facing complex (one-to-many, many-to-one, many-to-many) relationships. For example, given two facts "Actor 3 - Lead Actor - King of Comedy" and "Actor 3 - Lead Actor - Movie 1". Then according to the idea of the model, it will be: Actor 3 + Lead Actor ≈ King of Comedy, Actor 3 + Lead Actor ≈ Movie 1, which will make King of Comedy ≈ Movie 1, but these two movies are different entities and should be represented by different vectors. The main workflow of the TransE model is as Figure 10 shown.

[0147] Its loss function is shown in Equation 2-6, and subsequent researchers have improved TransE.

[0148]

[0149] The quality of the schema layer design is directly related to the quality of the constructed knowledge graph. How to add "high-quality" knowledge to the ontology library is the main problem to be solved. The semantic elements of the movie domain knowledge graph constructed by the present invention are as Figure 11 shown:

[0150] The schema layer design is mainly divided into the following steps:

[0151] 1) Determine the abstract classes; abstract the entities in the movie knowledge graph into classes. For example, abstract the movie name into the movie class (MOVIE), abstract the movie actors, directors, and screenwriters into the person class (PERSON), abstract the movie type into the movie type class (GENRES), and abstract the movie id into "MovieID" to obtain a total of 4 entity classes.

[0152] Determine the domain and range of the classes in the ontology; the ontology classes, class attributes, domain, and range defined by the present invention are shown in Tables 3.1 and 3.2 respectively:

[0153] Table 3.1 Ontology class design

[0154]

[0155] Table 3.2 Attributes and constraints of classes in the ontology

[0156]

[0157] 3) Instantiate the ontology: After the schema layer design is completed, the ontology needs to be instantiated, and the schema layer is filled at the data level by obtaining the movie domain knowledge graph data.

[0158] The main purpose of constructing the movie knowledge graph by the present invention is for the personalized recommendation in the following text. The construction of the movie knowledge graph is as Figure 12 shown.

[0159] 3.2.1 Movie data crawling

[0160] The acquisition of movie data is carried out by using a Python web crawler, and the request library, Beautiful Soup library and regular expressions in Python are called for acquisition. The request library is mainly used to send requests to URL links and receive the response results of the requests. The Beautiful Soup library is used to parse the requested HTML page and parse out the content required for constructing the graph. Since the URL link of each movie in the ml-latest-small dataset is the IMDB website link "https: / / www.imdb.com / " concatenated with " / ", "title / tt" and the corresponding imdbId in sequence, the imdbIds in the links.csv file can be read in sequence to crawl the relevant information of the corresponding movies. Taking the movie "Movie 3" corresponding to MovieId 189333 as an example, the relevant information of this movie can be accessed through the concatenated link "https: / / www.imdb.com / title / tt4912910 / ", and relevant information such as the movie name, movie director and movie actors are selected for crawling, as Figure 12 shown

[0161] The relevant information such as the movie name, movie director, movie actors and movie type crawled are saved into the locally classified MongoDB database respectively. Taking the movie director as an example, as Figure 13 shown, the database for saving directors is divided into two columns, which save the movie name and the corresponding movie director respectively, facilitating subsequent structured export. The processing methods for movie type and movie actors are the same as that of the movie director, and will not be elaborated here. So far, the relevant information of the required movies has been crawled. Figure 14 The crawled director information is stored in mongodb.

[0162] This invention mainly uses load csv to import data and uses Cypher language to query data. The specific methods are as follows:

[0163] The four groups of files of movie actors, directors, movie types and MovieID are placed in the import folder of the neo4j installation directory in the form of.csv files.

[0164] 2) In the bin directory, open the command line window and execute the "neo4j.bat console" command to open the neo4j database service.

[0165] 3) Enter the visualization interface of database management and input the command "LOAD csv WITH HEADERS FROM ' / filename.csv' AS csv" in sequence, and the data can be imported into Neo4j for storage. This invention uses a python program to import four groups of data simultaneously at one time. The program screenshot is as Figure 15 , 16 shown.

[0166] After the data storage is completed, the movie domain knowledge graph is constructed, as Figure 17 shown. You can input Cypher statements in the visualization management interface to search for knowledge. Taking the movie "Movie 3" as an example, by inputting the command "Match(m:Movie{title:'Movie 3'})--(n)return m,n" in the Neo4j command bar, all the corresponding relationships and tail entities with this movie as the head entity can be queried, as Figure 18 shown. The entity types and entity quantities of the movie domain knowledge graph constructed in this invention are shown in Table 3.4. The total number of entities in the relationship types and relationship quantities is 35,172, and the total number of relationships is 72,508. The specific quantities of each entity and relationship can be viewed in Table 3.4 and Table 3.5. The movie domain knowledge graph implemented in this invention is crawled from the movie information on the IMDB website based on the ml-latest-small dataset and implemented according to the model requirements of this invention.

[0167] Table 3.4 Entity Information of Movie Domain Knowledge Graph

[0168]

[0169] Table 3.5 Relationship Information of Movie Domain Knowledge Graph

[0170]

[0171] Using the MovieID in the ml-latest-small dataset, movie data (director, lead actor, editor, movie type, etc.) was crawled from the movie website IMDB. Through the design of the schema layer of the movie domain knowledge graph, the semi-structured data crawled was transformed into structured data through knowledge extraction and stored in the database for persistent storage. Thus, the movie domain knowledge graph was constructed, providing data and technical support for the implementation of the subsequent movie recommendation system.

[0172] For the TransE model, for a fact triple (h, r, t), the relationship r is regarded as the translational distance between the head entity and the tail entity after vector space mapping. Then, the scoring function shown in formula (2-5) is used to continuously adjust the quality of the model, making h + r in the fact triple as close as possible to t, that is, h + r ≈ t;

[0173] Although the TransE model has advantages such as fast training and easy implementation, it also has problems in dealing with complex one-to-many and many-to-many relationships. The subsequent TransH model proposed by researchers has well solved this problem. The comparison diagram of the TransE and TransH models is as Figure 19 shown:

[0174] Compared with TransE, for each type of relationship r in the triple, the TransH model gives a hyperplane Wr, defines the relationship vector dr on the hyperplane Wr, and maps the original head entity h and tail entity t onto the hyperplane, denoted as hr and tr. The correct triple is required to satisfy the following formula:

[0175] h r + d r = t r (4-2)

[0176] In TransH, if there is a triple (h1, r, t) and (h2, r, t) for both the h1 and h2 vectors. Through the mapping of the hyperplane of relationship r in TranH, there is:

[0177] h 1r + d t = t r (4-3)

[0178] h 2r + d r = t r (4-4)

[0179] That is to say, the mappings of h1 and h2 on the hyperplane are the same or approximate. However, h1 and h2 themselves can be not close, that is, they can be distinguished. As Figure 20 shown:

[0180] Although TransH adds the conversion of each relationship vector to the hyperplane space compared with TransE, in fact, the overall learning parameters only increase by one item Wr compared with TransE. So its algorithm efficiency is still relatively high. Its main process is the same as that of TransE. Compared with TransE, when updating the vector representation, for each relationship, one more mapping vector Wr is learned. The projected head entity and tail entity can be expressed as:

[0181]

[0182] In addition, the scoring function and loss function for learning are different. The scoring function and the calculation function are as shown in Formulas (4-7) and (4-8):

[0183]

[0184] L = Σ (h,r,t)∈△ Σ (h′,r′,t′)∈△′ [γ + f r (h, t) - f r (h′, t′)] + (4-8)

[0185] Although the TransH model solves the problems of one-to-many and many-to-many complex relationships, through the Figure 10 process shown, we can learn that the model may introduce the problem of incorrect labels during the negative sampling process. The negative sampling strategy of the original TransH model is: First, set a certain probability to replace the head entity or the tail entity. The purpose is to make the head entity more likely to be replaced when the relationship is one-to-many, and make the tail entity more likely to be replaced when the relationship is many-to-one. This replacement strategy can improve the accuracy of negative sample generation compared with the random replacement method of TransE.

[0186] Figure 21 The flowchart of the improved TransH model is shown, which fully demonstrates how to optimize each link of data preprocessing and model training by combining the virtual damping type phase-locked strategy and the negative sample replacement mechanism, so as to improve the robustness and performance of the overall system.

[0187] The main idea of the model is: (1) Vectorize the movie entities in the knowledge graph constructed by the present invention through the translation distance model, and calculate the semantic similarity between movies; (2) Construct a content similarity matrix for content-based recommendation and a movie rating similarity matrix for collaborative filtering recommendation, and fuse the two similarity matrices with the semantic similarity of the knowledge graph respectively; (3) Mix the results of the two single models through weighted mixing and adjust the fusion weight factor. (4) Use the least squares method to reduce the relative error of the model and perform TOP-N recommendation. Its flowchart is as Figure 22 shown.

[0188] Similarity calculation is to calculate the distance between two feature vectors through a distance calculation formula (the farther the distance, the smaller the similarity; the closer the distance, the greater the similarity). For example, both M and N are K-dimensional feature vectors, expressed as M=(m1,m2,....,mk) and N=(n1,n2,....,nk). Common methods for calculating the distance between two entities M and N include: Euclidean distance, cosine distance, Pearson correlation coefficient, etc. Among them, the cosine distance formula is often used to calculate sparse data matrices, and the Euclidean distance is often used to calculate dense data matrices.

[0189] (1) Calculate the semantic similarity of movies

[0190] Through the improved TransH model, the vectorized representations of movie entities and movie relationships in the movie domain knowledge graph constructed in the present invention are shown in formula (4-9):

[0191] mi = [mi1, mi2, …, mik]T(4-9)

[0192] Where i represents the i-th movie, k represents the embedding dimension, and mik represents the value of movie mi in the k-th dimension. Since the matrix composed of movies is dense, the Euclidean distance is used to calculate the semantic similarity between movies. The Euclidean distance expression is shown in (4-10):

[0193]

[0194] When d(ma, mb) between two entities is large, it indicates a large Euclidean distance and a small similarity; when it is small, it indicates a small Euclidean distance and a large similarity. Then the similarity expression is shown in (4-11):

[0195]

[0196] It can be seen from the formula that the similarity value range is [0, 1]. When the similarity approaches 1, it indicates a large association; when the similarity approaches 0, it indicates a small association. By calculating the similarity between every two movies, a movie semantic similarity matrix can be obtained. The movie semantic similarity matrix is shown in Table 4.1:

[0197] Table 4.1 Movie semantic similarity matrix

[0198]

[0199] (2) Movie content similarity

[0200] First, define the set composed of all movies in the movie dataset as M = {m1, m2, …, mn}; the set composed of all users as U = {u1, u2, …, uk}; the set composed of all movie genres as T = {t1, t2, …, tr}, where tr represents the type of movie genre. If a movie contains this type of genre, then tr = 1; if it does not contain this movie genre, then tr = 0. Each movie can be represented by a genre vector, and the similarity between two movies can be calculated through this vector, thus obtaining the similarity situation between the two movies.

[0201] According to the above definition, the movie genre vector of movie ma can be represented as Ta = {Ta1, Ta2, …, Tar}, and movie mb can be represented as Tb = {Tb1, Tb2, …, Tbr}. After obtaining the genre vectors of the two movies, the present invention uses the cosine similarity formula to calculate the similarity between movie ma and mb. Its calculation formula is shown in Equation (4-12):

[0202]

[0203] If r (rating) represents the user's rating of a movie, then the rating of user u for movie ma can be represented by rua. If the user has not watched movie mb, then based on the rating of ma and the movie similarity calculated previously, the predicted rating of the user for movie mb can be predicted. The calculation formula for the rating of rub can be represented as shown in Equation (4-13):

[0204]

[0205] Among them, simCB(Ta, Tb) is the movie similarity calculated between movie ma and movie mb according to the cosine similarity calculation formula, rua represents the historical rating of user u for movie ma, and N(b, k) represents the set of the top k movies with the highest similarity to movie mb. After obtaining the predicted rating of the user for each movie, a recommendation list is obtained by ranking according to the rating size, and then the top N movies in the recommendation list are selected according to the rating ranking for TOP-N recommendation.

[0206] (3) Movie rating similarity

[0207] Define the set of all movies in the movie dataset as \(M = \{m_1, m_2, \ldots, m_n\}\); the set of all users as \(U = \{u_1, u_2, \ldots, u_k\}\); and use \(r_{ia}\) to represent the rating given by user \(u_i\) to movie \(m_a\). To obtain the rating data of all users for movies, it is necessary to calculate the similarity between movies pairwise. Denote the set of ratings given by all users in the dataset for a certain movie as a rating vector. Assuming there are \(k\) users in the dataset, then for movie \(m_a\), its rating vector can be expressed as \(r_a=\{r_{1a}, r_{2a}, \ldots, r_{ka}\}\); similarly, for movie \(m_b\), its rating vector is \(r_b = \{r_{1b}, r_{2b}, \ldots, r_{kb}\}\). After representing all movies with rating vectors, the similarity between movies can be calculated. The present invention adopts the cosine similarity calculation formula based on the Pearson correlation coefficient. Its calculation formula is shown in Equation (4-14):

[0208]

[0209] where \(U_a\) and \(U_b\) respectively represent the sets of users who have watched movie \(m_a\) and movie \(m_b\), \(U_{a,b}\) represents the set of users who have watched both movie \(m_a\) and movie \(m_b\), \(r_{ua}\) and \(r_{ub}\) respectively represent the rating data of users in the user set for movie \(m_a\) and movie \(m_b\), represents the average value of the ratings of all users for all movies in the dataset. The purpose of adding the average value to the formula is to reduce the error caused by different users having different evaluation criteria. The rating given by user \(u\) to movie \(m_a\) can be expressed as \(r_{ua}\), and from this, the predicted rating \(r_{ub}\) of user \(u\) for movie \(m_b\) can be calculated. Its calculation formula is shown in Equation (4-15):

[0210]

[0211] After obtaining the movie semantic similarity matrix, the movie content similarity matrix, and the movie rating similarity matrix respectively, when fusing the movie semantic similarity matrix with the movie content similarity matrix and the movie rating similarity matrix respectively, it is first necessary to determine the fusion ratio. The movie semantic similarity formula is shown in (4-11), the movie content similarity formula is shown in (4-12), and the movie rating similarity is shown in (4-14). Assuming that the fusion ratios of the semantic similarity based on the knowledge graph and the similarities of the two traditional recommendation algorithms are \(a_1\) and \(a_2\) respectively, then the similarity expressions of the traditional recommendation models fused with the movie semantic similarity based on the knowledge graph are:

[0212] sim CKG (m a ,m b ) = α1sin KG (m a ,m b)+(1-α1)sin CB (m a ,m b ) (4-16)

[0213] sim CFKG (m a ,m b )=α2sin KG (m a ,m b )+(1-α2)sin CF (m a ,m b ) (4-17)

[0214] Then, on the premise of knowing the rating of movie ma by user u, the rating formula for movie mb can be expressed as:

[0215]

[0216] Among them, R ub is the predicted rating of movie mb by user u, is the average rating of movie ma by all users in the dataset. Then, the rating calculation formula for content-based recommendation after integrating the semantic similarity of the knowledge graph and the rating calculation formula for collaborative filtering recommendation after integrating the semantic similarity of the knowledge graph are shown in Equations (4-19) and (4-20) respectively:

[0217]

[0218]

[0219] 4.2.3 Data Fitting

[0220] The recommendation result of the hybrid recommendation model is obtained by fusing the results of each single model according to a certain weight ratio w. Since the error brought by each single model is unpredictable, it is necessary to adjust the weight w to reduce the relative error J brought by the fusion of single models.

[0221] Suppose there are n single models, and the results obtained through these n single recommendation models are y(xi)j (j = 1, 2,..., n), where xi (i = 1, 2,..., k) represents the i-th user rating data, k represents the total number of user ratings in the dataset, and ω = (ω1, ω2,..., ωn) represents the weight of each single recommendation model. Then, the formula for the rating prediction result obtained by fitting the data through each single model is as follows:

[0222] Φ(x i )=ω1y(x i )1+ω2y(x i )2+…+ωn y(x i ) n (4 - 4)

[0223] Among them, In the process of determining the weights ω = (ω1, ω2, …, ωn) of the hybrid recommendation model, the relative error method can be used to solve for the weights. This is beneficial to the output stability of the hybrid recommendation model. After obtaining the scoring data Φ(xi) of the hybrid recommendation model, assuming the true scoring data is Y(xi) and the relative error is J, the relative error calculation formula is as follows:

[0224]

[0225] In Equation (4 - 5), the value of the weights ω = (ω1, ω2, …, ωp) can be obtained by solving the minimum value of the relative error J. Denote R as a matrix with all elements in an n - row and 1 - column matrix being 1, then the calculation formula for the weights ω is as shown in Equation (3 - 7):

[0226]

[0227] Among them, U is the inverse matrix of the diagonal matrix composed of the true scores Y(xi) (i = 1, 2, …, n), as shown in Equation (4 - 7); Φ is the matrix composed of the actual calculated scores obtained by each single recommendation model y(xi)j (j = 1, 2, …, p), as shown in Equation (4 - 8).

[0228]

[0229] Φ = [y(x i )1, y(x i )2, …, y(x i ) p )] m×n (i = 1, 2, …, n) (4 - 8)

[0230] After fitting the model with data, in the present invention, collaborative filtering recommendation based on the knowledge graph and content recommendation based on the knowledge graph are fitted to obtain the fusion factors α1 and α2. To make the predicted scores more accurate and the obtained model more stable, the present invention also uses the least - squares method to make the predicted scores closer to the actual scores. The least - squares fitting formula is as shown in 4 - 9:

[0231] y = ax + b (4 - 9)

[0232] Then, using the least - squares method to make the mean - square error of the model smaller, its formula can be expressed as:

[0233] y(x i ) = αΦ(xi) + b (4 - 10)

[0234] Among them, y(x i ) represents the final predicted score, and x i represents the i-th user rating data in the dataset. Φ(xi) is the predicted rating data after obtaining the two model weights α1 and α2 through data fitting. To make y(x i ) closer to Y(xi) (the true rating data), it is actually transformed into the problem of solving α and b.

[0235]

[0236]

[0237] Taking the partial derivatives of α and b in (4-12) respectively, the optimal solution expressions of the parameters α and b can be obtained as shown in (4-13) and (4-14):

[0238]

[0239] Substituting a * , b * into formula (4-10), the final predicted score formula (4-15) of the model can be obtained:

[0240]

[0241] Using formula (4-15), the final predicted score of the user for the movie can be obtained. Sort the rating data from high to low and select the top n to recommend to the user.

[0242] The dataset in the present invention is divided into two parts:

[0243] The first part is the publicly available movie dataset ml-latest-small. The detailed information of the data in this dataset is shown in Table 4.2.

[0244] Table 4.2 Content Table of ml-latest-small

[0245]

[0246] Part of the data in the dataset is as Figure 23 shown (from left to right are the movie name and type information, user rating information, and movie link information):

[0247] The second part is the movie domain knowledge graph dataset constructed in Chapter 3. This dataset mainly crawls information such as the leading actors, directors, and movie genres of each movie using a crawler on the IMDB website based on the imdbId corresponding to the movies in the ml-latest-small dataset. To avoid an overly large dataset, during the crawling process, only 3 leading actors, 1 director, and all movie genre information are crawled for each movie. This dataset is mainly used to calculate semantic similarity based on the knowledge graph for auxiliary recommendation.

[0248] 4.3.2 Evaluation Metrics

[0249] Common evaluation criteria for measuring the quality of knowledge representation learning models are Hit Rate (Hit) and Mean Rank (Meanrank). Hit refers to the hit rate of the predicted rank in the real data. Commonly seen are Hit@10, Hit@50, and Hit@100. Hit@10 refers to the number of times the predicted value of the recommendation system ranks among the real values in the dataset. If the movie predicted to be liked by the user ranks among the top ten movies that the user really likes, the value is incremented by one. The larger the value, the better the prediction effect. The same principle applies to Hit@50 and Hit@100. The average rank is the average of the ranks of the prediction results in the real results. The smaller the value, the better the model effect.

[0250] Precision, recall, and F1-score are the three most commonly used metrics for evaluating the performance of a recommendation system. Their calculation formulas are shown in Equations (4-9), (4-10), and (4-11) respectively. The meaning of each formula is shown in Table 4.3:

[0251]

[0252] Table 4.3 Confusion Matrix

[0253]

[0254] As shown in Equation (4-9), precision refers to the proportion of movies liked by the recommended users among the total recommended movies; recall refers to the proportion of movies liked by the recommended users among all the movies liked by the users; the F1-score is the harmonic mean of precision and recall. Its purpose is to find a balance point to make both precision and recall reach the best. The larger the values of these three metrics, the better the model effect. In the present invention, a score greater than 3.5 is set as liked by the user, and less than 3.5 is set as not liked.

[0255] At the same time, in order to verify the stability of the model, the present invention also selects the Mean Absolute Error (MAE) as the metric for evaluating the model stability. Its calculation formula is shown in Equation (4-12):

[0256]

[0257] Among them, n represents the number of ratings, y(xi) represents the predicted rating of the model, and Y(xi) represents the true rating. The smaller the MAE value, the smaller the gap between the predicted value and the true value of the model, and the better and more stable the model performance.

[0258] The improved TransH model, the TransH model, and the TransE model were compared to determine the effectiveness of the models in different dimensions. The main parameters during the experiment included the learning rate l, the vector embedding dimension embeding_dim, the margin, the number of triples processed each time BatchSize, and the number of iterations n. The learning rate l = 0.01, the margin = 1, and the embedding vector dimensions were (100, 200, 300, 400, 500, 600), BatchSize = 100, and the iteration was 400 times. The hit rates at different dimensions were calculated respectively, and the Meanrank was used to verify the effectiveness of the model.

[0259] From Figure 24 and Figure 25 As shown, the hit rate of the improved TransH model first increases and then decreases with the increase of the embedding dimension, reaching the peak hit@10 = 0.57 when embeding_dim = 300. The trends of the TransH and TransE models are roughly the same. In terms of the Meanrank index, it first decreases and then increases with the increase of the embedding dimension, and the improved TransH model is always lower than the TransH and TransE models in terms of the Meanrank index. Compared with the traditional Trans model, the improved model can make effective recommendations for user behavior with higher hit rates and average ranks, indicating the effectiveness of the model.

[0260] In addition to verifying the improved TransH model of the present invention, an experimental analysis was also carried out on the hybrid recommendation model. The hybrid recommendation model based on the knowledge graph proposed by the present invention was compared with the content-based recommendation model, the item-based collaborative filtering recommendation, and the user-based collaborative filtering recommendation respectively. For the convenience of subsequent expression and experimental comparison, the recommendation models used in the present invention are listed in the table in the form of English abbreviations as shown in Table 4.4.

[0261] Table 4.4 English Abbreviations of Algorithms

[0262]

[0263] Compare and verify mainly from the following aspects: (1) Determine the fusion ratios α1 and α2 of the movie domain semantic similarity obtained based on the movie domain knowledge graph with the similarity of content-based recommendation and item-based collaborative filtering recommendation respectively. (2) Verify the effectiveness of the hybrid recommendation model. (3) Compare the mean absolute errors of several models and some performance data of several traditional recommendation algorithms.

[0264] The specific experimental results are as follows:

[0265] (1) Determine the fusion ratios α1, α2;

[0266] When the number of recommended movies K = 10 is fixed, to determine the fusion ratio α1 of the movie semantic similarity based on the knowledge graph and the CB similarity, let α1 start from 0 and increase to 1 in steps of 0.1 for experiments respectively. When α1 = 0, it is content-based recommendation. When α1 = 1, it is movie semantic similarity, and the similarity of content-based recommendation is not used for recommendation. The experimental results of CB with different fusion ratios are as Figure 26 shown

[0267] Figure 26 The horizontal axis represents the value of the fusion ratio α1, and the vertical axis represents the precision, recall, and F1 value of CB with different fusion ratios. From Figure 26 it can be seen that in the ml-latest-small dataset, when the number of recommended movies K remains fixed, as the fusion ratio α1 gradually increases, the precision of the fused algorithm first increases and then decreases, and the precision reaches the maximum value when α1 = 0.8; the recall and F1 value of the fused algorithm both first increase and then gradually tend to be stable. The larger the F1 value, the more robust the model. When the F values are not very different, since the precision reaches the maximum at α1 = 0.8, it indicates that the accuracy of the recommendation is the highest at this time. Therefore, according to the experimental results, the fusion ratio α1 = 0.8 can be set.

[0268] When the number of recommended movies K = 10 is fixed, to determine the fusion ratio α2 of the movie semantic similarity based on the knowledge graph and ItemCF, similarly let α2 start from 0 and increase to 1 in steps of 0.1 for experiments respectively. When α2 = 0, it is item-based collaborative filtering recommendation without fusing movie semantic similarity. When α2 = 1, it indicates complete fusion of movie semantic similarity, and the similarity of ItemCF is not used for recommendation. The experimental results of ItemCF with different fusion ratios are as Figure 26 shown.

[0269] Figure 27 The horizontal axis represents the value of the fusion ratio α2, and the vertical axis represents the precision, recall, and F1 value of ItemCF with different fusion ratios. From Figure 27It can be seen that in the ml-latest-small dataset, when the fixed number of recommended movies K remains unchanged, as the fusion ratio α2 gradually increases, the precision of the fused algorithm first increases and then decreases, and the precision reaches the maximum value when α2 = 0.5; the recall rate and F1 value of the fused algorithm both show a trend of first increasing and then decreasing. The larger the F1 value, the more robust the model. Since the F1 value reaches the maximum when α2 = 0.5 and the precision is also the highest at this time, according to the experimental results, the fusion ratio α2 can be set to 0.5

[0270] Verify the effectiveness of the Mix-KG model.

[0271] Based on the fusion ratios α1 and α2 determined in experiment (1), while fixing the fusion ratios α1 = 0.8 and α2 = 0.5, when the number of recommended movies K is set to 10, 20, 30, 40, 50, 60, and 70 respectively, conduct experimental comparisons of CB, Item-CF, User-CF, and CFCKG in the ml-latest-small dataset. The precision of the four algorithms is as Figure 28 shown, the recall rate is as Figure 29 shown, and the F1 value is as Figure 30 shown.

[0272] In the ml-latest-small dataset, when the fusion ratios α1 and α2 are fixed, as the number of recommended movies K increases, the precision of the four algorithms shows a downward trend, and the recall rate shows an upward trend. As the number of recommended movies K increases, the F1 value of the four algorithms shows a slow upward trend. When the four algorithms have the same K value, the Mix-KG model has the highest precision, recall rate, and F1 value. Followed by the User-CF model and the Item-CF model, and the CB model performs the worst. Among them, the F1 values of Mix-KG and CB reach the maximum when K = 60, and the F1 values of User-CF and Item-CF reach the maximum when K = 50. The larger the F1 value, the more robust the model. The Mix-KG model proposed in the present invention shows good experimental results in terms of precision, recall rate, and F1 value and other indicators.

[0273] (3) Error verification

[0274] Based on the fusion ratios α1 and α2 determined in experiment (1), while fixing the fusion ratios α1 = 0.8 and α2 = 0.5, when the number of recommended movies K is set to 10, 20, 30, 40, 50, 60, and 70 respectively, conduct error experimental comparisons of CB, Item-CF, User-CF, and Mix-KG in the ml-latest-small dataset. The mean absolute error MAE of the four algorithms is as Figure 31 shown.

[0275] From Figure 31 It can be seen that in the dataset ml-latest-small, when the fusion ratios α1 and α2 are fixed and at the same K value, the mean absolute error (MAE) of the Mix-KG model proposed in the present invention is the smallest. The MAE value of User-CF is slightly smaller than that of Item-CF but larger than the MAE value of Mix-KG. The above three algorithms generally achieve the minimum MAE value when K = 20; the MAE value of CB is the largest, achieving the minimum value approximately when K = 30, but it is still larger than the MAE values of the other three algorithms. The smaller the MAE value, the smaller the error of the recommendation model when calculating the obtained scores, and the better the effect of TOP-N recommendation. The mean absolute error (MAE) of the Mix-KG model proposed in the present invention is smaller than the other three algorithms, reflecting the superiority of the Mix-KG model.

[0276] Firstly, the main principle of the TransH model is introduced, and an improved negative sampling optimization strategy for the TransH model is proposed. Through the improved TransH model, the movie semantic similarity matrix, movie content similarity matrix, and movie rating similarity matrix are obtained respectively. The optimal fusion ratio between the models is determined through experiments. Through comparative experiments, the advantages of the hybrid model compared to the single model are proven, and better results are achieved in the evaluation metrics, proving the effectiveness of the model.

[0277] 5. Design and Implementation of the Movie Recommendation System

[0278] 5.1 System Functional Requirement Analysis

[0279] Functional requirement analysis is a key link in system development. Its purpose is to define the modules and functions involved in the system development process. Functional requirement analysis generally starts from the specific objects using the system and conducts functional requirement analysis for different user roles from the perspective of the users. For this system, the main roles are administrators and users.

[0280] For the user role, first, the user needs to enter the username and password to log in to the system. If it is a new user, registration is required before logging in to the system. The user can also modify personal information, search for movies, view the watched movies and display them in the form of a knowledge graph, view movies online and the relationships between related movies, rate and evaluate movies, and receive real-time recommendations. The user case diagram is as Figure 32 shown:

[0281] For the administrator role, the administrator permissions need to be granted by the super account. The initial super account directly writes to the database. The administrator can view user account-related information, reset user passwords, edit user information, add users, and delete users. At the same time, the administrator can also manage movie information and maintain movie information, including editing, adding, and deleting.

[0282] The administrator use case diagram is shown in Figure 33:

[0283] 5.2 System Design

[0284] 5.2.1 System Architecture Design

[0285] The movie recommendation system designed in the present invention adopts a B / S architecture. The front end uses the vue framework, and the back end uses node.js. The system architecture and the functions implemented by each layer are as Figure 34 shown.

[0286] The main function of the application layer is to collect user requests and return the results of user requests. It is an important part of the interaction between the system and users. The business layer is below the application layer and above the data layer. Its main role is to process user requests, process the data returned by the data layer, and return the processed data to the business layer for display. The main logic of the system is implemented in the business layer. For the system, the quality of the business layer design directly affects the system usage experience. The main role of the data layer is to connect to the database, read and write the database, and data caching. The underlying database is mainly for persistent storage.

[0287] The system functions of the present invention are divided into the following three parts: user management, movie management, and administrator. The detailed functions of each module are as Figure 35 shown. The system modules designed in the present invention are independent of each other, which is convenient for adding functions and maintaining the system in the later stage. At the same time, it can improve the system stability and ensure that the problems of one module will not affect the use of other modules.

[0288] (1) User Management Module

[0289] User management mainly includes the user login function. New users need to contact the administrator to register an account and fill in their personal information. After successful registration, they can log in to the system. The user rating function allows users to rate and comment on their historical movie viewings and fill in comments. The historical movie viewing records mainly save the historical movie viewing records of 610 users in the dataset, as well as the given ratings. For new users, they have no historical movie viewing records.

[0290] (2) Movie Management Module

[0291] The movie management module mainly includes functions such as movie search, movie recommendation, and movie relationship query. The movie search function supports users to search for movies according to the movie name and display the main information of the movies in the form of a knowledge graph, including information such as movie stars, actors, genres, directors, etc. Entering the movie IDs of two movies also supports searching for the relationship between the two movies; the movie recommendation function is based on the user's historical viewing records, obtains the user's interest preferences through the hybrid recommendation model proposed in Chapter 4 of this paper, and selects ten movies that the user may be interested in from the movies to be recommended for recommendation.

[0292] The administrator module mainly realizes the maintenance functions of user and movie information. The user management function includes new user registration, user information modification, and password reset to ensure that users can log in smoothly and manage their personal information; the movie information maintenance function supports administrators to add, modify, and delete movie records to ensure the accuracy and integrity of movie data in the database in real time.

[0293] In terms of database design, the entities included in the system, such as users, movie genres, movies, administrators, and user ratings, are determined through an E-R diagram, and the attributes of each entity and their mutual relationships are defined in detail. According to the E-R diagram design, the system has established user information tables, movie information tables, administrator information tables, user rating tables, movie genre tables, and movie review tables. Each table clearly defines the primary key, data type, and field constraints to ensure data consistency and efficient query.

[0294] The system function design covers three modules: user management, movie management, and administrator. The user management module realizes login, historical viewing record query, and user evaluation; the movie management module supports movie search, movie relationship query, and personalized movie recommendation, and displays movie-related information in the form of a knowledge graph; the administrator module is responsible for the creation and maintenance of user and movie information to ensure the continuous update of system data. Each function module is displayed through an intuitive interface diagram, such as the user login page, home page, historical viewing knowledge graph, and movie recommendation interface, etc.

[0295] System development and testing are carried out in a strict development environment. The test environment includes key components such as Windows 10, 16GB of memory, Neo4j, MongoDB, MySQL, Python, Node.js, and Vue. After environment configuration and multi-module testing (such as user login, movie search, recommendation results, and administrator information maintenance, etc.), the system runs stably, successfully realizes a movie recommendation system based on the knowledge graph hybrid recommendation model, and achieves the expected results in terms of personalized recommendation, data storage, and information display.

[0296] Open VScode, enter the command "npm run sever" in the front-end directory to start the front-end service, enter "node server.js" in the back-end directory, open the MongoDB client to establish a database connection, and open Neo4j to start the graph database. All relevant services of the system are started. Next, enter "localhost:8080 / " in the browser to enter the login page of this system. Since the original data of the system uses the data in the dataset "ml-latest-small", which contains 100,000 rating records of 610 users for different movies, we use user 3 as an example to test this system during testing.

[0297] The user management module implements core functions such as user login, display of historical movie viewing records, and user reviews. After entering the account password through the login page, the user successfully enters the system home page; if the information is incorrect, a login failure prompt is displayed. The system can display the user's historical rating records and visually show the association relationships between movies, directors, leading actors, genres, etc. in the user's movie viewing history in the form of a graph. At the same time, the user is allowed to review the movies watched.

[0298] The movie management module provides functions for movie search, movie recommendation, and movie relationship query. Users can quickly obtain detailed movie information by entering the movie name through the online search function; after clicking "Start Recommendation", the system uses the results of the offline generated hybrid recommendation model based on the knowledge graph to generate a personalized recommendation list for the user and display it in tabular form; in addition, users can also query the association relationship between two movies by entering the movie ID, and the system presents the relevant connections in the form of a graph.

[0299] The administrator module is mainly used for the maintenance of user information and movie information. The administrator can perform operations such as new user registration, modification of user information, password reset, and addition, editing, and deletion of movie information to ensure the timely update and accuracy of system data. Each function interface has been tested, showing that the system is stable and efficient in information management.

[0300] Overall, the movie recommendation system based on the hybrid recommendation model of the knowledge graph realizes the personalized and visual display of movie recommendations through constructing a knowledge graph in the movie field, improving the hybrid recommendation algorithm (integrating content-based, collaborative filtering, and semantic similarity), and a perfect system architecture design. The system performs better than traditional algorithms in terms of precision, recall rate, and F1 value, and at the same time provides improvement directions for subsequent cloud online recommendations, knowledge graph expansion, and introduction of time factors.

[0301] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A method for applying a hybrid recommendation model based on knowledge graph in a movie recommendation system, characterized in that: The method comprises the following steps: S101, building a knowledge graph in the film field; S102, constructing an improved hybrid recommendation model based on the movie field knowledge graph; S103, implementing a movie recommendation system, and applying the hybrid recommendation model obtained in step S102 to movie recommendations.

2. The method according to claim 1, characterized in that Step S101 further includes: (a) We selected a public dataset containing movie names, movie genres, user ratings, and external link information, and collected information about the director, starring actor, screenwriter, and semi-structured movie genre from the IMDB website using web crawler technology based on the MovieID of each movie; (b) Remove duplicates, fill missing values, and convert the format of the collected data to generate a structured movie dataset and store it in CSV format; (c) Extracting entities and their relationships in the structured data into triples, and using the LOAD CSV command to import the data into the Neo4j graph database to complete the construction of the knowledge graph.

3. The method according to claim 2, characterized in that Step S101 also includes designing a model layer of the knowledge graph in the film field, and the model layer design includes: (a) Abstract movie knowledge into movie (MOVIE), person (PERSON), movie type (GENRES) and movie identifier (ID) entity categories; (b) Determine the attributes and relationships between entity categories, and define the value range and constraints of each attribute; (c) The designed ontology rules are used to guide the instantiation of entities and relationships in the subsequent data layer.

4. The method according to claim 1, characterized in that Step S102 includes: (a) Using the constructed movie domain knowledge graph, the movie entities and relationships are vectorized through the translation distance model (such as TransE and its improved TransH model) to calculate the movie semantic similarity; (b) Using movie content features and user rating data, we construct a content-based recommendation similarity matrix and a collaborative filtering-based similarity matrix. (c) performing weighted fusion of the movie semantic similarity obtained in step (a) and the similarity matrix obtained in step (b) according to predetermined fusion ratios α1 and α2 to obtain a single recommendation model scoring result based on the knowledge graph; (d) By adjusting the weights of the results of each single model and using the least squares fitting method to reduce the relative error of the fusion process, the prediction score of the final hybrid recommendation model is obtained.

5. The method according to claim 4, characterized in that In the training process of the improved TransH model, the negative sample generation strategy is optimized by adopting the following steps: (a) Based on the four types of relationships in the graph, the entities in the movie domain are clustered by relationship type using the k-means clustering method, with a fixed number of clusters of 4; (b) When generating negative example triplets, entities of different categories from the current positive example are selected as replacement objects for the head entity or tail entity to improve the accuracy of negative sample generation and reduce the risk of introducing incorrect labels.

6. The method according to claim 1, characterized in that Step S103 includes: (a) Design the system architecture according to the requirements of the movie recommendation system and divide it into construction module, improvement module and implementation module; (b) Design and implement movie management, graph display, movie search, movie relationship query, user evaluation and movie recommendation functional modules, and implement visual display of recommendation results in the system.

7. The method according to claim 4, characterized in that Step S102 also includes using the obtained movie entity vectors to calculate the semantic similarity between movies using the Euclidean distance formula, and respectively using cosine similarity or a method based on the Pearson correlation coefficient to calculate the movie content similarity and movie rating similarity, and then fusing them according to the following formula: sim_CKG(ma,m_b)=α1·sim_KG(ma,m_b)+(1–α1)·sim_CB(ma,m_b) sim_CFKG(ma,m_b)=α2·sim_KG(ma,m_b)+(1–α2)·sim_CF(ma,m_b) Among them, α1 and α2 are fusion ratio parameters determined by experiments.

8. The method according to claim 1, characterized in that The method further includes evaluating the recommendation performance of the hybrid recommendation model, and the evaluation step includes: (a) Use hit rate (Hit@N), mean rank (Meanrank), precision, recall and F1 value indicators to quantify the model recommendation effect; (b) The mean absolute error (MAE) is used to evaluate the error between the model's predicted score and the true score, and the optimal fusion parameters and the stability of the hybrid model are determined through comparative experiments.

9. A movie recommendation system based on a hybrid recommendation model of knowledge graph, characterized in that: The system includes: (a) A knowledge graph construction module in the film field, which is used to collect, process and integrate film-related data, construct a knowledge graph containing movie, character, film type entities and their attributes and relationships, and store the data in a graph database; (b) An improved hybrid recommendation model module is used to use the constructed knowledge graph to vectorize movie entities through an improved TransH model and negative sample optimization strategy, and calculate the movie semantic similarity, content-based similarity, and score-based similarity respectively, and then perform weighted fusion according to a predetermined fusion ratio, and reduce the prediction error through data fitting, so as to generate the final recommendation score; (c) Movie management module, which is used to realize movie search, movie relationship query, user evaluation, user historical viewing record management and visual display of recommendation results, so as to realize personalized recommendation of movies.

10. A computer system, characterized in that: The system includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following functions are realized: (a) Constructing a knowledge graph in the film field, including collecting film-related data from public datasets and external websites, using web crawlers and data preprocessing techniques to generate structured film data, and importing the data into a graph database to form a knowledge graph containing film, character, and film genre entities and their relationships; (b) Implement a hybrid recommendation model based on knowledge graph. This model uses the improved TransH model and negative sample generation optimization strategy to vectorize movie entities, calculate the movie semantic similarity, content similarity and rating similarity respectively, and then perform weighted fusion according to the preset fusion ratio to obtain the final recommendation score through data fitting; (c) Provide a user interaction interface to enable user login, movie information query, movie relationship display, user rating and comment input, and display the TOP-N movie recommendation results to the user in a graphical manner based on the output of the recommendation model.

Citation Information

Patent Citations

  • Personalized intelligent clothes matching recommendation method combined with knowledge graph

    CN112612973A

  • Movie recommendation method fusing knowledge graph and collaborative filtering

    CN115774818A

  • Course recommendation system based on knowledge graph and graph attention network

    CN115840853A