A method, device, equipment and storage medium for expanding user interest portraits
By utilizing knowledge graphs and context prediction algorithms, the similarity between interest labels is calculated and user interest profiles are expanded, and the problem of limited diversity of user interest profiles in the prior art is solved, and a richer personalized service is achieved.
Patent Information
- Application Number
- CN202011233447.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-11-06
AI Technical Summary
When generating user interest portraits, the prior art only based on the user's click logs, resulting in the diversity of user interest portraits being limited and cannot be effectively expanded, which in turn affects the richness of personalized services.
By obtaining the target knowledge graph, generating the target entity sequence, and using a context prediction algorithm to determine the entity vector, calculate the similarity between the interest labels based on the similarity of the entity vector and the mapping relationship of the interest labels, thereby expanding the user interest portrait.
It realizes the rapid and accurate expansion of user interest portraits, provides richer personalized services, and improves user experience.
Smart Images

Figure CN112232889B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Artificial Intelligence (AI) technology, and in particular, to a method, apparatus, device, and storage medium for expanding a user interest profile. Background Art
[0002] A user interest profile is essentially a collection of user interest tags, which can reflect the content that a user is interested in. In the era of network big data, many network platforms need to provide corresponding personalized services for users based on the user interest profile, such as personalized recommendation, personalized search, accurate advertisement push, intelligent marketing, etc. Nowadays, how to accurately determine the user interest profile has become the focus of attention of many network platforms.
[0003] The related technology currently mainly generates a user interest profile by analyzing the click logs of users. Specifically, according to the click situation of users on the content on the network platform, weights can be configured for the tags corresponding to the clicked content, and then the tags with relatively high corresponding weights can be selected to form the user interest profile of the user. For example, assuming that user A often clicks on articles or videos related to basketball, the server can add the tag "basketball" to the user interest profile of user A.
[0004] However, the above-mentioned method for generating a user interest profile has the following problems: Generating a user interest profile only based on the click logs of users will limit the diversity of the user interest profile and is not conducive to the expansion of the user interest profile; correspondingly, the personalized services provided for users based on the user interest profile generated in this way will also be relatively single, affecting the user experience. Summary of the Invention
[0005] The embodiments of this application provide a method, apparatus, device, and storage medium for expanding a user interest profile, which can quickly and accurately expand the user interest profile and is conducive to network platforms providing richer personalized services for users.
[0006] In view of this, the first aspect of this application provides a method for expanding a user interest profile, and the method includes:
[0007] Obtain a target knowledge graph; the target knowledge graph is used to represent the association relationship between target entities, and the target entities are entities related to the target network platform;
[0008] Generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having an association relationship in the target knowledge graph;
[0009] Determine the entity vectors corresponding to the target entities in the target knowledge graph based on the context prediction algorithm according to the target entity sequence;
[0010] Determine a first similarity between interest tags on the target network platform according to the similarity between entity vectors corresponding to target entities in the target knowledge graph and the mapping relationship between the target entities and interest tags on the target network platform;
[0011] Based on the first similarity, expand the user interest profile on the target network platform.
[0012] A second aspect of the present application provides a user interest profile expansion device, the device includes:
[0013] A knowledge graph acquisition module, configured to acquire a target knowledge graph; the target knowledge graph is used to represent the association relationship between target entities, and the target entities are entities related to the target network platform;
[0014] An entity sequence generation module, configured to generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having an association relationship in the target knowledge graph;
[0015] An entity vector determination module, configured to determine entity vectors corresponding to the target entities in the target knowledge graph based on the context prediction algorithm according to the target entity sequence;
[0016] A first label similarity determination module, configured to determine a first similarity between interest tags on the target network platform according to the similarity between entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and interest tags on the target network platform;
[0017] A user profile expansion module, configured to expand the user interest profile on the target network platform based on the first similarity.
[0018] A third aspect of the present application provides a device, the device includes a processor and a memory:
[0019] The memory is used to store a computer program;
[0020] The processor is configured to execute the steps of the user interest profile expansion method as described in the first aspect above according to the computer program.
[0021] A fourth aspect of the present application provides a computer-readable storage medium, the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the steps of the user interest profile expansion method as described in the first aspect above.
[0022] A fifth aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the user interest portrait expansion method described in the first aspect above.
[0023] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0024] The embodiments of the present application provide a method for expanding a user interest portrait. The method innovatively proposes a solution for expanding a user interest portrait based on a knowledge graph. Specifically, in the user interest portrait expansion method provided by the embodiments of the present application, first, a target knowledge graph for characterizing the association relationship between target entities is obtained, where the target entities are entities related to a target network platform; then, a target entity sequence is formed by using multiple target entities having an association relationship in the target knowledge graph, and an entity vector corresponding to the target entity in the target knowledge graph is determined based on a context prediction algorithm according to the formed target entity sequence; furthermore, according to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest tags on the target network platform, the similarity between the interest tags on the target network platform is determined; finally, the user interest portrait on the target network platform is expanded based on the similarity between the interest tags. The above method is based on a knowledge graph covering a large number of entities and relationships between entities, determines the similarity between entities in the knowledge graph, and converts the similarity between entities into the similarity between interest tags according to the mapping relationship between entities and interest tags, and then expands the user interest portrait based on the similarity between interest tags; in this way, rapid and accurate expansion of the user interest portrait is achieved, and furthermore, it is beneficial for the network platform to provide richer personalized services based on the expanded user interest portrait. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of an application scenario of the user interest portrait expansion method provided by the embodiments of the present application;
[0026] Figure 2 It is a schematic flowchart of a user interest portrait expansion method provided by the embodiments of the present application;
[0027] Figure 3 It is a schematic diagram of an exemplary triple in the basic knowledge graph provided by the embodiments of the present application;
[0028] Figure 4 It is a schematic diagram of training a skip-gram model provided by the embodiments of the present application;
[0029] Figure 5 Schematic diagram for training SLIM based on the basic user portrait matrix provided by the embodiments of the present application;
[0030] Figure 6 Schematic flow chart of another user interest portrait extension method provided by the embodiments of the present application;
[0031] Figure 7 Schematic structural diagram of the first user interest portrait extension device provided by the embodiments of the present application;
[0032] Figure 8 Schematic structural diagram of the second user interest portrait extension device provided by the embodiments of the present application;
[0033] Figure 9 Schematic structural diagram of the third user interest portrait extension device provided by the embodiments of the present application;
[0034] Figure 10 Schematic structural diagram of the fourth user interest portrait extension device provided by the embodiments of the present application;
[0035] Figure 11 Schematic structural diagram of the terminal device provided by the embodiments of the present application;
[0036] Figure 12 Schematic structural diagram of the server provided by the embodiments of the present application. Detailed implementation manners
[0037] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0038] In the description, claims and the above-mentioned drawings of this application, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0039] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0040] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0041] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0042] The solution provided by the embodiments of this application relates to the technology of expanding user interest portraits in artificial intelligence, and will be specifically described through the following embodiments:
[0043] The related technology currently mainly generates user interest portraits by analyzing users' click logs, and generating user interest portraits in this way may limit the diversity of user interest portraits, and further affect the personalized services provided by the network platform for users.
[0044] In view of the problems existing in the above related technologies, the embodiments of the present application provide a method for expanding a user interest profile. This method can determine the similarity between interest tags on a network platform based on a knowledge graph covering a large number of entities and the relationships between entities, and then expand the existing user interest profile on the network platform accordingly.
[0045] Specifically, in the method for expanding a user interest profile provided by the embodiments of the present application, first, a target knowledge graph for characterizing the association relationships between target entities is obtained. Here, the target entities are entities related to the target network platform. Then, a target entity sequence is formed by using multiple target entities with association relationships in the target knowledge graph, and based on a context prediction algorithm, entity vectors corresponding to the target entities in the target knowledge graph are determined according to the formed target entity sequence. Furthermore, according to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest tags on the target network platform, the similarity between the interest tags on the target network platform is determined. Finally, the user interest profile on the target network platform is expanded based on the similarity between the interest tags.
[0046] The above method for expanding a user interest profile is based on a knowledge graph covering a large number of entities and the relationships between entities, determines the similarity between entities in the knowledge graph, and according to the mapping relationship between entities and interest tags, converts the similarity between entities into the similarity between interest tags, and then expands the user interest profile based on the similarity between interest tags. In this way, the user interest profile can be expanded quickly and accurately, and thus, it is beneficial for the network platform to provide richer personalized services for users based on the expanded user interest profile.
[0047] It should be understood that the method for expanding a user interest profile provided by the embodiments of the present application can be applied to electronic devices with data processing capabilities, such as terminal devices or servers. Among them, the terminal device can specifically be a computer, a smart phone, a tablet computer, a personal digital assistant (Personal Digital Assitant, PDA), etc.; the server can specifically be an application server or a Web server. In actual deployment, it can be an independent server, or a cluster server or a cloud server.
[0048] To facilitate the understanding of the method for expanding a user interest profile provided by the embodiments of the present application, the following takes the execution subject of the method for expanding a user interest profile as a server as an example to exemplarily introduce the application scenarios applicable to the method for expanding a user interest profile.
[0049] See Figure 1 , Figure 1 is a schematic diagram of the application scenario of the method for expanding a user interest profile provided by the embodiments of the present application. AsFigure 1 As shown, in this application scenario, it includes a server 110, a database 120, and a database 130. The server 110 can access the database 120 and the database 130 through a network, or the database 120 and the database 130 can also be integrated in the server 110. Among them, the server 110 is used to execute the user interest portrait extension method provided in the embodiments of the present application, the database 120 is used to store a knowledge graph, and the database 130 is used to store the user interest portraits on the target network platform.
[0050] In practical applications, the server 110 can retrieve a target knowledge graph from the database 120. The target knowledge graph can represent the association relationships between target entities, where the target entities are entities related to the target network platform. The target network platform can be a network platform that provides personalized services to users based on user interest portraits. For example, it can be a network platform that needs to recommend information such as articles, videos, audio, and products to users, or a network platform that needs to push advertisements to users. The present application does not make any limitations on the target network platform here.
[0051] After the server 110 obtains the target knowledge graph, it can use the random walk algorithm to generate a number of target entity sequences based on the target knowledge graph. Each target entity sequence is essentially a sequence composed of multiple target entities with association relationships in the target knowledge graph. Then, the server 110 can use a context prediction algorithm (such as the skip-gram algorithm, etc.) to determine the entity vectors corresponding to each target entity in the target knowledge graph according to the generated target entity sequences. Furthermore, the server 110 can calculate the similarity between the entity vectors corresponding to each target entity in the target knowledge graph, that is, calculate the similarity between the entity vectors corresponding to every two target entities in the target knowledge graph, and convert the calculated similarity between the entity vectors into the similarity between interest tags according to the mapping relationship between each target entity and each interest tag on the target network platform, and record this similarity as the first similarity between interest tags.
[0052] Finally, the server 110 can retrieve the user interest portraits on the target network platform from the database 130 and expand the user interest portraits based on the first similarity between the above-mentioned interest tags, that is, expand the interest tags that were not included before in the user interest portraits according to the similarity between the interest tags.
[0053] Optionally, in order to more accurately expand the user interest portrait on the target network platform, after the server 110 retrieves the user interest portrait on the target network platform from the database 130, it can determine the second similarity between each interest tag on the target network platform based on the retrieved user interest portrait; Exemplarily, the server 110 can train a Sparse Linear Model (SLIM) based on the retrieved user interest portrait, and then use the trained SLIM to represent the second similarity between each interest tag.
[0054] When the server 110 has determined both the first similarity and the second similarity between each interest tag, the server 110 can determine the target similarity between each interest tag according to the first similarity and the second similarity between each interest tag, and then expand the user interest portrait based on the target similarity between each interest tag.
[0055] It should be understood that Figure 1 The application scenarios shown are only examples. In actual applications, the user interest portrait expansion method provided by the embodiments of the present application can also be applied to other application scenarios. For example, the user interest portrait expansion method provided by the embodiments of the present application can be executed by a terminal device. The present application does not make any limitation on the application scenarios of this user interest portrait expansion method.
[0056] The user interest portrait expansion method provided by the present application will be introduced in detail below through method embodiments.
[0057] See Figure 2 , Figure 2 is a schematic flowchart of the user interest portrait expansion method provided by the embodiments of the present application. For ease of description, the following embodiments will still be introduced by taking the execution subject of this user interest portrait expansion method as the server as an example. As Figure 2 shown, the user interest portrait expansion method includes the following steps:
[0058] Step 201: Obtain a target knowledge graph; the target knowledge graph is used to represent the association relationship between target entities, and the target entities are entities related to the target network platform.
[0059] A knowledge graph is a collection of knowledge that represents the association relationships between nodes in a graph form. The nodes in the knowledge graph correspond to entities, and the association relationships between the nodes correspond to the association relationships between the entities. For example, "Zhao Moumou" is an entity in the knowledge graph, and "Li Moumou" is another entity in the knowledge graph. In the knowledge graph, "Zhao Moumou" and "Li Moumou" are associated through the "spouse" relationship, and "Zhao Moumou" - "spouse" - "Li Moumou" forms a triple. In practical applications, a knowledge graph with more than one type of included nodes or more than one link relationship between nodes can also be called a Heterogeneous Information Network (HIN).
[0060] In the technical solution provided in the embodiment of the present application, the server can first obtain a target knowledge graph, which can represent the association relationships between target entities. Here, the target entities are entities related to the target network platform. Considering that the knowledge graph constructed based on all the information in the network is usually very large, expanding the user interest portrait on the target network platform based on this knowledge graph requires a large amount of computing power, and there is a lot of information in the network that is not relevant to the target network platform. If such information is also taken into account when expanding the user interest portrait on the target network platform, it will only consume unnecessary computing power. Based on this, the method provided in the embodiment of the present application is based on the target knowledge graph used to represent the association relationships between target entities related to the target network platform to expand the user interest portrait on the target network platform.
[0061] It should be understood that in practical applications, the server can autonomously extract the target knowledge graph from the basic knowledge graph. Here, the basic knowledge graph is a knowledge graph constructed based on all the information in the network; it can also directly obtain the target knowledge graph from other devices. The present application does not make any limitations on the implementation manner of the server to obtain the target knowledge graph.
[0062] The following introduces the implementation manner of the server extracting the target knowledge graph from the basic knowledge graph.
[0063] The server can select entities that meet the preset conditions as target entities from the basic knowledge graph. Here, the preset conditions include at least one of the following: the entity type is a preset type, and the entity popularity exceeds the preset popularity threshold; then, according to the association relationships of the selected target entities in the basic knowledge graph, the target knowledge graph is determined.
[0064] Specifically, the basic knowledge graph can be composed of several triples with association relationships. Each triple consists of a head entity, an entity relationship, and a tail entity. Figure 3 For the schematic diagram of an exemplary triple in the basic knowledge graph, such asFigure 3 As shown, the head entity included in this triple is "Zhao Moumou", the entity relationship is "spouse", and the tail entity is "Li Moumou". Each entity in the basic knowledge graph also includes a set of attribute information corresponding to it. Exemplarily, the attribute information included in each entity includes but is not limited to entity type, entity name, entity popularity, etc.
[0065] When the server extracts the target knowledge graph from the basic knowledge graph, it can select entities that meet the preset conditions from the basic knowledge graph as target entities. Exemplarily, the server can select entities with a preset entity type from the basic knowledge graph as target entities. Taking the target network platform as a video playback platform as an example, the server can set the preset type to include people, movies, TV dramas, variety shows, etc., and then select entities with the above preset type as target entities from the basic knowledge graph; Exemplarily, the server can also select entities whose entity popularity exceeds a preset popularity threshold from the basic knowledge graph. Still taking the target network platform as a video playback platform as an example, the server can set the preset popularity threshold to 500, and then select entities with an entity popularity exceeding 500 from the basic knowledge graph.
[0066] It should be understood that in practical applications, the server can filter target entities only based on entity type, or only based on entity popularity, or can also filter target entities based on both entity type and entity popularity, or the server can also filter target entities based on other entity attribute information. This application does not make any limitation on the preset conditions based on which target entities are selected from the basic knowledge graph.
[0067] After the server selects target entities related to the target network platform from the basic knowledge graph, it can extract the association relationships of the selected target entities in the basic knowledge graph. Furthermore, based on the selected target entities and the association relationships of the target entities in the basic knowledge graph, a target knowledge graph suitable for expanding the user interest portrait for the target network platform is constructed.
[0068] Step 202: Generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having an association relationship in the target knowledge graph.
[0069] After the server obtains the target knowledge graph, it can generate a number of target entity sequences based on this target knowledge graph. Specifically, the server can use a series of target entities having an association relationship with each other to form a target entity sequence based on the association relationships between the target entities in the target knowledge graph.
[0070] In practical applications, the server may use the Random Walk algorithm to generate the above-mentioned target entity sequence based on the target knowledge graph. The Random Walk algorithm is essentially a mathematical statistical model. Based on the Random Walk algorithm, a series of trajectories can usually be generated, and each step in the walking process is random. Considering that the randomness of the Random Walk algorithm may generate a large number of relatively long target entity sequences, therefore, in order to limit the number and length of the generated target entity sequences to a certain extent, conditions for random walks can be set during the process of generating the target entity sequence based on the Random Walk algorithm.
[0071] Exemplarily, when generating the target entity sequence based on the target knowledge graph, it can be achieved through at least one of the following methods:
[0072] Through the Random Walk algorithm, generate the target entity sequence based on the target entities with direct association relationships in the target knowledge graph. The direct association relationship can also be called a first-degree relationship. For example, for "Zhou Moumou" - starred in - "Movie A" - actor - "Zhao Moumou" - co-starred with - "Zheng Moumou" in the target knowledge graph, "Zhou Moumou" has a direct association relationship with "Movie A", "Movie A" has a direct association relationship with "Zhao Moumou", and "Zhao Moumou" also has a direct association relationship with "Zheng Moumou". Correspondingly, the target entity sequence composed of this group of target entities with direct association relationships is "Zhou Moumou" - "Movie A" - "Zhao Moumou" - "Zheng Moumou".
[0073] Through the Random Walk algorithm, generate the target entity sequence based on the target entities belonging to the same upper-level range in the target knowledge graph. That is, several target entities corresponding to the same upper-level word in the target knowledge graph can be used to form the target entity sequence. For example, for "Zhao Moumou" - female star - "Li Moumou" - singer - "Zhou Moumou" in the target knowledge graph, "Zhao Moumou" and "Li Moumou" belong to the same upper-level range "female star", and "Li Moumou" and "Zhou Moumou" belong to the same upper-level range "singer". Correspondingly, the target entity sequence composed of this group of target entities belonging to the same upper-level range is "Zhao Moumou" - "Li Moumou" - "Zhou Moumou".
[0074] It should be understood that in practical applications, when using the Random Walk algorithm to generate the target entity sequence based on the target knowledge graph, other conditions for restricting the association relationships between the target entities in the target entity sequence can also be set according to actual needs. The present application does not make any limitations on the restrictive conditions set for the Random Walk algorithm here.
[0075] Step 203: Based on the context prediction algorithm, determine the entity vectors corresponding to the target entities in the target knowledge graph according to the target entity sequence.
[0076] After the server generates several target entity sequences based on the target knowledge graph, it can use the context prediction algorithm to correspondingly determine the entity vector corresponding to the target entity in each generated target entity sequence. In this way, the entity vectors corresponding to each target entity in the target knowledge graph are traversed and determined in the above manner. Since the entity vector corresponding to the above target entity is determined by the server through the context prediction algorithm based on the target entity sequence including the target entity, the entity vector corresponding to the target entity can reflect its relevance to other target entities to a certain extent.
[0077] The following introduces the specific implementation method for determining the entity vector corresponding to the target entity.
[0078] The server can first perform one-hot encoding on each target entity in the target knowledge graph to obtain the basic vector corresponding to each target entity; then, train the skip-gram model based on the basic vector corresponding to the target entity in the target entity sequence. During the training of the skip-gram model, the embedding word vector of the target entity can be continuously adjusted; finally, the embedding of the target entity after the training of the skip-gram model is used as the entity vector corresponding to the target entity.
[0079] The skip-gram model is a model used to predict the output words adjacent to the input word within a preset window according to the input word.
[0080] In the technical solution provided in the embodiment of the present application, in order to convert the target entity in the target knowledge graph into a form that can be recognized by the machine, the server needs to first perform one-hot encoding on the target entity in the target knowledge graph to obtain the basic vector corresponding to each target entity. Suppose there are 10,000 target entities in the target knowledge graph. After performing one-hot encoding on these 10,000 target entities, the basic vector corresponding to each target entity should be a 10,000-dimensional vector. The value of each dimension in the basic vector can only be 0 or 1. Suppose the occurrence position of the target entity "Zhao Moumou" in the target knowledge graph is the third, then the basic vector corresponding to "Zhao Moumou" should be a 10,000-dimensional vector with the value of the third dimension being 1 and the values of other dimensions being 0.
[0081] Since the basic vectors obtained through one-hot encoding cannot reflect the similarity between target entities, and the embodiments of the present application need to obtain a dense vector (i.e., the entity vector corresponding to the target entity) that can reflect the relevance between target entities, therefore, the embodiments of the present application need to initialize the embedding corresponding to the target entity, and then, during the process of training the skip-gram model using the basic vectors of the target entities in the target entity sequence, continuously update the embedding corresponding to the target entity, that is, adjust the weights in the vector. After completing the training of the skip-gram model, the entity vector that can reflect the relevance between target entities can be obtained accordingly.
[0082] Figure 4 is a schematic diagram of training a skip-gram model. As Figure 4 shown, during the process of predicting, through the skip-gram model, the target entity adjacent to a certain target entity within a preset window in the target entity sequence, the basic vector corresponding to the target entity will be input into the skip-gram model first. Through the hidden layer in the skip-gram model, the target entity can be mapped to the corresponding embedding, and then, through the output layer in the skip-gram model, the probabilities of all target entities in the target knowledge graph being adjacent to the target entity are output. Among them, the vector obtained after the basic vector of the target entity is processed by the hidden layer in the skip-gram model is the entity vector corresponding to the target entity required in the embodiments of the present application.
[0083] Step 204: Determine a first similarity between the interest tags on the target network platform according to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest tags on the target network platform.
[0084] After the server calculates the entity vectors corresponding to the respective target vectors in the target knowledge graph, it can combine each pair of target entities in the target knowledge graph and calculate the similarity between the entity vectors corresponding to each two target entities. Furthermore, the server can obtain the mapping relationship between the target entities and the interest tags on the target network platform, convert the similarity between the entity vectors corresponding to the target entities into the similarity between the interest tags on the target network platform, and record the similarity between the interest tags determined in this way as the first similarity between the interest tags.
[0085] It should be understood that the mapping relationship between the above-mentioned target entity and the interest tags on the target network platform can be determined according to the attribute information such as the entity name of the target entity. This mapping relationship can be determined temporarily when determining the first similarity between the interest tags, or can be determined in advance. This application does not make any limitation on the determination method and determination timing of the mapping relationship between the target entity and the interest tags.
[0086] It should be understood that in practical applications, if two target entities are mapped to the same interest tag, then when determining the first similarity between the interest tags, the similarity between the entity vectors corresponding to these two target entities respectively can be not considered. In other words, the first similarity between the above-mentioned interest tags is substantially determined based on the similarity between the entity vectors corresponding to the target entities mapped to different interest tags.
[0087] The following introduces the specific implementation method for determining the first similarity between the interest tags on the above-mentioned target network platform.
[0088] The server can first determine an entity similarity matrix according to the entity vectors corresponding to each target entity in the target knowledge graph. Each element in this entity similarity matrix is used to represent the similarity between the target entity corresponding to the row where the element is located and the target entity corresponding to the column where the element is located. Furthermore, the server can convert the above entity similarity matrix into a first tag similarity matrix according to the mapping relationship between each target entity and each interest tag on the target network platform. Each element in this first tag similarity matrix is used to represent the similarity between the interest tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located.
[0089] Specifically, after the server determines the entity vectors corresponding to each target entity in the target knowledge graph, it can combine each pair of target entities in the target knowledge graph to obtain a number of target entity pairs; then, for each target entity pair, calculate the cosine similarity between the entity vectors corresponding to the two target entities respectively as the similarity corresponding to this target entity pair; furthermore, construct an entity similarity matrix based on the similarities corresponding to each target entity pair respectively. This entity similarity matrix takes each target entity as both the row and the column. The element located in the i-th row and the j-th column in the entity similarity matrix is actually the cosine similarity between the entity vector of the target entity corresponding to the i-th row and the entity vector of the target entity corresponding to the j-th column.
[0090] It should be understood that in practical applications, in addition to using the cosine similarity between entity vectors to construct the entity similarity matrix, the server can also use the similarity between entity vectors calculated based on other algorithms to construct this entity similarity matrix. This application does not make any limitation on the similarity algorithm used when calculating the similarity between entity vectors.
[0091] After the server constructs the entity similarity matrix, it can convert the entity similarity matrix into a first label similarity matrix used to represent the similarity between interest labels according to the mapping relationship between each target entity and each interest label on the target network platform. It should be understood that when the server converts the first label similarity matrix, for the similarity between the entity vectors corresponding to two target entities mapped to the same interest label, this similarity can be discarded and not included in the first label similarity matrix.
[0092] It should be noted that in practical applications, in order to more accurately expand the user interest portrait on the target network platform based on the similarity between the interest labels on the target network platform, the method provided in the embodiments of the present application can not only determine the similarity between the interest labels on the target network platform from the dimension of the knowledge graph, but also determine the similarity between the interest labels on the target network platform from the existing user interest portrait on the target network platform.
[0093] That is, the server can determine the second similarity between the interest labels on the target network platform according to the user interest portrait of the users on the target network platform. Specifically, the server can obtain the user interest portraits of some or all users on the target network platform, and then analyze the similarity between the interest labels based on the obtained user interest portraits. For example, assuming that the user interest portrait of user A includes interest label 1, interest label 2, interest label 3, and interest label 4, and the user interest portrait of user B includes interest label 2, interest label 3, interest label 4, and interest label 5, then when the server analyzes the user interest portraits, it can consider that interest label 1 and interest label 5 have a certain similarity; based on the above basic idea, the server can determine the similarity between the interest labels on the target network platform according to the existing user interest portraits on the target network platform, and the similarity between the interest labels determined in this way can be recorded as the second similarity between the interest labels.
[0094] Next, the specific implementation method for determining the second similarity between the interest labels on the above target network platform will be introduced.
[0095] The server can first construct a basic user profile matrix based on the user interest profiles on the target network platform. Each element in the basic user profile matrix is used to represent the degree of interest of the user corresponding to the row where the element is located in the interest tag corresponding to the column where the element is located. Furthermore, the server can train a Sparse Linear Model (SLIM) based on the basic user profile matrix, and use the trained SLIM as the second tag similarity matrix. Each element in the second tag similarity matrix is used to represent the similarity between the tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located.
[0096] Specifically, the server can construct a basic user profile matrix R(user-tag) according to the user interest profiles of all users on the target network platform. The rows in the basic user profile matrix represent users (user), and the columns represent interest tags (tag). Each element in the basic user profile matrix represents the degree of interest of the user corresponding to the row where the element is located in the interest tag corresponding to the column where the element is located. In other words, a row of elements in the basic user profile matrix can represent the degree of interest of a user in each interest tag on the target network platform, and a column of elements in the basic profile matrix can represent the degree of interest of all users on the target network platform in an interest tag.
[0097] Figure 5 FIG. is a schematic diagram for training SLIM based on the basic user profile matrix. Its basic principle is to train a matrix W such that the product of the basic user profile matrix R and the matrix W is still approximately equal to the basic user profile matrix R. To avoid obtaining a trivial and useless solution with diagonal elements being 1 and other elements being 0 during training, the diagonal elements of the matrix W need to be kept as 0 during training. The matrix W obtained through the above method is actually the second tag similarity matrix required by this application. Each element in the second tag similarity matrix can represent the similarity between the interest tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located. For example, the element located in the i-th row and j-th column of the second tag similarity matrix can represent the similarity between the i-th interest tag and the j-th interest tag.
[0098] It should be noted that in practical applications, in addition to obtaining the above-mentioned second tag similarity matrix based on the user interest profiles on the target network platform by training the SLIM model, other methods can also be used to obtain the above-mentioned second tag similarity matrix. For example, matrix factorization, embedding, etc. can be used to generate the second tag similarity matrix based on the user interest profiles on the target network platform. This application does not make any limitations on the method of generating the second tag similarity matrix.
[0099] Step 205: Expand the user interest profile on the target network platform based on the first similarity.
[0100] After the server determines the first similarity between each interest label on the target network platform, it can expand the user interest profile on the target network platform based on this first similarity between interest labels. The basic principle is to expand in the user interest profile interest labels that are relatively similar to the original interest labels therein based on the first similarity between interest labels.
[0101] When the first similarity between each interest label on the target network platform is represented by the above-mentioned first label similarity matrix, the server can expand the user interest profile on the target network platform in the following way:
[0102] Determine the expanded user profile matrix according to the first label similarity matrix and the basic user profile matrix; the basic user profile matrix is constructed based on the existing user interest profiles on the target network platform, and each element in the basic user profile matrix and the expanded user profile matrix is used to represent the degree of interest of the user corresponding to the row where the element is located in the interest label corresponding to the column where the element is located.
[0103] Specifically, the server can multiply the basic user profile matrix R by the first label similarity matrix P to obtain the expanded user profile matrix R'. Among them, the basic user profile matrix R can be constructed by the server according to the user interest profiles of all users on the target network platform. The rows in the basic user profile matrix R correspond to users, and the columns correspond to interest labels. The element R in the basic user profile matrix R ij is used to represent the degree of interest of the i-th user in the j-th interest label. The first label similarity matrix P is determined by the server based on the similarity between the entity vectors corresponding to each target entity in the target knowledge graph. The rows and columns in the first label similarity matrix P both correspond to interest labels. The element P in the first label similarity matrix P ij is used to represent the similarity between the i-th interest label and the j-th interest label. The expanded user profile matrix R' obtained by multiplying the basic user profile matrix R by the first label similarity matrix P can reflect the interest labels expanded in the user interest profile. The expanded user profile matrix R' is similar to the basic user profile matrix R, where the rows correspond to users and the columns correspond to interest labels. The element R' in the expanded user profile matrix R' ij is used to represent the degree of interest of the i-th user in the j-th interest label after the user profile expansion process.
[0104] If the server has previously determined the second similarity between each interest tag on the target network platform based on the user interest portrait on the target network platform, the server can then expand the user interest portrait on the target network platform based on the first similarity and the second similarity between each interest tag.
[0105] Specifically, the server can fuse the first similarity and the second similarity between each interest tag on the target network platform to obtain the target similarity between each interest tag that matches the user interests and hobbies on the target network platform and the relationships between entities in the target knowledge graph. Furthermore, the server can expand the user interest portrait on the target network platform according to the target similarity between each interest tag.
[0106] In the case where the first similarity between each interest tag on the target network platform is represented by the above-mentioned first tag similarity matrix, and the second similarity between each interest tag is represented by the above-mentioned second tag similarity matrix, the server can expand the user interest portrait on the target network platform in the following manner:
[0107] Perform a weighted processing on the first tag similarity matrix and the second tag similarity matrix to obtain a target tag similarity matrix; furthermore, determine an extended user portrait matrix according to the target tag similarity matrix and the basic user portrait matrix.
[0108] Specifically, the server can first perform a weighted summation processing on the first tag similarity matrix P and the second tag similarity matrix W according to the pre-set weights; for example, assume that the server sets a weight value x1 for the first tag similarity matrix P and a weight value x2 for the second tag similarity matrix W, then the server can calculate the target tag similarity matrix Q through the following formula:
[0109] Q = P * x1 + W * x2
[0110] It should be understood that the weight values x1 and x2 here are determined by the server according to the degree of attention to the first tag similarity matrix and the second tag similarity matrix. If more attention is paid to the similarity between interest tags determined based on the relationships between entities in the target knowledge graph when expanding the user interest portrait, x1 can be set to be greater than x2. If more attention is paid to the similarity between interest tags determined based on the user interest portrait on the target network platform when expanding the user interest portrait, x2 can be set to be greater than x1. This application does not make specific limitations on the set weight values x1 and x2 here.
[0111] Furthermore, the server can multiply the basic user portrait matrix R by the target tag similarity matrix Q to obtain the extended user portrait matrix R'.
[0112] It should be noted that in practical applications, after the server determines the extended user portrait matrix through any of the above methods, it is necessary to expand the existing user interest portraits on the target network platform based on this extended user portrait matrix. Exemplarily, the server can expand the user interest portrait of the i-th user on the target network platform based on the elements in the i-th row of the extended user portrait matrix. In order to enable the server to more conveniently expand the existing user interest portraits based on this extended user portrait matrix, after generating the extended user portrait matrix, the server can perform optimization processing on the generated extended user portrait matrix through the following implementation methods.
[0113] In a possible implementation method, the server can, for the elements in the same position in the extended user portrait matrix and the basic user portrait matrix, determine whether the element in the basic user portrait matrix at this position is greater than the first preset threshold. If so, the element in the extended user portrait matrix at this position can be set to 0.
[0114] Specifically, if it is determined according to the basic user portrait matrix that user A is interested in interest tag 1, and it is also determined according to the extended user portrait matrix that user A is interested in interest tag 1, then the server needs to set the element in the extended user portrait matrix that represents the degree of user A's interest in interest tag 1 to 0, so as to avoid adding the original interest tag to the user portrait matrix again when subsequently expanding the user interest tags based on the extended user portrait matrix.
[0115] Generally, in the basic user portrait matrix, if an element is not 0, then the interest tag corresponding to the column where this element is located should be in the user interest portrait of the user corresponding to the row where this element is located. In this case, the server needs to set the above first preset threshold to 0. Of course, in some cases, in the basic user portrait matrix, only when an element is greater than the preset value a, the interest tag corresponding to the column where this element is located will be in the user interest portrait of the user corresponding to the row where this element is located. In this case, the server needs to set the above first preset threshold to a. This application does not make specific limitations on this first preset threshold here.
[0116] In another possible implementation method, the server can, for each element in the extended user portrait matrix, determine whether this element is less than or equal to the second preset threshold. If so, set this element in the extended user portrait matrix to 0.
[0117] Specifically, in order to ensure the sparsity of the extended user portrait matrix, the server needs to set a second preset threshold between 0 and 1. For the elements in the extended user portrait matrix that are less than or equal to this second preset threshold, the server needs to set them to 0 accordingly.
[0118] It should be noted that if the above-mentioned second preset threshold is set too low, the extended user portrait matrix will be too dense, which is not conducive to subsequent storage and calculation; if the above-mentioned second preset threshold is set too high, the number of extended interest tags will be too small, and the effect of expanding the user interest portrait will not be obvious. Therefore, in practical applications, the above-mentioned second preset threshold can be determined according to the results of A / B testing (AB test) and the engineering requirements for implementation speed. This application does not make specific limitations on this second preset threshold here.
[0119] After processing the extended user portrait matrix through the above two implementation methods, the interest tags corresponding to the columns where the non-zero items are located in each row of the extended user portrait matrix are actually the interest tags extended for the user corresponding to that row.
[0120] It should be understood that in practical applications, in addition to optimizing the extended user portrait matrix through the above two implementation methods, other methods can also be used to optimize the extended user portrait matrix. This application does not make any limitations on the methods for optimizing the extended user portrait matrix here.
[0121] The above user interest portrait expansion method is based on a knowledge graph covering a large number of entities and relationships between entities, determines the similarity between entities in the knowledge graph, and converts the similarity between entities into the similarity between interest tags according to the mapping relationship between entities and interest tags. Furthermore, the user interest portrait is expanded based on the similarity between interest tags. In this way, the user interest portrait can be expanded quickly and accurately, and thus it is beneficial for the network platform to provide richer personalized services based on the expanded user interest portrait.
[0122] To facilitate further understanding of the user interest portrait expansion method provided by the embodiments of this application, still taking the server as the execution subject as an example, in combination with Figure 6 the shown flowchart, an overall exemplary introduction to the user interest portrait expansion method provided by the embodiments of this application is given.
[0123] As Figure 6 shown, the user interest portrait expansion method provided by the embodiments of this application is mainly implemented through four steps, namely Step 1 - generating a similarity matrix P using the knowledge graph (i.e., the first label similarity matrix in the above text), Step 2 - generating a similarity matrix W using the user interest portrait (i.e., the second label similarity matrix in the above text), Step 3 - fusing the similarity matrices, and Step 4 - generating an extended user interest portrait. The following is an introduction to these four steps respectively.
[0124] Step 1 - generating a similarity matrix P using the knowledge graph:
[0125] Extract valid information from the basic knowledge graph to form a target knowledge graph: The basic knowledge graph consists of a number of triples with associated relationships (head entity / entity relationship / tail entity), where each entity contains a set of attribute information. Typical attribute information includes entity type, entity name, entity popularity, etc. This application mainly extracts target entities with entity types of person, movie, TV drama, variety show, etc. and high entity popularity from the basic knowledge graph to form a target knowledge graph.
[0126] Random walk based on the target knowledge graph: Randomly walk according to the relationships between target entities in the target knowledge graph to form a number of target entity sequences. Exemplarily, a target entity sequence can be generated by randomly walking based on the first-degree relationship between target entities. For example, "Zhou Moumou" - "acted in" - "Movie A" - "actor" - "Zhao Moumou" - "partnered with" - "Zheng Moumou", where "Zhou Moumou", "Movie A", "Zhao Moumou", and "Zheng Moumou" can form a target entity sequence; a target entity sequence can also be generated by randomly walking based on the hypernyms of target entities. For example, "Zhao Moumou" - "female star" - "Li Moumou" - "singer" - "Zhou Moumou", where "Zhao Moumou", "Li Moumou", and "Zhou Moumou" can form a target entity sequence.
[0127] Train the entity vector embedding of the target entity: Use the skip-gram algorithm to train the embedding of each target entity in the target knowledge graph on the target entity sequences generated by random walk.
[0128] Calculate the similarity between the entity vectors of the target entities: Calculate the cosine similarity between all pairs of target entities to obtain a similarity matrix from target entity to target entity (i.e., entity similarity matrix).
[0129] Generate the interest label similarity matrix P: Map the target entities to interest labels according to attribute information such as the entity names of the target entities to obtain the similarity matrix P from interest label to interest label.
[0130] Step 2 - Generate the similarity matrix W using the user interest portrait:
[0131] Input the user interest portrait: Construct a basic user portrait matrix R (user-tag) according to the user interest portrait on the target network platform, where the rows represent users and the columns represent interest labels. The element values in the basic user portrait matrix R represent the degree of interest of users in the interest labels.
[0132] Train the sparse linear model: Use the basic user portrait matrix R to train the label similarity matrix W. After multiplying the basic user portrait matrix R by the label similarity matrix W, it is still approximately equal to the basic user portrait matrix R; the element W in the label similarity matrix Wij It represents the similarity between the i-th interest tag and the j-th interest tag. When training the tag similarity matrix W, the diagonal elements of the tag similarity matrix W need to be kept as 0, aiming to avoid trivial solutions (i.e., a matrix with diagonal elements being 0 and other elements being 0) during training.
[0133] Step 3 - Fuse the similarity matrix:
[0134] Set the weights [x1, x2] to obtain the final target tag similarity matrix Q = P * x1 + W * x2.
[0135] Step 4 - Generate the extended user interest profile:
[0136] Multiply the basic user profile matrix R by the target tag similarity matrix Q to obtain the extended user profile matrix R'.
[0137] For all elements in the extended user profile matrix R', where U represents the number of users on the target network platform and T represents the number of interest tags on the target network platform.
[0138] When R[i, j] > 0, set R'[i, j] = 0; to avoid expanding the user's original interest tags based on this extended user profile matrix.
[0139] When R'[i, j] threshold, set R'[i, j] = 0; threshold is a number between 0 and 1. The setting of threshold is to ensure the sparsity of the extended user profile matrix R'. If the threshold is set too low, the extended user profile matrix R' will be too dense, which is not conducive to subsequent storage and calculation; if the threshold is set too high, the number of extended interest tags will be too small, and the effect of expanding the user interest profile will not be obvious. Therefore, this threshold can be determined according to the AB test results and the engineering requirements for implementation speed.
[0140] The non-zero terms in a row of the extended user profile matrix R' are the extended interest tags of the corresponding user in that row.
[0141] For the user interest profile expansion method described above, this application also provides a corresponding user interest profile expansion device to enable the application and implementation of the above user interest profile expansion method in practice.
[0142] See Figure 7 , Figure 7 is the above Figure 2Schematic diagram of a user interest profile expansion device 700 corresponding to the user interest profile expansion method shown. As Figure 7 shown, the user interest profile expansion device 700 includes:
[0143] A knowledge graph acquisition module 701, configured to acquire a target knowledge graph; the target knowledge graph is used to represent the association relationship between target entities, and the target entities are entities related to the target network platform;
[0144] An entity sequence generation module 702, configured to generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having an association relationship in the target knowledge graph;
[0145] An entity vector determination module 703, configured to determine an entity vector corresponding to the target entity in the target knowledge graph based on a context prediction algorithm according to the target entity sequence;
[0146] A first label similarity determination module 704, configured to determine a first similarity between interest labels on the target network platform according to the similarity between entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest labels on the target network platform;
[0147] A user profile expansion module 705, configured to expand the user interest profile on the target network platform based on the first similarity.
[0148] Optionally, based on the user interest profile expansion device shown in Figure 7 refer to Figure 8 , Figure 8 Schematic diagram of another user interest profile expansion device 800 provided by an embodiment of the present application. As Figure 8 shown, the device further includes:
[0149] A second label similarity determination module 801, configured to determine a second similarity between interest labels on the target network platform according to the user interest profile of the user on the target network platform;
[0150] Then the user profile expansion module 705 is specifically configured to:
[0151] Expand the user interest profile on the target network platform based on the first similarity and the second similarity.
[0152] Optionally, based on the user interest profile expansion device shown in Figure 7 the first label similarity determination module 704 is specifically configured to:
[0153] Determine an entity similarity matrix based on the entity vectors corresponding to each of the target entities in the target knowledge graph; each element in the entity similarity matrix is used to represent the similarity between the entity corresponding to the row where the element is located and the entity corresponding to the column where the element is located;
[0154] According to the mapping relationship between each of the target entities and each interest tag on the target network platform, convert the entity similarity matrix into a first tag similarity matrix; each element in the first tag similarity matrix is used to represent the similarity between the interest tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located;
[0155] Then the user profile expansion module 705 is specifically configured to:
[0156] Determine an extended user profile matrix according to the first tag similarity matrix and the basic user profile matrix; the basic user profile matrix is constructed according to the user interest profiles of users on the target network platform; each element in the basic user profile matrix and the extended user profile matrix is used to represent the degree of interest of the user corresponding to the row where the element is located in the interest tag corresponding to the column where the element is located.
[0157] Optionally, based on the Figure 8 shown user interest profile expansion device, the second tag similarity determination module 801 is specifically configured to:
[0158] Train a sparse linear regression model based on the basic user profile matrix, and use the sparse linear regression model as the second tag similarity matrix; each element in the second tag similarity matrix is used to represent the similarity between the interest tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located;
[0159] Then the user profile expansion module 705 is specifically configured to:
[0160] Perform weighted processing on the first tag similarity matrix and the second tag similarity matrix to obtain a target tag similarity matrix;
[0161] Determine the extended user profile matrix according to the target tag similarity matrix and the basic user profile matrix.
[0162] Optionally, based on the Figure 7 or Figure 8 shown user interest profile expansion device, refer to Figure 9 , Figure 9 which is a schematic structural diagram of another user interest profile expansion device 900 provided by an embodiment of the present application, as shown in Figure 9 shown, this device further includes:
[0163] The first matrix correction module 901 is configured to, for elements at the same position in the extended user portrait matrix and the basic user portrait matrix, determine whether the element at the position in the basic user portrait matrix is greater than a first preset threshold. If so, set the element at the position in the extended user portrait matrix to 0.
[0164] Optionally, based on Figure 7 or Figure 8 the user interest portrait expansion device shown, refer to Figure 10 , Figure 10 which is a schematic structural diagram of another user interest portrait expansion device 900 provided by an embodiment of the present application. As shown in Figure 10 , the device further includes:
[0165] The second matrix correction module 1001 is configured to, for each element in the extended user portrait matrix, determine whether the element is less than or equal to a second preset threshold. If so, set the element in the extended user portrait matrix to 0.
[0166] Optionally, based on Figure 7 the user interest portrait expansion device shown, the entity vector determination module 703 is specifically configured to:
[0167] Perform one-hot encoding on each of the target entities in the target knowledge graph to obtain a basic vector corresponding to each of the target entities;
[0168] Train a skip-gram model based on the basic vectors corresponding to the target entities in the target entity sequence, and adjust the embedding word vector embedding of the target entity during the training process;
[0169] Use the embedding of the target entity after the training of the skip-gram model is completed as the entity vector corresponding to the target entity.
[0170] Optionally, based on Figure 7 the user interest portrait expansion device shown, the entity sequence generation module 702 is specifically configured to:
[0171] Generate the target entity sequence based on the target entities having a direct association relationship in the target knowledge graph through a random walk algorithm;
[0172] And / or generate the target entity sequence based on the target entities belonging to the same upper range in the target knowledge graph through the random walk algorithm.
[0173] Optionally, based on Figure 7Based on the user interest profile expansion device shown above, the knowledge graph acquisition module 701 is specifically configured to:
[0174] Select entities that meet preset conditions from the basic knowledge graph as the target entities; the preset conditions include at least one of the following: the entity type is a preset type, and the entity popularity exceeds a preset popularity threshold;
[0175] Determine the target knowledge graph according to the association relationship of the target entities in the basic knowledge graph.
[0176] Based on the knowledge graph covering a large number of entities and relationships between entities, the above user interest profile expansion device determines the similarity between entities in the knowledge graph, and according to the mapping relationship between entities and interest tags, converts the similarity between entities into the similarity between interest tags, and then expands the user interest profile based on the similarity between interest tags. In this way, the user interest profile can be quickly and accurately expanded, and furthermore, it is beneficial for the network platform to provide richer personalized services for users based on the expanded user interest profile.
[0177] An embodiment of the present application also provides a device for expanding a user interest profile. This device may specifically be a terminal device or a server. The terminal device and the server provided in the embodiment of the present application will be introduced from the perspective of hardware implementation below.
[0178] See Figure 11 , Figure 11 is a schematic structural diagram of the terminal device provided in the embodiment of the present application. As Figure 11 shown, for the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. This terminal may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (full English name: Personal Digital Assistant, English abbreviation: PDA), a point of sales (full English name: Point of Sales, English abbreviation: POS), an in-vehicle computer, etc. Taking the terminal as a computer as an example:
[0179] Figure 11 Shown is a block diagram of a part of the structure of a computer related to the terminal provided in the embodiment of the present application. Refer to Figure 11, the computer includes components such as a Radio Frequency (RF) circuit 1110, a memory 1120, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a wireless fidelity (WiFi) module 1170, a processor 1180, and a power supply 1190. Those skilled in the art can understand that Figure 11 the computer structure shown in
[0180] does not limit the computer, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functional applications and data processing of the computer by running the software programs and modules stored in the memory 1120. The memory 1120 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer (such as audio data, a phone book, etc.). In addition, the memory 1120 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0181] The processor 1180 is the control center of the computer, connects various parts of the entire computer through various interfaces and lines, and executes various functions of the computer and processes data by running or executing the software programs and / or modules stored in the memory 1120, and calling the data stored in the memory 1120. Optionally, the processor 1180 can include one or more processing units; preferably, the processor 1180 can integrate an application processor and a modulation / demodulation processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modulation / demodulation processor mainly processes wireless communication. It can be understood that the above modulation / demodulation processor may not be integrated into the processor 1180.
[0182] In the embodiments of the present application, the processor 1180 included in the terminal further has the following functions:
[0183] Obtain a target knowledge graph; the target knowledge graph is used to represent the association relationship between target entities, and the target entities are entities related to the target network platform;
[0184] Generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having an association relationship in the target knowledge graph;
[0185] Based on the context prediction algorithm, according to the target entity sequence, determine the entity vector corresponding to the target entity in the target knowledge graph;
[0186] According to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest tags on the target network platform, determine the first similarity between the interest tags on the target network platform;
[0187] Based on the first similarity, expand the user interest profile on the target network platform.
[0188] Optionally, the processor 1180 is further configured to execute the steps of any implementation manner of the user interest profile expansion method provided in the embodiments of the present application.
[0189] See Figure 12 , Figure 12 FIG. is a schematic structural diagram of a server 1200 provided in an embodiment of the present application. The server 1200 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1222 (for example, one or more processors) and a memory 1232, and one or more storage media 1230 (for example, one or more mass storage devices) for storing application programs 1242 or data 1244. Among them, the memory 1232 and the storage media 1230 may be transient storage or persistent storage. The program stored in the storage media 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1222 may be configured to communicate with the storage media 1230 and execute a series of instruction operations in the storage media 1230 on the server 1200.
[0190] The server 1200 may further include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0191] The steps performed by the server in the above embodiments may be based on the Figure 12 server structure shown.
[0192] Among them, the CPU 1222 is used to execute the following steps:
[0193] Obtain a target knowledge graph; the target knowledge graph is used to represent the association relationships between target entities, and the target entities are entities related to a target network platform;
[0194] Generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having association relationships in the target knowledge graph;
[0195] Determine the entity vectors corresponding to the target entities in the target knowledge graph based on the context prediction algorithm according to the target entity sequence;
[0196] Determine a first similarity between the interest tags on the target network platform according to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest tags on the target network platform;
[0197] Expand the user interest profile on the target network platform based on the first similarity.
[0198] Optionally, the CPU 1222 can also be used to execute the steps of any implementation manner of the user interest profile expansion method provided in the embodiments of the present application.
[0199] The embodiments of the present application further provide a computer-readable storage medium for storing a computer program, and the computer program is used to execute any implementation manner of the user interest profile expansion method described in the foregoing embodiments.
[0200] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any implementation manner of the user interest profile expansion method described in the foregoing embodiments.
[0201] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0202] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0203] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0205] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. And the foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs and other various media that can store computer programs.
[0206] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0207] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for expanding a user interest profile, characterized in that, The method includes: Obtain a target knowledge graph; the target knowledge graph is used to represent the association relationships between target entities, and the target entities are entities related to a target network platform; Generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having association relationships in the target knowledge graph; Based on a context prediction algorithm, determine entity vectors corresponding to the target entities in the target knowledge graph according to the target entity sequence; According to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and interest tags on the target network platform, determine a first similarity between the interest tags on the target network platform; Based on the first similarity, expand the user interest profile on the target network platform; The step of determining the first similarity between the interest tags on the target network platform according to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and interest tags on the target network platform includes: Determine an entity similarity matrix according to the entity vectors corresponding to the respective target entities in the target knowledge graph; each element in the entity similarity matrix is used to represent the similarity between the target entity corresponding to the row where the element is located and the target entity corresponding to the column where the element is located; According to the mapping relationship between each target entity and each interest tag on the target network platform, convert the entity similarity matrix into a first tag similarity matrix; each element in the first tag similarity matrix is used to represent the similarity between the interest tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located; Then the step of expanding the user interest profile on the target network platform based on the first similarity includes: Determine an extended user profile matrix according to the first tag similarity matrix and a basic user profile matrix; the basic user profile matrix is constructed according to the user interest profiles of users on the target network platform; each element in the basic user profile matrix and the extended user profile matrix is used to represent the degree of interest of the user corresponding to the row where the element is located in the interest tag corresponding to the column where the element is located.
2. The method according to claim 1, wherein The method further includes: Determine a second similarity between the interest tags on the target network platform according to the user interest profiles of users on the target network platform; Then the step of expanding the user interest profile on the target network platform based on the first similarity includes: Expand the user interest profile on the target network platform based on the first similarity and the second similarity.
3. The method according to claim 1, characterized in that, The method further includes: Train a sparse linear model based on the basic user profile matrix, and use the sparse linear model as a second tag similarity matrix; each element in the second tag similarity matrix is used to represent the similarity between the interest tag corresponding to the row where the element is located and the interest tag corresponding to the column where the element is located; Then, determining the extended user profile matrix according to the first label similarity matrix and the basic user profile matrix includes: Performing weighted processing on the first label similarity matrix and the second label similarity matrix to obtain a target label similarity matrix; Determining the extended user profile matrix according to the target label similarity matrix and the basic user profile matrix.
4. The method according to claim 1 or 3, characterized in that The method further includes: For elements at the same position in the extended user profile matrix and the basic user profile matrix, determining whether the element at the position in the basic user profile matrix is greater than a first preset threshold. If so, setting the element at the position in the extended user profile matrix to 0.
5. The method according to claim 1 or 3, characterized in that, The method further includes: For each element in the extended user profile matrix, determining whether the element is less than or equal to a second preset threshold. If so, setting the element in the extended user profile matrix to 0.
6. The method according to claim 1, wherein The determining, by the context prediction algorithm, the entity vector corresponding to the target entity in the target knowledge graph according to the target entity sequence includes: Performing one-hot encoding on each target entity in the target knowledge graph to obtain a basic vector corresponding to each target entity; Training a skip-gram model based on the basic vector corresponding to the target entity in the target entity sequence, and adjusting the embedding of the target entity during the training process; Taking the embedding of the target entity after the training of the skip-gram model is completed as the entity vector corresponding to the target entity.
7. The method according to claim 1, wherein The generating the target entity sequence based on the target knowledge graph includes at least one of the following: Generating the target entity sequence based on the target entities having a direct association relationship in the target knowledge graph through a random walk algorithm; Generating the target entity sequence based on the target entities belonging to the same upper range in the target knowledge graph through a random walk algorithm.
8. The method according to claim 1, characterized in that The obtaining the target knowledge graph includes: Selecting entities that meet preset conditions from the basic knowledge graph as the target entities; the preset conditions include at least one of the following: the entity type is a preset type, and the entity popularity exceeds a preset popularity threshold; Determining the target knowledge graph according to the association relationship of the target entities in the basic knowledge graph.
9. A user interest profile expansion device, characterized in that, The device includes: A knowledge graph acquisition module, configured to acquire a target knowledge graph; the target knowledge graph is used to represent the association relationship between target entities, and the target entities are entities related to a target network platform; An entity sequence generation module, configured to generate a target entity sequence based on the target knowledge graph; the target entity sequence is a sequence composed of multiple target entities having an association relationship in the target knowledge graph; An entity vector determination module, configured to determine the entity vector corresponding to the target entity in the target knowledge graph according to the target entity sequence based on the context prediction algorithm; The first label similarity determination module is used to determine the first similarity between the interest labels on the target network platform according to the similarity between the entity vectors corresponding to the target entities in the target knowledge graph and the mapping relationship between the target entities and the interest labels on the target network platform; The user portrait expansion module is used to expand the user interest portrait on the target network platform based on the first similarity; Specifically, the first label similarity determination module is used to: Determine an entity similarity matrix according to the entity vectors corresponding to each of the target entities in the target knowledge graph; each element in the entity similarity matrix is used to represent the similarity between the entity corresponding to the row where the element is located and the entity corresponding to the column where the element is located; According to the mapping relationship between each of the target entities and each interest label on the target network platform, convert the entity similarity matrix into a first label similarity matrix; each element in the first label similarity matrix is used to represent the similarity between the interest label corresponding to the row where the element is located and the interest label corresponding to the column where the element is located; Specifically, the user portrait expansion module is used to: Determine an extended user portrait matrix according to the first label similarity matrix and the basic user portrait matrix; the basic user portrait matrix is constructed according to the user interest portraits of the users on the target network platform; each element in the basic user portrait matrix and the extended user portrait matrix is used to represent the degree of interest of the user corresponding to the row where the element is located in the interest label corresponding to the column where the element is located.
10. The device according to claim 9, characterized in that, The device further includes: The second label similarity determination module is used to determine the second similarity between the interest labels on the target network platform according to the user interest portraits of the users on the target network platform; Specifically, the user portrait expansion module is used to: Expand the user interest portrait on the target network platform based on the first similarity and the second similarity.
11. The device according to claim 10, characterized in that, Specifically, the second label similarity determination module is used to: Train a sparse linear regression model based on the basic user portrait matrix, and use the sparse linear regression model as the second label similarity matrix; each element in the second label similarity matrix is used to represent the similarity between the interest label corresponding to the row where the element is located and the interest label corresponding to the column where the element is located; Specifically, the user portrait expansion module is used to: Perform weighted processing on the first label similarity matrix and the second label similarity matrix to obtain a target label similarity matrix; Determine the extended user portrait matrix according to the target label similarity matrix and the basic user portrait matrix.
12. The device according to claim 9 or 11, characterized in that, The device further includes: The first matrix correction module is used to, for the elements in the same position in the extended user portrait matrix and the basic user portrait matrix, determine whether the element in the basic user portrait matrix at the position is greater than a first preset threshold. If so, set the element in the extended user portrait matrix at the position to 0.
13. The device according to claim 9 or 11, characterized in that, The device further includes: The second matrix correction module is used to determine, for each element in the extended user profile matrix, whether the element is less than or equal to a second preset threshold. If so, the element in the extended user profile matrix is set to 0.
14. The device according to claim 9, characterized in that, The entity vector determination module is specifically configured to: perform one-hot encoding on each of the target entities in the target knowledge graph to obtain a respective basic vector corresponding to each of the target entities; train a skip-gram model based on the basic vectors corresponding to the target entities in the target entity sequence, and adjust the embedding word vector of the target entity during the training process; use the embedding of the target entity after the training of the skip-gram model is completed as the entity vector corresponding to the target entity.
15. The device according to claim 9, characterized in that, The entity sequence generation module is specifically configured to: generate the target entity sequence based on the target entities having a direct association relationship in the target knowledge graph through a random walk algorithm; and / or generate the target entity sequence based on the target entities belonging to the same upper range in the target knowledge graph through the random walk algorithm.
16. The device according to claim 9, characterized in that, The knowledge graph acquisition module is specifically configured to: select entities that meet preset conditions from the basic knowledge graph as the target entities; the preset conditions include at least one of the following: the entity type is a preset type, and the entity popularity exceeds a preset popularity threshold; determine the target knowledge graph according to the association relationship of the target entities in the basic knowledge graph.
17. An electronic device, characterized in that, The device includes a processor and a memory; the memory is used to store a computer program; the processor is used to execute the user interest profile extension method according to any one of claims 1 to 8 based on the computer program.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the user interest profile extension method according to any one of claims 1 to 8.
19. A computer program product, characterized in that, The computer program product includes instructions that, when running on a computer device, cause the computer device to execute the user interest profile extension method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Knowledge graph vector determination method and device, terminal equipment and medium
CN110309316A