Knowledge graph construction, information recommendation method and device, and computer equipment
By acquiring user-read text, linking and expanding entities, calculating similarity, and building an interest knowledge graph, the problem of low accuracy in user interest representation is solved, achieving more accurate user interest representation and recommendation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-10-21
- Publication Date
- 2026-05-22
AI Technical Summary
Existing technologies that characterize user interests through user profiling suffer from issues such as omission of interests and low accuracy.
By acquiring the reading text of target users within a preset time period, entity links and expansions are performed, the similarity between instances and concepts is calculated, and an interest knowledge graph is established, including an instance entity set, an expanded instance entity set, and an expanded concept entity set. The expanded entity weights are calculated using entity weights and similarity, and finally, the interest knowledge graph of the target users is established.
It improves the accuracy of user interest representation and ensures the precision of recommended information.
Smart Images

Figure CN116010611B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a knowledge graph construction, information recommendation method, apparatus, computer device, storage medium, and computer program product. Background Technology
[0002] With the development of artificial intelligence technology, knowledge graph technology has emerged. Knowledge graphs are a modern theory that combines theories and methods from applied mathematics, computer graphics, information visualization, and information science with bibliometric citation analysis and co-occurrence analysis. They utilize visualized graphs to vividly display the core structure, development history, cutting-edge fields, and overall knowledge architecture of a discipline, achieving multidisciplinary integration. Currently, when characterizing user interests, user profiles are typically used to represent user interest characteristics. However, using user profiles to represent user interest characteristics may result in omissions and low accuracy in representing user interests. Summary of the Invention
[0003] Therefore, it is necessary to provide a knowledge graph construction, information recommendation method, device, computer equipment, storage medium, and computer program product that can improve the accuracy of user interest representation in order to address the above-mentioned technical problems.
[0004] A knowledge graph construction method, the method comprising:
[0005] Obtain the reading text of the target user within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text;
[0006] Based on each instance entity in the instance entity set, we perform instance expansion to obtain an expanded instance entity set. Based on the expanded instance entity set and the instance entity set, we obtain the target instance entity set. Based on each target instance entity in the target instance entity set, we perform concept expansion to obtain an expanded concept entity set.
[0007] Obtain the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0008] Calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and use the weight of each instance entity and the instance similarity to calculate the weight of the extended instance entity, thus obtaining the weight of each extended instance entity.
[0009] Calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. Use the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0010] An interest knowledge graph corresponding to the target user identifier is constructed based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0011] A knowledge graph construction apparatus, the apparatus comprising:
[0012] The instance acquisition module is used to obtain the reading text of the target user identifier within a preset time period, and perform entity linking based on the reading text to obtain the instance entity set corresponding to the reading text.
[0013] The extension module is used to extend instances based on each instance entity in the instance entity set to obtain an extended instance entity set, obtain a target instance entity set based on the extended instance entity set and the instance entity set, and extend concepts based on each target instance entity in the target instance entity set to obtain an extended concept entity set.
[0014] The weight acquisition module is used to acquire the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0015] The instance weight calculation module is used to calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity. It uses the weight of each instance entity and the instance similarity to calculate the weight of the extended instance entity, and obtains the weight of each extended instance entity.
[0016] The concept weight calculation module is used to calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. It uses the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0017] The graph building module is used to build an interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0018] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:
[0019] Obtain the reading text of the target user within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text;
[0020] Based on each instance entity in the instance entity set, we perform instance expansion to obtain an expanded instance entity set. Based on the expanded instance entity set and the instance entity set, we obtain the target instance entity set. Based on each target instance entity in the target instance entity set, we perform concept expansion to obtain an expanded concept entity set.
[0021] Obtain the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0022] Calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and use the weight of each instance entity and the instance similarity to calculate the weight of the extended instance entity, thus obtaining the weight of each extended instance entity.
[0023] Calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. Use the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0024] An interest knowledge graph corresponding to the target user identifier is constructed based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0025] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0026] Obtain the reading text of the target user within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text;
[0027] Based on each instance entity in the instance entity set, we perform instance expansion to obtain an expanded instance entity set. Based on the expanded instance entity set and the instance entity set, we obtain the target instance entity set. Based on each target instance entity in the target instance entity set, we perform concept expansion to obtain an expanded concept entity set.
[0028] Obtain the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0029] Calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and use the weight of each instance entity and the instance similarity to calculate the weight of the extended instance entity, thus obtaining the weight of each extended instance entity.
[0030] Calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. Use the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0031] An interest knowledge graph corresponding to the target user identifier is constructed based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0032] A computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0033] Obtain the reading text of the target user within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text;
[0034] Based on each instance entity in the instance entity set, we perform instance expansion to obtain an expanded instance entity set. Based on the expanded instance entity set and the instance entity set, we obtain the target instance entity set. Based on each target instance entity in the target instance entity set, we perform concept expansion to obtain an expanded concept entity set.
[0035] Obtain the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0036] Calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and use the weight of each instance entity and the instance similarity to calculate the weight of the extended instance entity, thus obtaining the weight of each extended instance entity.
[0037] Calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. Use the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0038] An interest knowledge graph corresponding to the target user identifier is constructed based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0039] The aforementioned knowledge graph construction method, apparatus, computer equipment, storage medium, and computer program product acquire reading text of a target user identifier within a preset time period, perform entity linking based on the reading text to obtain a set of instance entities corresponding to the reading text; expand each instance entity in the instance entity set to obtain an expanded instance entity set; obtain a target instance entity set based on the expanded instance entity set and the instance entity set; and expand concepts based on each target instance entity in the target instance entity set to obtain an expanded concept entity set. Then, obtain the weights of each instance entity, and calculate the weights of each expanded instance entity using the instance entity weights and instance similarity. Finally, calculate the weights of each expanded concept entity using the expanded instance entity weights, the instance entity weights, and the concept similarity. Finally, use the target instance entity set, the expanded concept entity set, the instance entity weights, the expanded instance entity weights, and the expanded concept entity weights to construct an interest knowledge graph corresponding to the target user identifier, thereby improving the accuracy of the constructed interest knowledge graph in representing user interests.
[0040] An information recommendation method, the method comprising:
[0041] Receive information recommendation instructions, which carry user identifiers and query statements;
[0042] Extract query keywords from the query statement to obtain target keywords;
[0043] The interest knowledge graph corresponding to the user identifier is obtained by acquiring the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, and obtaining the weight of each instance entity in the reading text. The weight of each expanded entity in the expanded entity set is calculated based on the weight of each instance entity in the reading text. The graph is built using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0044] The target interest entity is identified from the interest knowledge graph, and recommendation information is obtained based on the target keywords and target interest entity. The recommendation information is then returned to the terminal corresponding to the user identifier.
[0045] An information recommendation device, the device comprising:
[0046] The instruction receiving module is used to receive information recommendation instructions, which carry user identifiers and query statements.
[0047] The extraction module is used to extract query keywords from query statements to obtain target keywords;
[0048] The graph acquisition module is used to acquire the interest knowledge graph corresponding to the user identifier. The interest knowledge graph is obtained by acquiring the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, acquiring the weight of each instance entity in the reading text, calculating the weight of each expanded entity in the expanded entity set based on the weight of each instance entity in the reading text, and building the graph using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0049] The recommendation module is used to identify target interest entities from the interest knowledge graph, obtain recommendation information based on target keywords and target interest entities, and return the recommendation information to the terminal corresponding to the user identifier.
[0050] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:
[0051] Receive information recommendation instructions, which carry user identifiers and query statements;
[0052] Extract query keywords from the query statement to obtain target keywords;
[0053] The interest knowledge graph corresponding to the user identifier is obtained by acquiring the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, and obtaining the weight of each instance entity in the reading text. The weight of each expanded entity in the expanded entity set is calculated based on the weight of each instance entity in the reading text. The graph is built using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0054] The target interest entity is identified from the interest knowledge graph, and recommendation information is obtained based on the target keywords and target interest entity. The recommendation information is then returned to the terminal corresponding to the user identifier.
[0055] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0056] Receive information recommendation instructions, which carry user identifiers and query statements;
[0057] Extract query keywords from the query statement to obtain target keywords;
[0058] The interest knowledge graph corresponding to the user identifier is obtained by acquiring the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, and obtaining the weight of each instance entity in the reading text. The weight of each expanded entity in the expanded entity set is calculated based on the weight of each instance entity in the reading text. The graph is built using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0059] The target interest entity is identified from the interest knowledge graph, and recommendation information is obtained based on the target keywords and target interest entity. The recommendation information is then returned to the terminal corresponding to the user identifier.
[0060] A computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0061] Receive information recommendation instructions, which carry user identifiers and query statements;
[0062] Extract query keywords from the query statement to obtain target keywords;
[0063] The interest knowledge graph corresponding to the user identifier is obtained by acquiring the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, and obtaining the weight of each instance entity in the reading text. The weight of each expanded entity in the expanded entity set is calculated based on the weight of each instance entity in the reading text. The graph is built using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0064] The target interest entity is identified from the interest knowledge graph, and recommendation information is obtained based on the target keywords and target interest entity. The recommendation information is then returned to the terminal corresponding to the user identifier.
[0065] The aforementioned information recommendation method, apparatus, computer equipment, storage medium, and computer program product receive an information recommendation instruction carrying a user identifier and a query statement; extract query keywords from the query statement to obtain target keywords; then determine the corresponding target interest entity from the interest knowledge graph corresponding to the user identifier; and finally obtain recommendation information using the target keywords and target interest entities. Since the entities in the interest knowledge graph can more accurately represent the user's interests, the obtained recommendation information is more accurate. Attached Figure Description
[0066] Figure 1 This is an application environment diagram of a knowledge graph construction method in one embodiment;
[0067] Figure 2 This is a flowchart illustrating a knowledge graph construction method in one embodiment;
[0068] Figure 3 This is a schematic diagram illustrating the process of obtaining an instance entity set in one embodiment;
[0069] Figure 4 This is a flowchart illustrating the entity linking process in a specific embodiment.
[0070] Figure 5 This is a flowchart illustrating entity disambiguation in one embodiment;
[0071] Figure 6 This is a schematic diagram of the entity disambiguation framework in a specific embodiment;
[0072] Figure 7 This is a schematic diagram of the process for obtaining an extended instance entity set in one embodiment;
[0073] Figure 8 This is a schematic diagram of the process for obtaining an extended conceptual entity set in one embodiment;
[0074] Figure 9 This is a schematic diagram illustrating the process of obtaining the first extended concept entity set and the second extended concept entity set in one embodiment;
[0075] Figure 10 This is a schematic diagram of the process for obtaining the extended instance entity weights in one embodiment;
[0076] Figure 11 This is a schematic diagram of the process for obtaining the weights of extended conceptual entities in one embodiment;
[0077] Figure 12 This is a schematic diagram illustrating the process of obtaining the weights of the first and second extended concept entities in one embodiment.
[0078] Figure 13 This is a schematic diagram of the process for obtaining the instance weight decay factor in one embodiment;
[0079] Figure 14 This is a schematic diagram illustrating the relationship between time interval and instance weight decay factor in a specific embodiment;
[0080] Figure 15 This is a schematic diagram of an interest knowledge graph in a specific embodiment;
[0081] Figure 16 for Figure 15 A schematic diagram of the concept entity weights of the interest knowledge graph in a specific embodiment;
[0082] Figure 17 for Figure 15A schematic diagram of another interest knowledge graph in a specific embodiment;
[0083] Figure 18 for Figure 15 A schematic diagram of concept entity weights in another interest knowledge graph in a specific embodiment;
[0084] Figure 19 This is a flowchart illustrating an information recommendation method in one embodiment;
[0085] Figure 20 This is a flowchart illustrating a knowledge graph construction method in a specific embodiment.
[0086] Figure 21 This is a schematic diagram of the process of constructing an interest knowledge graph in a specific embodiment;
[0087] Figure 22 This is a schematic diagram of a personal interest knowledge graph in a specific embodiment;
[0088] Figure 23 This is a structural block diagram of a knowledge graph construction device in one embodiment;
[0089] Figure 24 This is a structural block diagram of an information recommendation device in one embodiment;
[0090] Figure 25 This is an internal structural diagram of a computer device in one embodiment;
[0091] Figure 26 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0092] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0093] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0094] The solutions provided in this application involve technologies such as knowledge graphs in artificial intelligence, and are specifically illustrated through the following embodiments:
[0095] The knowledge graph construction method provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 can send knowledge graph construction instructions to the server. Server 104, based on these instructions, can retrieve reading text from the data storage system containing the target user's identifier within a preset time period. Based on the reading text, it performs entity linking to obtain a set of instance entities corresponding to the reading text. Server 104 expands the instance entities in the instance entity set to obtain an expanded instance entity set. Based on the expanded instance entity set and the instance entity set, it obtains a target instance entity set. Based on the target instance entity set, it performs concept expansion to obtain an expanded concept entity set. Server 104 obtains the weight of each instance entity, which refers to the number of times each instance entity appears in the reading text and the target reading text within the target time period before the preset time period. Server 104 calculates the instance similarity between each expanded instance entity and the associated instance entities within each instance entity. Using the instance entity weights and instance similarity, it calculates the expanded instance entity weights to obtain the weights of each expanded instance entity. Server 104 calculates the conceptual similarity between each extended conceptual entity in the extended conceptual entity set and the associated target instance entities in each target instance entity set. It calculates the extended conceptual entity weights using the weights of each instance entity, each extended instance entity, and the conceptual similarity. Server 104 then establishes an interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended conceptual entity set, the weights of each instance entity, each extended instance entity, and each extended conceptual entity. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0096] In one embodiment, such as Figure 2 As shown, a knowledge graph construction method is provided, which can be applied to... Figure 1Taking a server as an example, it can be understood that this method can also be applied to a terminal, or to a system that includes both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the following steps are included:
[0097] Step 202: Obtain the reading text of the target user identifier within a preset time period, and perform entity linking based on the reading text to obtain the instance entity set corresponding to the reading text.
[0098] The target user identifier is used to uniquely identify the user to be included in the knowledge graph. The preset time period refers to a pre-set time frame, such as a day, a week, a month, etc. The reading text refers to the text that the user has browsed; this text can be of various types, such as news text, advertising text, chat conversation text, etc. Entity links are used to link entity words found in the text to labeled entities in the knowledge base. The instance entity set refers to the collection of instance entities in the reading text. An instance entity is an entity of instance type. A knowledge graph is a vast knowledge network where nodes represent entities, and edges between nodes represent relationships between entities. Entities include both concepts and instances.
[0099] Specifically, the server can retrieve text viewed by the target user within a preset time period from a database to obtain the reading text. Alternatively, the server can collect text viewed by the target user within the same time period from the internet to obtain the reading text. The server can also obtain all text viewed by the target user within the preset time period from the business service provider and use all of this text as the reading text. Then, the server performs entity linking on the reading text to obtain the instance entity set corresponding to the reading text. In one embodiment, the server can also perform named entity recognition on the reading text to obtain the instance entity set corresponding to the reading text.
[0100] Step 204: Based on each instance entity in the instance entity set, perform instance expansion to obtain an expanded instance entity set. Based on the expanded instance entity set and the instance entity set, obtain the target instance entity set. Based on each target instance entity in the target instance entity set, perform concept expansion to obtain an expanded concept entity set.
[0101] The extended instance entity set refers to the collection of all extended instance entities. An extended instance entity is an instance entity obtained by extending an existing instance entity based on the relationships between the existing instance entity and pre-defined instance entities. The target instance entity set refers to the collection of all target instance entities. Each target instance entity includes both the extended instance entities in the extended instance entity set and the instance entities in the instance entity set. The extended concept entity set refers to the collection of all extended concept entities. An extended concept entity is a concept entity obtained by extending an existing instance entity based on the relationships between the target instance entity and pre-defined concept entities and instance entities. A concept entity is an entity of a concept type.
[0102] Specifically, the server uses the instance entities in the instance entity set to expand the instance entities according to the pre-defined associations between instance entities, resulting in instance entities associated with each instance entity in the instance entity set, thus obtaining the expanded instance entity set. Based on the expanded instance entity set and the instance entity set, the target instance entity set is obtained. Then, using the target instance entities in the target instance entity set, the server uses the pre-defined associations between instance entities and conceptual entities to expand the conceptual entities, resulting in conceptual entities associated with each target instance entity in the target instance entity set, thus obtaining the expanded conceptual entity set.
[0103] Step 206: Obtain the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0104] The target time period before the preset time period refers to the pre-set time period preceding the preset time period, and the target reading text refers to all text viewed by the user within the target reading time period. For example, if the preset time period is May 12th, the target time period could be the 10 days prior to May 12th, i.e., the period from May 2nd to May 11th. The target reading text is obtained by retrieving all text viewed by the user within the period from May 2nd to May 11th. Instance entity weight refers to the number of times an instance entity appears in both the reading text and the target reading text. Entity weight is used to characterize the user's level of interest in that entity.
[0105] Specifically, the server obtains the occurrence count of each instance entity and uses this count as the weight of each instance entity. The server can pre-count the occurrence count of each instance entity in the target reading text within the target time period, then count the occurrence count of each instance entity in the reading text, and finally calculate the sum of these occurrence counts, using this final sum as the instance entity weight.
[0106] Step 208: Calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and use the weights of each instance entity and the instance similarity to calculate the weights of the extended instance entities to obtain the weights of each extended instance entity.
[0107] Instance similarity is used to characterize the degree of similarity between extended instance entities and associated instance entities. Extended instance entity weight refers to the entity weight of the extended instance type.
[0108] Specifically, the server can vectorize each extended instance entity and each instance entity, resulting in extended instance entity vectors and instance entity vectors. The server uses a similarity algorithm to calculate the similarity between each extended instance entity vector and its associated instance entity vectors, obtaining the instance similarity level. This similarity algorithm can be cosine similarity, distance similarity, etc. The server then uses the weight of each instance entity to weight the corresponding instance similarity level, obtaining the extended instance entity weight for each extended instance entity.
[0109] In one embodiment, after expanding the instance entities in the instance entity set to obtain an expanded instance entity set, the server directly calculates the instance similarity between each expanded instance entity and its associated instance entities. It then uses the weights of each instance entity and the instance similarity to calculate the weights of the expanded instance entities, thus obtaining the weights of each expanded instance entity. After obtaining the weights of each expanded instance entity, the server obtains the target instance entity set based on the expanded instance entity set and the instance entity set. Finally, it expands the concept of each target instance entity in the target instance entity set to obtain an expanded concept entity set.
[0110] Step 210: Calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. Use the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0111] Conceptual similarity is used to characterize the degree of similarity between a concept instance entity and its associated target instance entity. Extended concept entity weight refers to the entity weight of the extended concept type.
[0112] Specifically, the server can vectorize each target instance entity and each extended concept entity, resulting in vectors for each extended concept entity and each target instance entity. The server uses a similarity algorithm to calculate the similarity between each extended concept entity vector and its associated target instance entity vector, obtaining the degree of concept similarity. This similarity algorithm can be cosine similarity, distance similarity, etc. The server then uses a weighted average of the weights of each instance entity, each extended instance entity, and the degree of concept similarity to calculate the weights of each extended concept entity.
[0113] Step 212: Establish an interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0114] Among them, the interest knowledge graph refers to the knowledge graph obtained by the target user's reading text within a preset time period. The interest knowledge graph includes each target instance entity, each extended concept entity, each instance entity weight, each extended instance entity weight, each extended concept entity weight, and the relationship between entities.
[0115] Specifically, the server establishes an initial knowledge graph based on the target instance entity set, the extended concept entity set, and the relationships between the entities. Then, it sets the weights of the entities in the initial knowledge graph according to the weights of each instance entity, each extended instance entity, and each extended concept entity, to obtain the interest knowledge graph corresponding to the target user identifier.
[0116] In the aforementioned knowledge graph construction method, the reading text of the target user identifier within a preset time period is obtained. Entity links are then established based on the reading text to obtain a set of instance entities corresponding to the reading text. Each instance entity in the instance entity set is then expanded to obtain an expanded instance entity set. The target instance entity set is obtained from the expanded instance entity set and the target instance entity set. Conceptual expansion is then performed on each target instance entity in the target instance entity set to obtain an expanded concept entity set. Next, the weights of each instance entity are obtained, and the weights of each expanded instance entity are calculated using these weights and instance similarity. Finally, the weights of each expanded concept entity are calculated using the target instance entity set, the expanded concept entity set, the weights of each instance entity, the weights of each expanded instance entity, and the weights of each expanded concept entity. Finally, the interest knowledge graph corresponding to the target user identifier is constructed using the target instance entity set, the expanded concept entity set, the weights of each instance entity, the weights of each expanded instance entity, and the weights of each expanded concept entity. This improves the accuracy of the constructed interest knowledge graph in representing user interests.
[0117] In one embodiment, such as Figure 3As shown, step 202 involves linking entities based on the reading text to obtain a set of instance entities corresponding to the reading text, including:
[0118] Step 302: Entity word recognition is performed based on the reading text to obtain each entity word.
[0119] Among them, entity words refer to words identified in the reading text that are not linked to nodes in the knowledge base.
[0120] Specifically, the server can use all names in a pre-configured alias table to perform string searches on the reading text and use a string matching algorithm to quickly identify the boundaries of each entity word. The string matching algorithm can use a multi-pattern matching algorithm, such as the Aho-Corasick algorithm (which scans the text by preprocessing the pattern string into a deterministic finite state automaton). The alias table stores words with slightly different names but consistent meanings, such as "Lu Xun" and "Zhou Shuren". In a specific embodiment, the alias table can be obtained through hyperlinks on a Chinese wiki page.
[0121] The server can also use sequence labeling algorithms to identify individual entity words in the reading text. Alternatively, the server can use named entity recognition models to identify individual entity words in the reading text.
[0122] Step 304: Based on each entity word, retrieve entities from the preset knowledge base to obtain the candidate entity set corresponding to each entity word.
[0123] The preset knowledge base refers to a pre-configured database containing entities. The candidate entity set refers to the set of all related entities for each entity term in the preset knowledge base. The candidate entity set includes all candidate entities.
[0124] Specifically, the server retrieves entities from a pre-defined knowledge base for each entity term, obtaining a set of all relevant entities and a candidate entity set for each entity term. Alternatively, the server can use an alias table for entity retrieval, obtaining separate candidate entity sets for each entity term.
[0125] Step 306: Perform entity disambiguation based on the candidate entity set corresponding to each entity word to obtain the entity corresponding to each entity word, and obtain the instance entity set corresponding to the reading text based on the entity corresponding to each entity word.
[0126] Entity disambiguation is used to sort candidate entities in the candidate entity set and select the most similar candidate entity as the result of instance linking.
[0127] Specifically, the server can also perform entity disambiguation on the candidate entity set corresponding to each entity word, that is, select the most similar entity from the candidate entity set as the entity corresponding to the entity word. By traversing all the candidate entity sets corresponding to entity words, the entities corresponding to each entity word are obtained, forming the instance entity set corresponding to the reading text. In a specific embodiment, such as... Figure 4 The diagram shows the process of entity linking. The process involves identifying entity words in the reading text to obtain entity word-annotated text, then recalling each candidate entity corresponding to the entity word from the knowledge base, calculating the similarity score between the entity word and each candidate entity to obtain the score of each candidate entity, and finally selecting the candidate entity with the highest score to obtain the target entity corresponding to that entity word.
[0128] In one embodiment, such as Figure 5 As shown, step 306 involves performing entity disambiguation based on the candidate entity sets corresponding to each entity word, to obtain the entities corresponding to each entity word, including:
[0129] Step 502: Determine the current entity word from the various entity words and obtain the entity text corresponding to the current entity word;
[0130] Step 504: Input the entity text and the corresponding candidate entity set into the entity disambiguation model. The entity disambiguation model maps the entity text and the corresponding candidate entity set into the vector space respectively, to obtain the entity word vector corresponding to the current entity word and the candidate entity vector set corresponding to the candidate entity set. Calculate the similarity between the entity word vector and the candidate entity vector in the candidate entity vector set respectively, and determine the current entity corresponding to the current entity word from the candidate entity set based on the similarity.
[0131] Here, the current entity word refers to the entity word whose corresponding entity is to be determined. Entity text refers to the context text corresponding to the current entity word. The entity disambiguation model is used to score the similarity between the input entity word and the candidate entities. This entity disambiguation model is pre-trained using a neural network algorithm, and it can be a dual-encoder model. The current entity refers to the entity corresponding to the current entity word.
[0132] Specifically, the server sequentially takes each entity word as the current entity word and retrieves the corresponding entity text from the reading text. Then, the entity text and the corresponding candidate entity set are input into the entity disambiguation model. The entity disambiguation model maps the entity text and the corresponding candidate entity set into vector spaces, obtaining the entity word vector corresponding to the current entity word and the candidate entity vector set corresponding to the candidate entity set. It calculates the similarity between the entity word vector and each candidate entity vector in the candidate entity vector set, and then selects the candidate entity with the lowest similarity score as the current entity corresponding to the current entity word. In a specific implementation example, as follows... Figure 6The diagram shows a framework for entity disambiguation. The entity word context and the candidate entity context are segmented into words to obtain segmentation results. The segmentation results are then encoded using an encoder and input into a feedforward neural network for vectorization to obtain entity word vectors and candidate entity vectors. The similarity between the entity word vectors and the candidate entity vectors is then calculated. Finally, the similarity between the entity word vector and each candidate entity vector is obtained, and the candidate entity with the highest similarity is selected as the entity corresponding to the entity word.
[0133] In the above embodiments, the instance entity set is obtained by entity linking, which improves the accuracy of the obtained instance entity set.
[0134] In one embodiment, such as Figure 7 As shown, step 204 involves expanding the instance entities in the instance entity set to obtain an expanded instance entity set, including:
[0135] Step 702: Determine the current instance entity from the instance entity set, and use the current instance entity to search for each associated candidate instance entity in the preset knowledge base according to the preset association relationship.
[0136] Here, "current instance entity" refers to the instance entity that needs to be expanded. "Preset association relationship" refers to the pre-set association relationship between instances, such as instances expanded from "related to" relationships. "Candidate instance entity" refers to instance entities that need further confirmation.
[0137] Specifically, the server can expand upon this knowledge base using attributes such as superordinate, subordinate, and related attributes. The server selects an instance entity from the instance entity set as the current instance entity, and then searches the preset knowledge base for candidate instance entities associated with the current instance entity according to preset association relationships.
[0138] Step 704: Calculate the instance similarity between the current instance entity and each candidate instance entity, and select the extended instance entity associated with the current instance entity from each candidate instance entity based on the instance similarity.
[0139] Among them, instance similarity is used to characterize the similarity between the current instance entity and the candidate instance entity. The higher the instance similarity, the more similar the current instance entity and the candidate instance entity are.
[0140] Specifically, the server can obtain the current instance entity vector corresponding to the current instance entity and the candidate instance entity vectors corresponding to each candidate instance entity. Then, it uses a cosine similarity algorithm to calculate the instance similarity between the current instance entity vector and each candidate instance entity vector. The instance similarity is then compared to a pre-set expansion stopping threshold. If the instance similarity is not lower than the expansion stopping threshold, the candidate instance entity corresponding to that instance similarity is selected as an expanded instance entity. If the instance similarity is lower than the expansion stopping threshold, expansion stops, meaning the candidate instance entity corresponding to that instance similarity is not selected as an expanded instance entity.
[0141] Step 706: Traverse each instance entity in the instance entity set to obtain the extended instance entity set.
[0142] Specifically, the server returns the steps of determining the current instance entity from the instance entity set, using the current instance entity to search for each associated candidate instance entity in the preset knowledge base according to the preset association relationship, until all instance entities in the instance entity set are traversed, and the extended instance entity set is obtained based on all extended instance entities obtained from the extension.
[0143] In the above embodiments, by calculating the instance similarity between the current instance entity and each candidate instance entity, and selecting the extended instance entity associated with the current instance entity from each candidate instance entity based on the instance similarity, the accuracy of the obtained extended instance entity set is improved.
[0144] In one embodiment, such as Figure 8 As shown, step 204 involves conceptual expansion based on each target instance entity in the target instance entity set to obtain an expanded conceptual entity set, including:
[0145] Step 802: Obtain instance relationships and subclass relationships, and expand the concepts based on each target instance entity in the preset knowledge base according to the instance relationships to obtain the first expanded concept entity set.
[0146] In this context, instance relationships refer to the relationships between instances and concepts in a knowledge graph, known as the instanceOf relationship. Subclass relationships refer to the relationships between parent and child concepts in a knowledge graph, known as the subClassOf relationship; parent and child concepts can have multiple levels. The first extended concept entity set refers to the set of first extended concept entities obtained through instance relationships.
[0147] Specifically, the server obtains instance relationships and subclass relationships from a preset knowledge base, and then searches for each conceptual entity associated with each target instance entity in the preset knowledge base according to the instance relationships to obtain the first extended conceptual entity set.
[0148] Step 804: Based on the subclass relationship, perform concept expansion in the preset knowledge base according to each first extended concept entity in the first extended concept entity set to obtain the second extended concept entity set.
[0149] The second extended conceptual entity set refers to the set of second extended conceptual entities obtained through subclass relations.
[0150] Specifically, the server searches the preset knowledge base for each concept entity associated with each first extended concept entity in the first extended concept entity set according to the subclass relationship, and obtains the second extended concept entity set.
[0151] Step 806: Obtain the extended concept entity set based on the first extended concept entity set and the second extended concept entity set.
[0152] Specifically, the server obtains the extended concept entity set based on all the first extended concept entities and all the second extended concept entities.
[0153] In one embodiment, such as Figure 9 As shown, step 802 involves expanding the concepts based on each target instance entity in the preset knowledge base according to the instance relationships, to obtain the first expanded concept entity set, which includes:
[0154] Step 902: Determine the current target instance entity from among the various target instance entities, and use the current target instance entity to search for the associated first candidate extended concept entities in the preset knowledge base according to the instance relationship.
[0155] Here, the current target instance entity refers to the target instance entity that needs to be expanded into a conceptual entity. The first candidate expanded conceptual entity refers to the first candidate expanded conceptual entity.
[0156] Specifically, the server uses the current target instance entity to search for all associated conceptual entities in the preset knowledge base according to the instance relationship, and takes each conceptual entity as the first candidate extended conceptual entity.
[0157] Step 904: Calculate the first concept similarity between the current target instance entity and each first candidate extended concept entity, and select the first extended concept entity corresponding to the current target instance entity from each first candidate extended concept entity based on the first concept similarity.
[0158] The first concept similarity is used to characterize the similarity between the current target instance entity and the first candidate extended concept entity. The higher the similarity, the more similar the current target instance entity and the first candidate extended concept entity are.
[0159] Specifically, the server can obtain the current target instance entity vector corresponding to the current target instance entity and the first candidate extended concept entity vectors corresponding to each first candidate extended concept entity from a preset knowledge base. Then, it uses a cosine similarity algorithm to calculate the cosine similarity between the current target instance entity vector and the first candidate extended concept entity vectors, obtaining the similarity degree of each first concept. The similarity degree of each first concept is compared with a pre-set extension stopping threshold. When the first concept similarity degree is not lower than the extension stopping threshold, the first candidate extended concept entity corresponding to that first concept similarity degree is selected as the first extended concept entity. When the first concept similarity degree is lower than the extension stopping threshold, the extension is stopped, meaning the first candidate extended concept entity corresponding to that first concept similarity degree is not selected as the first extended concept entity.
[0160] Step 906: Traverse each target instance entity in the target instance entity set to obtain the first extended concept entity set.
[0161] Specifically, the server returns the steps of determining the current target instance entity from each target instance entity, using the current target instance entity to search for the associated first candidate extended concept entities in the preset knowledge base according to the instance relationship, until each target instance entity in the target instance entity set is traversed, and the first extended concept entity set is obtained based on the selected first extended concept entities.
[0162] In one embodiment, such as Figure 9 As shown, step 804 involves expanding the concept in the preset knowledge base according to the subclass relationship, based on each first extended concept entity in the first extended concept entity set, to obtain the second extended concept entity set, including:
[0163] Step 908: Determine the current first extended concept entity from the first extended concept entity set, and use the current first extended concept entity to search for each associated second candidate extended concept entity in the preset knowledge base according to the subclass relationship.
[0164] Here, the current first extended conceptual entity refers to the first extended conceptual entity that currently needs further expansion. The second candidate extended conceptual entity refers to the candidate second extended conceptual entity.
[0165] Specifically, the server selects a first extended concept entity from the first extended concept entity set to obtain the current first extended concept entity. Then, according to the subclass relationship, it searches in the preset knowledge base for all concept entities associated with the current first extended concept entity and uses each associated concept entity as a second candidate extended concept entity.
[0166] Step 910: Calculate the second concept similarity between the current first extended concept entity and each second candidate extended concept entity, and select the second extended concept entity corresponding to the current first extended concept entity from each second candidate extended concept entity based on the second concept similarity.
[0167] The second concept similarity is used to characterize the degree of similarity between the current first extended concept entity and the second candidate extended concept entity. The higher the similarity, the more similar the current first extended concept entity and the second candidate extended concept entity are.
[0168] Specifically, the server can obtain the current first extended concept entity vector and the second candidate extended concept entity vector corresponding to each second candidate extended concept entity from a preset knowledge base. It then uses a cosine similarity algorithm to calculate the similarity between the current first extended concept entity vector and each second candidate extended concept entity vector, obtaining the similarity degree of each second concept. The similarity degree of each second concept is compared with a pre-set extension stopping threshold. When the similarity degree is not lower than the extension stopping threshold, the second candidate extended concept entity corresponding to that similarity degree is selected as the second extended concept entity. When the similarity degree is lower than the extension stopping threshold, extension is stopped, meaning the second candidate extended concept entity corresponding to that similarity degree is not selected as the second extended concept entity.
[0169] Step 912: Traverse each first extended concept entity in the first extended concept entity set to obtain the second extended concept entity set.
[0170] Specifically, the server returns a linear sequence of steps: determining the current first extended concept entity from the first extended concept entity set; using the current first extended concept entity to search for each associated second candidate extended concept entity in the preset knowledge base according to the subclass relationship; and so on, until all first extended concept entities in the first extended concept entity set are traversed. Based on the selected second extended concept entities, the second extended concept entity set is obtained.
[0171] In the above embodiments, a first extended concept entity set is obtained by expanding the concept based on each target instance entity in the preset knowledge base according to the instance relationship, and a second extended concept entity set is obtained by expanding the concept based on each first extended concept entity in the first extended concept entity set according to the subclass relationship in the preset knowledge base, thereby obtaining an extended concept entity set, which improves the accuracy of the obtained extended concept entity set.
[0172] In one embodiment, such as Figure 10As shown, in step 208, the instance similarity between each extended instance entity and the associated instance entities within each instance entity is calculated. The weights of the extended instance entities are then calculated using the weights of each instance entity and the instance similarity, resulting in the weights of each extended instance entity, including:
[0173] Step 1002: Determine the current extended instance entity from among the various extended instance entities, and determine the instance entity associated with the current extended instance entity from among the various instance entities.
[0174] Step 1004: Obtain the associated instance entity weights corresponding to the associated instance entities from the weights of each instance entity, and determine the associated instance similarity between the current extended instance entity and the associated instance entity from the instance similarity.
[0175] Among them, the current extended instance entity refers to the extended instance entity whose weight needs to be calculated at present.
[0176] Specifically, when the server needs to calculate the weight of an extended instance entity, it first selects the current extended instance entity from among all extended instance entities. Then, it retrieves the instance entities associated with the current extended instance entity from among all instance entities. In one embodiment, there can be multiple instance entities associated with the current extended instance entity; for example, at least two instance entities associated with the current extended instance entity are retrieved. Simultaneously, the instance entity weight corresponding to each instance entity associated with the current extended instance entity is retrieved. At this point, the similarity between the current extended instance entity and each associated instance entity is obtained from the instance similarity data, thus obtaining the instance similarity of each associated instance entity. This instance similarity refers to the similarity between the instance entity and the current extended instance entity. The current extended instance entity can be obtained by extending different instance entities.
[0177] Step 1006: Calculate the product of the associated instance entity weight and the similarity of the associated instances to obtain the extended instance weight corresponding to the current extended instance entity.
[0178] Specifically, when the current extended instance entity has one and only one associated instance entity, the server calculates the product of the associated instance entity weight and the similarity of the corresponding associated instance, and uses this product as the extended instance weight corresponding to the previous extended instance entity. When the current extended instance entity has multiple corresponding instance entities, the server calculates the product of the associated instance entity weight and the similarity of the corresponding associated instance, and then calculates the sum of all products to obtain the extended instance weight corresponding to the current extended instance entity. In a specific embodiment, the extended instance weight can be calculated using the formula (1) shown below.
[0179] Formula (1)
[0180] in, Indicates an extended instance entity. This indicates the extended instance weight. v represents the associated instance entity. Indicates and A collection of instance entities in the associated reading text. Indicates the instance entity weight. Represents extended instance entities Related instance entities and extended instance entities The degree of similarity between instances. In a specific embodiment, the degree of similarity between instances can be calculated using the formula (2) shown below.
[0181] Formula (2)
[0182] in, It represents the degree of similarity between instances, and can also be understood as the edge weight of a knowledge graph. Represents an instance entity. Word vectors representing instance entities, Represents the associated extended instance entity. This represents the word vector of the extended instance entity.
[0183] Step 1008: Traverse each extended instance entity to obtain the weight of each extended instance entity.
[0184] Specifically, the server returns the steps of determining the current extended instance entity from each extended instance entity and determining the instance entity associated with the current extended instance entity from each instance entity, until all extended instance entities are traversed and the extended instance entity weight corresponding to each extended instance entity is obtained.
[0185] In one embodiment, such as Figure 11 As shown, step 210 involves calculating the conceptual similarity between each extended conceptual entity in the extended conceptual entity set and the associated target instance entities in each target instance entity set. This is done by using the weights of each instance entity, the weights of each extended instance entity, and the conceptual similarity to calculate the weights of the extended conceptual entities, resulting in the weights of each extended conceptual entity. This includes:
[0186] Step 1102: Determine the first set of extended concept entities and the second set of extended concept entities from each extended concept entity. The first set of extended concept entities is obtained based on the target instance entity set through a preset instance relationship, and the second set of extended concept entities is obtained based on the first set of extended concept entities through a preset subclass relationship.
[0187] Specifically, the server first divides each extended concept entity into a first extended concept entity set and a second extended concept entity set. The first extended concept entity set is obtained from the target instance entities in the target instance entity set through pre-set instance relationships. The second extended concept entity set is obtained from the first extended concept entity set in the first extended concept entity set through pre-set subclass relationships.
[0188] Step 1104: Calculate the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entities associated with each target instance entity. Determine the target instance entity weight associated with each first extended concept entity from the weights of each instance entity and each extended instance entity. Calculate the weight of each first extended concept entity based on the first concept similarity and the associated target instance entity weight to obtain the weight of each first extended concept entity.
[0189] The first concept similarity is used to characterize the similarity between the first extended concept entity and the associated target instance entity. The target instance entity weight refers to the weight of the instance entity associated with the first extended concept entity. The first extended concept entity weight refers to the weight of the first extended concept entity.
[0190] Specifically, the server can use a cosine similarity algorithm to calculate the similarity between each first extended concept entity in the first extended concept entity set and the associated target instance entities in each target instance entity set, thus obtaining the similarity degree of each first concept. Then, from the weights of each instance entity and the weights of each extended instance entity, the target instance entity weights corresponding to the target instance entities associated with each first extended concept entity are determined. The weights of the first extended concept entities are then calculated using the similarity degree of each first concept and the weights of the associated target instance entities, thus obtaining the weights of each first extended concept entity.
[0191] Step 1106: Calculate the degree of second concept similarity between each second extended concept entity in the second extended concept entity set and the first extended concept entities associated with each first extended concept entity. Determine the weights of the first extended concept entities associated with each second extended concept entity from the weights of each first extended concept entity. Calculate the weights of the second extended concept entities based on the degree of second concept similarity and the weights of the associated first extended concept entities to obtain the weights of each second extended concept entity.
[0192] The second concept similarity is used to characterize the degree of similarity between the second extended concept entity and the associated first extended concept entity. The second extended concept entity weight refers to the weight of the second extended concept entity.
[0193] Specifically, the server can use a cosine similarity algorithm to calculate the similarity between each second extended concept entity in the second extended concept entity set and its associated first extended concept entity, thus obtaining the similarity degree of each second concept. Then, from the weights of each first extended concept entity, the weight of the first extended concept entity associated with each second extended concept entity is determined. Using the similarity degree of each second concept and the weight of the associated first extended concept entity, the weights of the second extended concept entities are calculated to obtain the weights of each second extended concept entity.
[0194] Step 1108: Obtain the weights of each extended concept entity based on the weights of each first extended concept entity and the weights of each second extended concept entity.
[0195] Specifically, after obtaining the weight of each first extended concept entity and the weight of each second extended concept entity, the server obtains the weight of the extended concept entity corresponding to each extended concept entity.
[0196] In one embodiment, such as Figure 12 As shown, step 1104 involves calculating the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entities associated with each target instance entity, determining the target instance entity weights associated with each first extended concept entity from the weights of each instance entity and each extended instance entity, and calculating the weights of each first extended concept entity based on the first concept similarity and the associated target instance entity weights, thereby obtaining the weights of each first extended concept entity, including:
[0197] Step 1202: Determine the current first extended concept entity from each first extended concept entity, and determine the current target instance entity associated with the current first extended concept entity from each target instance entity.
[0198] Here, the current first extended concept entity refers to the first extended concept entity that needs to be weighted. The current target instance entity refers to the target instance entity associated with the current first extended concept entity. There can be multiple current target instance entities, that is, at least two current target instance entities are identified as associated with the current first extended concept entity.
[0199] Specifically, the server can sequentially select first extended concept entities from each first extended concept entity to obtain the current first extended concept entity. Then, it can determine the target instance entity associated with the current first extended concept entity from each target instance entity to obtain the current target instance entity. In one embodiment, multiple target instance entities are determined to be associated with the current first extended concept entity, that is, at least two current target instance entities are obtained.
[0200] Step 1204: Calculate the similarity between the current first extended concept entity and the current target instance entity, and determine the current target instance entity weight corresponding to the current target instance entity from the weights of each instance entity and the weights of each extended instance entity.
[0201] Among them, the current first concept similarity is used to characterize the similarity between the current first extended concept entity and the current target instance entity.
[0202] Specifically, the server can perform word vectorization on the current first extended concept entity and the current target instance entity respectively, obtaining the current first extended concept entity vector and the current target instance entity vector, or it can directly obtain the current first extended concept entity vector and the current target instance entity vector from the database. Then, it uses a cosine similarity algorithm to calculate the cosine similarity between the current first extended concept entity vector and the current target instance entity vector, obtaining the current first concept similarity level. When there are multiple current target instance entities, the server calculates the cosine similarity between the current first extended concept entity vector and each current target instance entity vector, obtaining multiple current first concept similarity levels. Then, the server determines the current target instance entity weight corresponding to the current target instance entity from the weights of each instance entity and the weights of each extended instance entity.
[0203] Step 1206: Calculate the product of the current first concept similarity and the current target instance entity weight to obtain the first extended concept entity weight corresponding to the current first extended concept entity.
[0204] Specifically, the server multiplies the current first concept similarity with the current target instance entity weight, and uses the result as the first extended concept entity weight corresponding to the current first extended concept entity. In a specific embodiment, the first extended concept entity weight corresponding to the current first extended concept entity can be calculated using the formula (3) shown below.
[0205] Formula (3)
[0206] in, This represents the first extended conceptual entity. Indicates the weight of the first extended concept entity. Indicates and The associated set of target instance entities, which can be extended instance entities or instance entities in the reading text. V represents the associated target instance entities. This represents the target instance entity weight. Indicates the degree of similarity between the first concepts.
[0207] Step 1208: Traverse each first extended concept entity to obtain the weight of each first extended concept entity.
[0208] Specifically, the server returns the steps of determining the current first extended concept entity from each first extended concept entity and determining the current target instance entity associated with the current first extended concept entity from each target instance entity, until the traversal of each first extended concept entity is completed, and obtains the weight of the first extended instance entity corresponding to each first extended concept entity.
[0209] In one embodiment, such as Figure 12 As shown, step 1106 involves calculating the second concept similarity between each second extended concept entity in the second extended concept entity set and the associated first extended concept entities in each first extended concept entity set; determining the weights of the associated first extended concept entities for each second extended concept entity from the weights of each first extended concept entity; and calculating the weights of the second extended concept entities based on the second concept similarity and the associated first extended concept entity weights to obtain the weights of each second extended concept entity, including:
[0210] Step 1210: Determine the current second extended concept entity from each of the second extended concept entities, and determine the current first extended concept entity associated with the current second extended concept entity from each of the first extended concept entities.
[0211] Here, the current second extended concept entity refers to the second extended concept entity that needs to be weighted. The current first extended concept entity refers to the first extended concept entity associated with the current second extended concept entity. There can be multiple current first extended concept entities, that is, at least two current first extended concept entities are identified as associated with the current second extended concept entity.
[0212] Specifically, the server can sequentially select second extended concept entities from each of the various second extended concept entities to obtain the current second extended concept entity. Then, it can determine the first extended concept entity associated with the current second extended concept entity from each of the various first extended concept entities to obtain the current first extended concept entity. In one embodiment, multiple first extended concept entities are determined to be associated with the current second extended concept entity, that is, at least two current first extended concept entities are obtained.
[0213] Step 1212: Calculate the similarity between the current second extended concept entity and the current first extended concept entity, and determine the weight of the current first extended concept entity corresponding to the current first extended concept entity from the weights of each first extended concept entity.
[0214] Among them, the current second concept similarity is used to characterize the degree of similarity between the current second extended concept entity and the current first extended concept entity.
[0215] Specifically, the server can perform word vectorization on the current second extended concept entity and the current first extended concept entity respectively, obtaining the current second extended concept entity vector and the current first extended concept entity vector. Alternatively, it can directly obtain the current second extended concept entity vector and the current first extended concept entity vector from the database. Then, it uses a cosine similarity algorithm to calculate the cosine similarity between the current second extended concept entity vector and the current first extended concept entity vector, obtaining the current second concept similarity level. When there are multiple current first extended concept entities, the server calculates the cosine similarity between the current first extended concept entity vector and each of the current first extended concept entity vectors, obtaining multiple current second concept similarity levels. Then, the server determines the current first extended concept entity weight corresponding to the current first extended concept entity from the weights of each first extended concept entity.
[0216] Step 1214: Calculate the product of the current second concept similarity and the current first extended concept entity weight to obtain the second extended concept entity weight corresponding to the current second extended concept entity.
[0217] Specifically, the server multiplies the current similarity of the second concept with the weight of the current first extended concept entity, and uses the result of the multiplication as the weight of the second extended concept entity corresponding to the current second extended concept entity. In a specific embodiment, the weight of the second extended concept entity corresponding to the current second extended concept entity can be calculated using the formula (4) shown below.
[0218] Formula (4)
[0219] in, This represents the second extended conceptual entity. Indicates the weight of the second extended concept entity. Indicates and The set of first-level extended conceptual entities associated with the relationship. V represents the set of first-level extended conceptual entities associated with the relationship. This indicates the weight of the first extended concept entity. Indicates the degree of similarity between the two concepts.
[0220] Step 1216: Traverse each second extended concept entity to obtain the weight of each second extended concept entity.
[0221] Specifically, the server returns the steps of determining the current second extended concept entity from each second extended concept entity and determining the current first extended concept entity associated with the current second extended concept entity from each first extended concept entity, until the weight of the second extended instance entity corresponding to each second extended concept entity is obtained when traversing each second extended concept entity.
[0222] In the above embodiments, the accuracy of the obtained extended concept entity weights is improved by calculating the weights of the first and second extended concept entities respectively.
[0223] In one embodiment, after step 208, that is, after calculating the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and calculating the extended instance entity weights using the weights of each instance entity and the instance similarity, the method further includes the following step:
[0224] Obtain the instance weight decay factor, and based on the instance weight decay factor, perform weight decay on the weight of each instance entity and the weight of each extended instance entity to obtain the decay weight of each instance entity and the decay weight of each extended instance entity.
[0225] The instance weight decay factor is used to decay the weight of the instance entity and the weight of the extended instance entities, representing the change of entity weight over time. The instance entity decay weight refers to the decayed instance entity weight. The extended instance entity decay weight refers to the decayed extended instance entity weight.
[0226] Specifically, the server can directly obtain the instance weight decay factor from the database, or it can obtain the current time point and calculate the instance weight decay factor based on the current time point. Then, the instance weight decay factor is used to decay the weights of each instance entity and each extended instance entity. That is, the product of the instance weight decay factor and each instance entity weight is calculated to obtain the decayed weight of each instance entity, and the product of the instance weight decay factor and each extended instance entity weight is calculated to obtain the decayed weight of each extended instance entity.
[0227] Step 210, which involves calculating the weights of extended concept entities using the weights of each instance entity, the weights of each extended instance entity, and the concept similarity, includes the following steps:
[0228] The weights of extended concept entities are calculated based on the decay weights of each instance entity, the decay weights of each extended instance entity, and the degree of concept similarity, thus obtaining the decay weights of each extended concept entity.
[0229] Among them, the extended concept entity decay weight refers to the decayed weight of the extended concept entity.
[0230] Specifically, since the extended concept entity weight is calculated using the instance entity weight and the extended instance entity weight, when the instance entity weight and the extended instance entity weight decay, the resulting extended concept entity weight also decays accordingly, thus obtaining an extended concept entity decay weight.
[0231] Step 212, namely, establishing the interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity, includes the following steps:
[0232] A target interest knowledge graph corresponding to the target user identifier is established based on the target instance entity set, the extended concept entity set, the decay weight of each instance entity, the decay weight of each extended instance entity, and the decay weight of each extended concept entity.
[0233] Specifically, the server uses the decayed weights to build the knowledge graph, that is, it uses the target instance entity set, the extended concept entity set, the decayed weights of each instance entity, the decayed weights of each extended instance entity, and the decayed weights of each extended concept entity to build the target interest knowledge graph corresponding to the target user identifier.
[0234] In the above embodiments, entity weights are attenuated by using an instance weight attenuation factor to obtain attenuated entity weights. Then, the attenuated entity weights are used to build a target interest knowledge graph, so that the obtained target interest knowledge graph can more accurately reflect the user's interest characteristics.
[0235] In one embodiment, such as Figure 13 As shown, the instance weight decay factor is obtained, including:
[0236] Step 1302: Obtain the first target time point corresponding to each instance entity and the second target time point corresponding to each extended instance entity. The first target time point is the historical time point when the current instance entity was previously used as an extended instance entity, and the second target time point is the historical time point when the extended instance entity was previously used as an extended instance entity.
[0237] Specifically, "the current instance entity was previously used as an extended instance entity" refers to the instance entity that appeared in the extended instance entities obtained when the interest knowledge graph was previously generated, based on the instance entity in the corresponding reading text. In other words, both the first and second target time points are the time points when the corresponding entities were extended from the corresponding reading text during the previous generation of the interest knowledge graph.
[0238] Step 1304: Obtain the current time point, determine the first time interval based on the first target time point and the current time point, and determine the second time interval based on the second target time point and the current time point.
[0239] The current time point refers to the time when the interest knowledge graph was generated, which can be a specific time or a date.
[0240] Specifically, the server can obtain the system time to get the current time, and then calculate the time interval between the first target time point corresponding to each instance entity and the current time point to obtain the time interval corresponding to each instance entity, i.e., the first time interval. At the same time, the server can calculate the time interval between the second target time point corresponding to each extended instance entity and the current time point to obtain the time interval corresponding to each extended instance entity, i.e., the second time interval.
[0241] Step 1306: When the first time interval is within the preset weight decay time range, calculate the first initial weight decay factor based on the first time interval and the preset weight decay rate parameter, normalize the first initial weight decay factor, and obtain the instance weight decay factor corresponding to each instance entity.
[0242] The preset weight decay time range refers to the pre-set time range within which the entity weight decays. When the time interval is less than the preset weight decay time range, the entity weight does not decay; when the time interval is greater than the preset weight decay time range, the entity weight tends to zero. The preset weight decay rate parameter refers to the pre-set parameter used to control the rate of weight decay.
[0243] Specifically, when the server determines that the first time interval is within the preset weight decay time range, it calculates the first initial weight decay factor based on the first time interval and the preset weight decay rate parameter, normalizes the first initial weight decay factor, and obtains the instance weight decay factor corresponding to each instance entity.
[0244] Step 1308: When the second time interval is within the preset weight decay time range, calculate the second initial weight decay factor based on the second time interval and the preset weight decay rate parameter, normalize the second initial weight decay factor, and obtain the instance weight decay factor corresponding to each extended instance entity.
[0245] Specifically, when the server determines that the second time interval is within the preset weight decay time range, it calculates the second initial weight decay factor based on the second time interval and the preset weight decay rate parameter, normalizes the second initial weight decay factor, and obtains the instance weight decay factor corresponding to each extended instance entity.
[0246] In a specific embodiment, such as Figure 14 The diagram shows the relationship between the time interval and the instance weight decay factor. Before time interval t1, the instance entity weight does not decay and remains unchanged. Between t1 and t2, the instance entity weight decays, and after t2, the instance entity weight approaches 0. The instance weight decay factor can then be calculated using the formula (5) shown below.
[0247] Formula (5)
[0248] Here, x refers to the time interval, which can also be in days. t1 and t2 are set based on experience; for example, t1 can be 7 days and t2 can be 25 days. The weight decay rate is controlled by a parameter set based on experience, and can be 0.25. The range for controlling entity weights is between 0 and 1. Since instance entity weights do not decay from 0 to t1, then... .in, The definition is shown in the following formula (6):
[0249] Formula (6)
[0250] In the above embodiments, the entity weights are attenuated by using an instance weight decay factor, resulting in more accurate attenuated entity weights.
[0251] In one embodiment, after step 212, and after establishing the interest knowledge graph of the target user in the current time period based on each current instance entity, each extended instance entity, each extended concept entity, current instance weight, each extended instance weight, and extended concept weight, the following step is further included:
[0252] Obtain the interest knowledge graph corresponding to the target user identifier at each preset time point, and dynamically visualize the interest knowledge graph corresponding to the target user identifier at each preset time point.
[0253] The preset time point refers to the time point at which the interest knowledge graph is established. It can be pre-set, such as establishing the user's interest knowledge graph at one-month intervals or establishing the user's interest knowledge graph every day.
[0254] Specifically, the server obtains the interest knowledge graph corresponding to the target user at each preset time point, which is a pre-built interest knowledge graph constructed using any embodiment of the knowledge graph construction method. Then, the interest knowledge graph corresponding to the target user at each preset time point is dynamically visualized. Dynamic visualization refers to a dynamic display, such as displaying it in the form of animation or video playback, etc. In a specific embodiment, such as... Figure 15 The image shows user A's interest knowledge graph on May 8th. The size of the circles in the graph represents the weight of the corresponding nodes. Blank circles represent instance entity nodes obtained from the read text, black circles represent expanded instance entity nodes, and gray circles represent expanded concept entity nodes. Figure 16 The image shows a diagram illustrating the concept entity node weights in user A's interest knowledge graph on May 8th. When the user clicks the play button, the interest knowledge graph for each day from May 4th to May 10th is played sequentially. Figure 17 As shown, this is a display of user A's interest knowledge graph on May 21st, combined with, for example... Figure 18 The diagram showing the concept entity node weights of user A in the interest knowledge graph on May 21 clearly shows that the user's interests have shifted significantly, from concept entity A to concept entities B and C.
[0255] In one embodiment, such as Figure 19 As shown, an information recommendation method is provided, which can be applied to... Figure 1 Taking a server as an example, it can be understood that this method can also be applied to a terminal, or to a system that includes both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the following steps are included:
[0256] Step 1902: Receive information recommendation instruction, which carries a user identifier and a query statement.
[0257] The user identifier is used to uniquely identify a user. A query statement refers to a statement used to request recommendation information.
[0258] Specifically, the server receives an information recommendation instruction sent by the user terminal, which carries a user identifier and a query statement.
[0259] Step 1904: Extract query keywords from the query statement to obtain target keywords.
[0260] Specifically, the server can use a keyword extraction algorithm to extract query keywords from the query statement to obtain target keywords. These target keywords are used to represent the type of information to be recommended to the user, such as news, advertisements, products, videos, etc.
[0261] Step 1906: Obtain the interest knowledge graph corresponding to the user identifier. The interest knowledge graph is obtained by obtaining the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, obtaining the weight of each instance entity in the reading text, calculating the weight of each expanded entity in the expanded entity set based on the weight of each instance entity in the reading text, and building the graph using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0262] Specifically, the server obtains the interest knowledge graph corresponding to the user identifier. This interest knowledge graph is built based on the reading text within the most recent time period. For example, if the current information recommendation time is May 12th, the server can obtain the user's interest knowledge graph built based on the reading text on May 11th. The interest knowledge graph can be obtained using any of the above-described interest knowledge graph construction methods. For instance, by obtaining the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, obtaining the weights of each instance entity in the reading text, calculating the weights of each expanded entity in the expanded entity set based on the weights of each instance entity in the reading text, and using the instance entities in the reading text, the expanded entity set, the weights of each instance entity, and the weights of each expanded entity to build the interest knowledge graph.
[0263] Step 1908: Identify the target interest entity from the interest knowledge graph, obtain recommendation information based on the target keywords and the target interest entity, and return the recommendation information to the terminal corresponding to the user identifier.
[0264] Specifically, the server determines the target interest entity from the interest knowledge graph, which can be the entity with the highest entity weight in the interest knowledge graph. Then, it obtains recommendation information based on the target keywords and the target interest entity, and returns the recommendation information to the terminal corresponding to the user identifier.
[0265] The aforementioned information recommendation method involves receiving an information recommendation instruction, which carries a user identifier and a query statement. Query keywords are extracted from the query statement to obtain target keywords. Then, the corresponding target interest entities are determined from the interest knowledge graph corresponding to the user identifier. Finally, the target keywords and target interest entities are used to obtain recommendation information. Because the entities in the interest knowledge graph can more accurately represent the user's interests, the resulting recommendation information is more accurate.
[0266] In a specific embodiment, such as Figure 20 As shown, a knowledge graph construction method is provided, which specifically includes the following steps:
[0267] Step 2002: Obtain the reading text of the target user identifier within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text.
[0268] Step 2004: Determine the current instance entity from the instance entity set; use the current instance entity to search for each associated candidate instance entity in the preset knowledge base according to the preset association relationship; calculate the instance similarity between the current instance entity and each candidate instance entity; select the extended instance entity associated with the current instance entity from each candidate instance entity based on the instance similarity; traverse each instance entity in the instance entity set to obtain the extended instance entity set; and obtain the target instance entity set based on the extended instance entity set and the instance entity set.
[0269] Step 2006: Determine the current extended instance entity from each extended instance entity, and determine the instance entity associated with the current extended instance entity from each instance entity; obtain the associated instance entity weight corresponding to the associated instance entity from each instance entity weight, and determine the associated instance similarity between the current extended instance entity and the associated instance entity from the instance similarity; calculate the product of the associated instance entity weight and the associated instance similarity to obtain the extended instance weight corresponding to the current extended instance entity; traverse each extended instance entity to obtain the weight of each extended instance entity.
[0270] Step 2008: Obtain the instance weight decay factor, and perform weight decay on the weights of each instance entity and each extended instance entity based on the instance weight decay factor to obtain the decay weights of each instance entity and each extended instance entity.
[0271] Step 2010: Obtain instance relationships and subclass relationships; expand concepts based on each target instance entity in the preset knowledge base according to instance relationships to obtain a first expanded concept entity set; expand concepts based on each first expanded concept entity in the first expanded concept entity set in the preset knowledge base according to subclass relationships to obtain a second expanded concept entity set.
[0272] Step 2012: Calculate the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entity associated with each target instance entity. Determine the target instance entity attenuation weight associated with each first extended concept entity from the attenuation weight of each instance entity and the attenuation weight of each extended instance entity. Calculate the attenuation weight of the first extended concept entity based on the first concept similarity and the attenuation weight of the associated target instance entity to obtain the attenuation weight of each first extended concept entity.
[0273] Step 2014: Calculate the degree of second concept similarity between each second extended concept entity in the second extended concept entity set and the first extended concept entities associated with each first extended concept entity. Determine the attenuation weight of the first extended concept entity associated with each second extended concept entity from the attenuation weight of each first extended concept entity. Calculate the attenuation weight of the second extended concept entity based on the degree of second concept similarity and the attenuation weight of the associated first extended concept entity to obtain the attenuation weight of each second extended concept entity.
[0274] Step 2016: Based on the target instance entity set, the extended concept entity set, the decay weight of each instance entity, the decay weight of each extended instance entity, the decay weight of each first extended concept entity, and the decay weight of each second extended concept entity, establish the target interest knowledge graph corresponding to the target user identifier.
[0275] This application also provides an application scenario in which the above-described knowledge graph construction method is applied. Specifically, in the information service platform of an instant messaging application, an interest knowledge graph for each user can be established, such as: Figure 21The diagram illustrates the process of constructing a knowledge graph. The instant messaging application server retrieves information service articles viewed by the user within a preset time period on the information service platform. It then links these articles to obtain individual instance entities. Based on these instance entities, it expands them using the "related to" relationship to create instances. Finally, it uses all instances to expand them using the "instance of" and "subclass of" relationships to create nodes representing concepts. If the word vector similarity between the expanded entities and the original entities is below a threshold, the expansion stops. At this point, the initial knowledge graph is obtained based on all instances and all concepts. Then, the weights of all nodes in the initial knowledge graph are set, and these weights change over time. Specifically, the weight of each instance entity is obtained by acquiring the frequency of its occurrence. Then, the weight of the extended instance entities is calculated based on the similarity between the instance entities and their corresponding weights. Finally, the weight of the concept entities is calculated using the extended instance entity weights, thus obtaining the weight of each node in the knowledge graph. This results in the user's personal interest knowledge graph for the preset time period, which is then dynamically displayed. Figure 22 The diagram shows a partial illustration of the implemented personal interest knowledge graph. Then, based on the interest entity with the highest weight in the personal interest knowledge graph, information service articles related to that interest entity can be recommended to the user. It can also recommend advertisements, products, videos, live streams, etc., related to the interest entity with the highest weight in the personal interest knowledge graph. In a specific embodiment, this information recommendation method can be applied to cold-start recommendations for new services. Specifically, it can obtain an interest knowledge graph built based on user reading texts from the old service, and then, when the new service is launched, when a user uses the new service, it recommends information from the new service that the user is interested in based on the interest knowledge graph of the old service. For example, a user interest knowledge graph can be obtained from information service articles on an instant messaging application's information service platform. This user interest knowledge graph can then be applied to information recommendations for new services, such as live streaming information recommendations for a new live streaming service. When a user uses a newly launched live streaming platform, the user interest knowledge graph can be obtained from the information service articles on the information service platform. Then, the interest entity with the highest weight in the user interest knowledge graph can be obtained. Finally, live streaming information related to the interest entity with the highest weight can be recommended to the user's terminal, thereby making it easier for the user to use the newly launched live streaming platform and improving the user experience.
[0276] It should be understood that, although Figure 2-20The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated in this document, there is no strict order in which these steps are executed; they can be performed in other orders. Furthermore, Figure 2-20 At least some steps in a flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0277] In one embodiment, such as Figure 23 As shown, a knowledge graph construction device 2300 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: an instance acquisition module 2302, an expansion module 2304, a weight acquisition module 2306, an instance weight calculation module 2308, a concept weight calculation module 2310, and a knowledge graph construction module 2310, wherein:
[0278] Module 2302 is used to obtain the reading text of the target user identifier within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text.
[0279] The extension module 2304 is used to extend the instance based on each instance entity in the instance entity set to obtain an extended instance entity set, to obtain a target instance entity set based on the extended instance entity set and the instance entity set, and to extend the concept based on each target instance entity in the target instance entity set to obtain an extended concept entity set.
[0280] The weight acquisition module 2306 is used to acquire the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period.
[0281] The instance weight calculation module 2308 is used to calculate the instance similarity between each extended instance entity and the associated instance entities in each instance entity, and to calculate the weight of each extended instance entity using the weight of each instance entity and the instance similarity.
[0282] The concept weight calculation module 2310 is used to calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. It uses the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities and obtain the weights of each extended concept entity.
[0283] The graph building module 2312 is used to build an interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
[0284] In one embodiment, the instance obtains module 2302, which includes:
[0285] The recognition unit is used to identify entity words based on the reading text to obtain each entity word;
[0286] The recall unit is used to recall entities from a pre-set knowledge base based on each entity word, and obtain the candidate entity set corresponding to each entity word.
[0287] The disambiguation unit is used to perform entity disambiguation based on the candidate entity set corresponding to each entity word, to obtain the entity corresponding to each entity word, and to obtain the instance entity set corresponding to the reading text based on the entity corresponding to each entity word.
[0288] In one embodiment, the disambiguation unit is further configured to determine the current entity word from each entity word and obtain the entity text corresponding to the current entity word; input the entity text and the corresponding candidate entity set into the entity disambiguation model, the entity disambiguation model maps the entity text and the corresponding candidate entity set into the vector space respectively, to obtain the entity word vector corresponding to the current entity word and the candidate entity vector set corresponding to the candidate entity set, calculate the similarity between the entity word vector and the candidate entity vector in the candidate entity vector set respectively, and determine the current entity corresponding to the current entity word from the candidate entity set based on the similarity.
[0289] In one embodiment, the extension module 2304 is further configured to determine the current instance entity from the instance entity set, use the current instance entity to search for each associated candidate instance entity in the preset knowledge base according to the preset association relationship, calculate the instance similarity between the current instance entity and each candidate instance entity, select the extended instance entity associated with the current instance entity from each candidate instance entity based on the instance similarity, and traverse each instance entity in the instance entity set to obtain the extended instance entity set.
[0290] In one embodiment, the extension module 2304 includes:
[0291] The first concept expansion unit is used to obtain instance relationships and subclass relationships, and expands the concept based on each target instance entity in the preset knowledge base according to the instance relationships to obtain the first expanded concept entity set.
[0292] The second concept expansion unit is used to expand the concept based on each first expanded concept entity in the first expanded concept entity set in the preset knowledge base according to the subclass relationship, so as to obtain the second expanded concept entity set.
[0293] The extended concept is obtained by a unit, which is used to obtain an extended concept entity set based on the first extended concept entity set and the second extended concept entity set.
[0294] In one embodiment, the first concept extension unit is further configured to determine the current target instance entity from each target instance entity, use the current target instance entity to search for each associated first candidate extended concept entity in a preset knowledge base according to the instance relationship, calculate the first concept similarity between the current target instance entity and each first candidate extended concept entity, select the first extended concept entity corresponding to the current target instance entity from each first candidate extended concept entity based on the first concept similarity, and traverse each target instance entity in the target instance entity set to obtain the first extended concept entity set.
[0295] In one embodiment, the second concept extension unit is further configured to: determine the current first extended concept entity from the first extended concept entity set; use the current first extended concept entity to search for each associated second candidate extended concept entity in a preset knowledge base according to the subclass relationship; calculate the second concept similarity between the current first extended concept entity and each second candidate extended concept entity; select the second extended concept entity corresponding to the current first extended concept entity from each second candidate extended concept entity based on the second concept similarity; and traverse each first extended concept entity in the first extended concept entity set to obtain the second extended concept entity set.
[0296] In one embodiment, the instance weight calculation module 2308 is further configured to determine the current extended instance entity from each extended instance entity, and determine the instance entity associated with the current extended instance entity from each instance entity; obtain the associated instance entity weight corresponding to the associated instance entity from each instance entity weight, and determine the associated instance similarity between the current extended instance entity and the associated instance entity from the instance similarity; calculate the product of the associated instance entity weight and the associated instance similarity to obtain the extended instance weight corresponding to the current extended instance entity; and traverse each extended instance entity to obtain the weight of each extended instance entity.
[0297] In one embodiment, the concept weight calculation module 2310 includes:
[0298] The determining unit is used to determine a first set of extended concept entities and a second set of extended concept entities from various extended concept entities. The first set of extended concept entities is obtained based on the target instance entity set through a preset instance relationship, and the second set of extended concept entities is obtained based on the first set of extended concept entities through a preset subclass relationship.
[0299] The first calculation unit is used to calculate the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entity associated in each target instance entity, determine the target instance entity weight associated with each first extended concept entity from the weight of each instance entity and the weight of each extended instance entity, and calculate the weight of each first extended concept entity based on the first concept similarity and the weight of the associated target instance entity to obtain the weight of each first extended concept entity.
[0300] The second calculation unit is used to calculate the degree of second concept similarity between each second extended concept entity in the second extended concept entity set and the first extended concept entity associated with each first extended concept entity, determine the weight of the first extended concept entity associated with each second extended concept entity from the weight of each first extended concept entity, and calculate the weight of the second extended concept entity based on the degree of second concept similarity and the weight of the associated first extended concept entity to obtain the weight of each second extended concept entity.
[0301] The weighting unit is used to obtain the weights of each extended concept entity based on the weights of each first extended concept entity and the weights of each second extended concept entity.
[0302] In one embodiment, the first computing unit is further configured to: determine the current first extended concept entity from each of the first extended concept entities; determine the current target instance entity associated with the current first extended concept entity from each of the target instance entities; calculate the current first concept similarity between the current first extended concept entity and the current target instance entity; determine the current target instance entity weight corresponding to the current target instance entity from each instance entity weight and each extended instance entity weight; calculate the product of the current first concept similarity and the current target instance entity weight to obtain the first extended concept entity weight corresponding to the current first extended concept entity; and traverse each first extended concept entity to obtain each first extended concept entity weight.
[0303] In one embodiment, the second computing unit is further configured to: determine the current second extended concept entity from each of the second extended concept entities; determine the current first extended concept entity associated with the current second extended concept entity from each of the first extended concept entities; calculate the current second concept similarity between the current second extended concept entity and the current first extended concept entity; determine the current first extended concept entity weight corresponding to the current first extended concept entity from each of the weights of the first extended concept entities; calculate the product of the current second concept similarity and the current first extended concept entity weight to obtain the second extended concept entity weight corresponding to the current second extended concept entity; and traverse each second extended concept entity to obtain the weights of each second extended concept entity.
[0304] In one embodiment, the knowledge graph construction apparatus 2300 further includes:
[0305] The weight decay module is used to obtain the instance weight decay factor, and to decay the weight of each instance entity and the weight of each extended instance entity based on the instance weight decay factor to obtain the decay weight of each instance entity and the decay weight of each extended instance entity.
[0306] The concept weight calculation module 2310 is also used to calculate the weight of extended concept entities based on the decay weight of each instance entity, the decay weight of each extended instance entity, and the concept similarity, so as to obtain the decay weight of each extended concept entity.
[0307] The graph building module 2312 is also used to build a target interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the decay weight of each instance entity, the decay weight of each extended instance entity, and the decay weight of each extended concept entity.
[0308] In one embodiment, the weight decay module is further configured to obtain a first target time point corresponding to each instance entity and a second target time point corresponding to each extended instance entity, wherein the first target time point is the historical time point when the current instance entity was previously an extended instance entity, and the second target time point is the historical time point when the extended instance entity was previously an extended instance entity; obtain the current time point, determine a first time interval based on the first target time point and the current time point, and determine a second time interval based on the second target time point and the current time point; when the first time interval is within a preset weight decay time range, calculate a first initial weight decay factor based on the first time interval and a preset weight decay rate parameter, normalize the first initial weight decay factor to obtain the instance weight decay factor corresponding to each instance entity; when the second time interval is within a preset weight decay time range, calculate a second initial weight decay factor based on the second time interval and the preset weight decay rate parameter, normalize the second initial weight decay factor to obtain the instance weight decay factor corresponding to each extended instance entity.
[0309] In one embodiment, the knowledge graph construction apparatus 2300 further includes:
[0310] The display module is used to obtain the interest knowledge graph corresponding to the target user identifier at each preset time point, and to dynamically visualize the interest knowledge graph corresponding to the target user identifier at each preset time point.
[0311] In one embodiment, such as Figure 24As shown, an information recommendation device 2400 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: an instruction receiving module 2402, an extraction module 2404, a map acquisition module 2406, and a recommendation module 2408, wherein:
[0312] The instruction receiving module 2402 is used to receive information recommendation instructions, which carry user identifiers and query statements.
[0313] Extraction module 2404 is used to extract query keywords from query statements to obtain target keywords;
[0314] The graph acquisition module 2406 is used to acquire the interest knowledge graph corresponding to the user identifier. The interest knowledge graph is obtained by acquiring the reading text of the target user identifier within a preset time period, expanding the instances and concepts based on each instance entity in the reading text to obtain an expanded entity set, acquiring the weight of each instance entity in the reading text, calculating the weight of each expanded entity in the expanded entity set based on the weight of each instance entity in the reading text, and building the graph using the instance entities in the reading text, the expanded entity set, the weight of each instance entity, and the weight of each expanded entity.
[0315] The recommendation module 2408 is used to determine the target interest entity from the interest knowledge graph, obtain recommendation information based on the target keywords and the target interest entity, and return the recommendation information to the terminal corresponding to the user identifier.
[0316] Specific limitations regarding the knowledge graph construction device and information recommendation device can be found in the limitations on the knowledge graph construction method and information recommendation method above, and will not be repeated here. Each module in the aforementioned knowledge graph construction device and information recommendation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0317] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 25As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores entity data, knowledge graph data, etc. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a knowledge graph construction method or an information recommendation method.
[0318] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 26 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a knowledge graph construction method or an information recommendation method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0319] Those skilled in the art will understand that Figure 25 The structures shown in Figure 26 are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0320] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0321] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0322] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.
[0323] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Users may also refuse or can easily refuse the recommended information.
[0324] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0325] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0326] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for constructing a knowledge graph, characterized in that, The method includes: Obtain the reading text of the target user identifier within a preset time period, and perform entity linking based on the reading text to obtain the instance entity set corresponding to the reading text; Based on each instance entity in the instance entity set, instance expansion is performed to obtain an extended instance entity set. Based on the extended instance entity set and the instance entity set, a target instance entity set is obtained. Based on each target instance entity in the target instance entity set, concept expansion is performed to obtain an extended concept entity set. Obtain the weight of each instance entity, where the weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period. Calculate the instance similarity between each extended instance entity in the extended instance entity set and the associated instance entity in each instance entity, and use the weight of each instance entity and the instance similarity to calculate the weight of each extended instance entity to obtain the weight of each extended instance entity. Calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entities associated with each target instance entity. Use the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to calculate the weights of the extended concept entities to obtain the weights of each extended concept entity. An interest knowledge graph corresponding to the target user identifier is established based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
2. The method according to claim 1, characterized in that, The step of linking entities based on the read text to obtain the instance entity set corresponding to the read text includes: Entity word recognition is performed on the read text to obtain each entity word; Based on each entity word, entity retrieval is performed from a preset knowledge base to obtain a candidate entity set corresponding to each entity word; Entity disambiguation is performed based on the candidate entity sets corresponding to each entity word to obtain the entities corresponding to each entity word, and the instance entity set corresponding to the reading text is obtained based on the entities corresponding to each entity word.
3. The method according to claim 2, characterized in that, The entity disambiguation based on the candidate entity sets corresponding to each entity word to obtain the entities corresponding to each entity word includes: Determine the current entity word from among the various entity words, and obtain the entity text corresponding to the current entity word; The entity text and the corresponding candidate entity set are input into the entity disambiguation model. The entity disambiguation model maps the entity text and the corresponding candidate entity set into a vector space to obtain the entity word vector corresponding to the current entity word and the candidate entity vector set corresponding to the candidate entity set. The similarity between the entity word vector and the candidate entity vector in the candidate entity vector set is calculated. Based on the similarity, the current entity corresponding to the current entity word is determined from the candidate entity set.
4. The method according to claim 1, characterized in that, The step of expanding the instance entity set based on each instance entity in the instance entity set to obtain an expanded instance entity set includes: The current instance entity is determined from the set of instance entities, and the current instance entity is used to search for each associated candidate instance entity in the preset knowledge base according to the preset association relationship; Calculate the instance similarity between the current instance entity and each of the candidate instance entities, and select the extended instance entity associated with the current instance entity from the candidate instance entities based on the instance similarity. The extended instance entity set is obtained by traversing each instance entity in the instance entity set.
5. The method according to claim 1, characterized in that, The process of conceptually expanding each target instance entity in the target instance entity set to obtain an expanded conceptual entity set includes: Obtain instance relationships and subclass relationships, and expand concepts based on each target instance entity in a preset knowledge base according to the instance relationships to obtain a first set of expanded concept entities; According to the subclass relationship, the concept is extended in the preset knowledge base based on each first extended concept entity in the first extended concept entity set to obtain the second extended concept entity set. The extended concept entity set is obtained based on the first extended concept entity set and the second extended concept entity set.
6. The method according to claim 5, characterized in that, The step of expanding concepts based on the target instance entities in a preset knowledge base according to the instance relationships to obtain a first expanded concept entity set includes: The current target instance entity is determined from the various target instance entities, and the current target instance entity is used to search for the associated first candidate extended concept entities in the preset knowledge base according to the instance relationship; Calculate the first concept similarity between the current target instance entity and each of the first candidate extended concept entities, and select the first extended concept entity corresponding to the current target instance entity from each of the first candidate extended concept entities based on the first concept similarity. By traversing each target instance entity in the target instance entity set, the first extended concept entity set is obtained.
7. The method according to claim 5, characterized in that, The step of extending the concept in the preset knowledge base according to the subclass relationship, based on each of the first extended concept entities in the first extended concept entity set, to obtain the second extended concept entity set includes: The current first extended concept entity is determined from the first extended concept entity set, and the current first extended concept entity is used to search for each associated second candidate extended concept entity in the preset knowledge base according to the subclass relationship; Calculate the degree of second concept similarity between the current first extended concept entity and each of the second candidate extended concept entities, and select the second extended concept entity corresponding to the current first extended concept entity from the second candidate extended concept entities based on the degree of second concept similarity. By traversing through each first extended concept entity in the first extended concept entity set, a second extended concept entity set is obtained.
8. The method according to claim 1, characterized in that, The calculation of instance similarity between each extended instance entity in the extended instance entity set and the associated instance entities in each instance entity, and the calculation of extended instance entity weights using the weights of each instance entity and the instance similarity, to obtain the weights of each extended instance entity, includes: Determine the current extended instance entity from each extended instance entity in the extended instance entity set, and determine the instance entity associated with the current extended instance entity from each instance entity. The associated instance entity weights corresponding to the associated instance entities are obtained from the weights of each instance entity, and the association instance similarity between the current extended instance entity and the associated instance entity is determined from the instance similarity. Calculate the product of the associated instance entity weight and the similarity of the associated instance to obtain the extended instance weight corresponding to the current extended instance entity; By traversing each of the extended instance entities, the weight of each extended instance entity is obtained.
9. The method according to claim 1, characterized in that, The calculation of the conceptual similarity between each extended conceptual entity in the extended conceptual entity set and the associated target instance entities in each target instance entity set, and the calculation of extended conceptual entity weights using the weights of each instance entity, the weights of each extended instance entity, and the conceptual similarity, to obtain the weights of each extended conceptual entity, includes: A first set of extended concept entities and a second set of extended concept entities are determined from the various extended concept entities. The first set of extended concept entities is obtained based on the target instance entity set through a preset instance relationship, and the second set of extended concept entities is obtained based on the first set of extended concept entities through a preset subclass relationship. Calculate the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entity associated with each target instance entity; determine the target instance entity weight associated with each first extended concept entity from the weight of each instance entity and the weight of each extended instance entity; calculate the weight of each first extended concept entity based on the first concept similarity and the weight of the associated target instance entity; and obtain the weight of each first extended concept entity. Calculate the degree of second concept similarity between each second extended concept entity in the second extended concept entity set and the first extended concept entity associated with each first extended concept entity. Determine the weight of the first extended concept entity associated with each second extended concept entity from the weight of each first extended concept entity. Calculate the weight of the second extended concept entity based on the degree of second concept similarity and the weight of the associated first extended concept entity to obtain the weight of each second extended concept entity. The weights of each extended concept entity are obtained based on the weights of each first extended concept entity and the weights of each second extended concept entity.
10. The method according to claim 9, characterized in that, The calculation of the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entities associated with each target instance entity, determining the target instance entity weights associated with each first extended concept entity from the weights of each instance entity and the weights of each extended instance entity, and calculating the weights of each first extended concept entity based on the first concept similarity and the associated target instance entity weights to obtain the weights of each first extended concept entity includes: Determine the current first extended concept entity from each of the first extended concept entities, and determine the current target instance entity associated with the current first extended concept entity from each of the target instance entities; Calculate the similarity between the current first concept entity and the current target instance entity, and determine the current target instance entity weight corresponding to the current target instance entity from the weights of each instance entity and the weights of each extended instance entity; Calculate the product of the current first concept similarity and the current target instance entity weight to obtain the first extended concept entity weight corresponding to the current first extended concept entity; By traversing each of the first extended concept entities, the weights of each of the first extended concept entities are obtained.
11. The method according to claim 9, characterized in that, The calculation of the second concept similarity between each second extended concept entity in the second extended concept entity set and the first extended concept entities associated with each first extended concept entity set, determining the weights of the first extended concept entities associated with each second extended concept entity from the weights of each first extended concept entity, and calculating the weights of the second extended concept entities based on the second concept similarity and the weights of the associated first extended concept entities to obtain the weights of each second extended concept entity includes: Determine the current second extended concept entity from the various second extended concept entities, and determine the current first extended concept entity associated with the current second extended concept entity from the various first extended concept entities; Calculate the similarity between the current second extended concept entity and the current first extended concept entity, and determine the current first extended concept entity weight corresponding to the current first extended concept entity from the weights of each first extended concept entity; Calculate the product of the current second concept similarity and the current first extended concept entity weight to obtain the second extended concept entity weight corresponding to the current second extended concept entity; By traversing each of the second extended concept entities, the weights of each second extended concept entity are obtained.
12. The method according to claim 1, characterized in that, After calculating the instance similarity between each extended instance entity in the extended instance entity set and the associated instance entities in each instance entity set, and using the weights of each instance entity and the instance similarity to calculate the weights of the extended instance entities, the method further includes: Obtain the instance weight decay factor, and perform weight decay on the weights of each instance entity and each extended instance entity based on the instance weight decay factor to obtain the decay weight of each instance entity and the decay weight of each extended instance entity. The step of calculating the extended concept entity weights using the weights of each instance entity, the weights of each extended instance entity, and the concept similarity to obtain the extended concept entity weights includes: The weights of extended concept entities are calculated based on the attenuation weights of each instance entity, the attenuation weights of each extended instance entity, and the concept similarity, to obtain the attenuation weights of each extended concept entity. The step of establishing an interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity includes: Based on the target instance entity set, the extended concept entity set, the decay weight of each instance entity, the decay weight of each extended instance entity, and the decay weight of each extended concept entity, a target interest knowledge graph corresponding to the target user identifier is established.
13. The method according to claim 12, characterized in that, The process of obtaining the instance weight decay factor includes: Obtain the first target time point corresponding to each instance entity and the second target time point corresponding to each extended instance entity. The first target time point is the historical time point when the current instance entity was previously an extended instance entity, and the second target time point is the historical time point when the extended instance entity was previously an extended instance entity. Obtain the current time point, determine a first time interval based on the first target time point and the current time point, and determine a second time interval based on the second target time point and the current time point; When the first time interval is within the preset weight decay time range, the first initial weight decay factor is calculated based on the first time interval and the preset weight decay rate parameter. The first initial weight decay factor is normalized to obtain the instance weight decay factor corresponding to each instance entity. When the second time interval is within the preset weight decay time range, a second initial weight decay factor is calculated based on the second time interval and the preset weight decay rate parameter. The second initial weight decay factor is then normalized to obtain the instance weight decay factor corresponding to each extended instance entity.
14. The method according to claim 1, characterized in that, After establishing the interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity, the method further includes: Obtain the interest knowledge graph corresponding to the target user identifier at each preset time point, and dynamically visualize the interest knowledge graph corresponding to the target user identifier at each preset time point.
15. An information recommendation method, characterized in that, The method includes: Receive information recommendation instructions, wherein the information recommendation instructions carry a user identifier and a query statement; Target keywords are obtained by extracting query keywords from the query statement; Obtain the interest knowledge graph corresponding to the user identifier, wherein the interest knowledge graph is established by the knowledge graph construction method as described in any one of claims 1 to 14; The target interest entity is determined from the interest knowledge graph, and recommendation information is obtained based on the target keywords and the target interest entity. The recommendation information is then returned to the terminal corresponding to the user identifier.
16. A knowledge graph construction device, characterized in that, The device includes: The instance acquisition module is used to obtain the reading text of the target user identifier within a preset time period, perform entity linking based on the reading text, and obtain the instance entity set corresponding to the reading text. An extension module is used to extend instances based on each instance entity in the instance entity set to obtain an extended instance entity set, to obtain a target instance entity set based on the extended instance entity set and the instance entity set, and to extend concepts based on each target instance entity in the target instance entity set to obtain an extended concept entity set. The weight acquisition module is used to acquire the weight of each instance entity. The weight of each instance entity refers to the number of times each instance entity appears in the reading text and the target reading text in the target time period before the preset time period. The instance weight calculation module is used to calculate the instance similarity between each extended instance entity in the extended instance entity set and the associated instance entity in each instance entity, and to calculate the extended instance entity weight using the weight of each instance entity and the instance similarity to obtain the weight of each extended instance entity. The concept weight calculation module is used to calculate the concept similarity between each extended concept entity in the extended concept entity set and the target instance entity associated with each target instance entity. The extended concept entity weight is calculated using the weight of each instance entity, the weight of each extended instance entity, and the concept similarity to obtain the weight of each extended concept entity. The graph building module is used to build an interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the weights of each instance entity, the weights of each extended instance entity, and the weights of each extended concept entity.
17. The knowledge graph construction apparatus according to claim 16, characterized in that, The module obtained in the instance includes: The recognition unit is used to perform entity word recognition based on the reading text to obtain each entity word; The recall unit is used to recall entities from a preset knowledge base based on each entity word, and obtain a candidate entity set corresponding to each entity word. The disambiguation unit is used to perform entity disambiguation based on the candidate entity set corresponding to each entity word, to obtain the entity corresponding to each entity word, and to obtain the instance entity set corresponding to the reading text based on the entity corresponding to each entity word.
18. The knowledge graph construction apparatus according to claim 17, characterized in that, The disambiguation unit is further configured to determine the current entity word from each entity word and obtain the entity text corresponding to the current entity word; input the entity text and the corresponding candidate entity set into the entity disambiguation model, the entity disambiguation model maps the entity text and the corresponding candidate entity set into a vector space respectively to obtain the entity word vector corresponding to the current entity word and the candidate entity vector set corresponding to the candidate entity set, calculate the similarity between the entity word vector and the candidate entity vector in the candidate entity vector set respectively, and determine the current entity corresponding to the current entity word from the candidate entity set based on the similarity.
19. The knowledge graph construction apparatus according to claim 16, characterized in that, The extension module is also used to determine the current instance entity from the instance entity set, and use the current instance entity to search for each associated candidate instance entity in the preset knowledge base according to the preset association relationship; Calculate the instance similarity between the current instance entity and each of the candidate instance entities, and select the extended instance entity associated with the current instance entity from the candidate instance entities based on the instance similarity; traverse each instance entity in the instance entity set to obtain the extended instance entity set.
20. The knowledge graph construction apparatus according to claim 16, characterized in that, The expansion module includes: The first concept expansion unit is used to obtain instance relationships and subclass relationships, and expand the concept based on each target instance entity in the preset knowledge base according to the instance relationships to obtain the first expanded concept entity set. The second concept expansion unit is used to expand the concept in the preset knowledge base according to the subclass relationship based on each first expansion concept entity in the first expansion concept entity set to obtain the second expansion concept entity set. The extended concept obtaining unit is used to obtain the extended concept entity set based on the first extended concept entity set and the second extended concept entity set.
21. The knowledge graph construction apparatus according to claim 20, characterized in that, The first concept expansion unit is further configured to determine the current target instance entity from the various target instance entities, and use the current target instance entity to search for the associated first candidate expansion concept entities in the preset knowledge base according to the instance relationship; Calculate the first concept similarity between the current target instance entity and each of the first candidate extended concept entities. Based on the first concept similarity, select the first extended concept entity corresponding to the current target instance entity from the first candidate extended concept entities. Traverse each target instance entity in the target instance entity set to obtain the first extended concept entity set.
22. The knowledge graph construction apparatus according to claim 20, characterized in that, The second concept expansion unit is further configured to determine the current first expansion concept entity from the first expansion concept entity set, use the current first expansion concept entity to search for each associated second candidate expansion concept entity in the preset knowledge base according to the subclass relationship; calculate the second concept similarity between the current first expansion concept entity and each second candidate expansion concept entity, select the second expansion concept entity corresponding to the current first expansion concept entity from each second candidate expansion concept entity based on the second concept similarity; and traverse each first expansion concept entity in the first expansion concept entity set to obtain the second expansion concept entity set.
23. The knowledge graph construction apparatus according to claim 16, characterized in that, The instance weight calculation module is further configured to determine the current extended instance entity from each extended instance entity in the extended instance entity set, and determine the instance entity associated with the current extended instance entity from each instance entity; obtain the associated instance entity weight corresponding to the associated instance entity from the weights of each instance entity, and determine the associated instance similarity between the current extended instance entity and the associated instance entity from the instance similarity. Calculate the product of the associated instance entity weight and the similarity of the associated instance to obtain the extended instance weight corresponding to the current extended instance entity; traverse each extended instance entity to obtain the weight of each extended instance entity.
24. The knowledge graph construction apparatus according to claim 16, characterized in that, The concept weight calculation module includes: The determining unit is configured to determine a first set of extended concept entities and a second set of extended concept entities from the various extended concept entities. The first set of extended concept entities is obtained based on the target instance entity set through a preset instance relationship, and the second set of extended concept entities is obtained based on the first set of extended concept entities through a preset subclass relationship. The first calculation unit is used to calculate the first concept similarity between each first extended concept entity in the first extended concept entity set and the target instance entity associated with each target instance entity, determine the target instance entity weight associated with each first extended concept entity from the weight of each instance entity and the weight of each extended instance entity, and calculate the weight of each first extended concept entity based on the first concept similarity and the weight of the associated target instance entity to obtain the weight of each first extended concept entity. The second calculation unit is used to calculate the degree of second concept similarity between each second extended concept entity in the second extended concept entity set and the first extended concept entity associated with each first extended concept entity, determine the weight of the first extended concept entity associated with each second extended concept entity from the weight of each first extended concept entity, and calculate the weight of the second extended concept entity based on the degree of second concept similarity and the weight of the associated first extended concept entity to obtain the weight of each second extended concept entity. The weighting unit is used to obtain the weights of each extended concept entity based on the weights of each first extended concept entity and the weights of each second extended concept entity.
25. The knowledge graph construction apparatus according to claim 24, characterized in that, The first calculation unit is further configured to determine a current first extended concept entity from the various first extended concept entities, and determine a current target instance entity associated with the current first extended concept entity from the various target instance entities; calculate the current first concept similarity between the current first extended concept entity and the current target instance entity, and determine the current target instance entity weight corresponding to the current target instance entity from the weights of the various instance entities and the weights of the various extended instance entities; Calculate the product of the current first concept similarity and the current target instance entity weight to obtain the first extended concept entity weight corresponding to the current first extended concept entity; traverse each first extended concept entity to obtain the weight of each first extended concept entity.
26. The knowledge graph construction apparatus according to claim 24, characterized in that, The second calculation unit is further configured to determine the current second extended concept entity from the various second extended concept entities, and determine the current first extended concept entity associated with the current second extended concept entity from the various first extended concept entities; calculate the current second concept similarity between the current second extended concept entity and the current first extended concept entity, and determine the current first extended concept entity weight corresponding to the current first extended concept entity from the weights of the various first extended concept entities; Calculate the product of the current second concept similarity and the current first extended concept entity weight to obtain the second extended concept entity weight corresponding to the current second extended concept entity; traverse each second extended concept entity to obtain the weight of each second extended concept entity.
27. The knowledge graph construction apparatus according to claim 16, characterized in that, The device further includes: The weight decay module is used to obtain the instance weight decay factor, and to perform weight decay on the weights of each instance entity and the weights of each extended instance entity based on the instance weight decay factor, so as to obtain the decay weight of each instance entity and the decay weight of each extended instance entity. The concept weight calculation module is also used to calculate the weight of extended concept entities based on the decay weight of each instance entity, the decay weight of each extended instance entity, and the concept similarity, so as to obtain the decay weight of each extended concept entity. The graph building module is also used to build a target interest knowledge graph corresponding to the target user identifier based on the target instance entity set, the extended concept entity set, the decay weight of each instance entity, the decay weight of each extended instance entity, and the decay weight of each extended concept entity.
28. The knowledge graph construction apparatus according to claim 27, characterized in that, The weight decay module is further configured to obtain a first target time point corresponding to each instance entity and a second target time point corresponding to each extended instance entity, wherein the first target time point is the historical time point when the current instance entity was previously an extended instance entity, and the second target time point is the historical time point when the extended instance entity was previously an extended instance entity; obtain the current time point, determine a first time interval based on the first target time point and the current time point, and determine a second time interval based on the second target time point and the current time point; When the first time interval is within the preset weight decay time range, the first initial weight decay factor is calculated based on the first time interval and the preset weight decay rate parameter. The first initial weight decay factor is normalized to obtain the instance weight decay factor corresponding to each instance entity. When the second time interval is within the preset weight decay time range, a second initial weight decay factor is calculated based on the second time interval and the preset weight decay rate parameter. The second initial weight decay factor is then normalized to obtain the instance weight decay factor corresponding to each extended instance entity.
29. The knowledge graph construction apparatus according to claim 16, characterized in that, The device further includes: The display module is used to obtain the interest knowledge graph corresponding to the target user identifier at each preset time point, and to dynamically visualize the interest knowledge graph corresponding to the target user identifier at each preset time point.
30. An information recommendation device, characterized in that, The device includes: The instruction receiving module is used to receive information recommendation instructions, which carry a user identifier and a query statement. The extraction module is used to extract query keywords from the query statement to obtain target keywords; The graph acquisition module is used to acquire the interest knowledge graph corresponding to the user identifier, wherein the interest knowledge graph is established by the knowledge graph construction method as described in any one of claims 1 to 14; The recommendation module is used to determine the target interest entity from the interest knowledge graph, obtain recommendation information based on the target keywords and the target interest entity, and return the recommendation information to the terminal corresponding to the user identifier.
31. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 15.
32. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 15.
33. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 15.