Method and System for Constructing User Portrait Based on Knowledge Graph
Through the method based on the knowledge graph, user information is collected and processed to form a structured knowledge system, the problem of low efficiency in user portrait construction in the existing technology is solved, and efficient processing of massive user data and accurate user portrait construction are achieved.
Patent Information
- Application Number
- CN202211640703.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-12-20
AI Technical Summary
The existing user profile technology lacks efficient methods to process massive user behavior data, resulting in insufficient scalability and inability to effectively build user profiles.
A user portrait construction method based on knowledge graph is adopted, and a structured and networked knowledge system is formed through steps such as collecting user information, information extraction, knowledge integration, knowledge processing and graph database preservation, and then a user portrait is built.
It improves the efficiency of processing massive user data, improves the speed and accuracy of user portrait construction, and meets the needs of efficient processing of massive user behavior data.
Smart Images

Figure CN115982379B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and system for constructing a user profile based on a knowledge graph. Background Art
[0002] User Profile is widely used in operations and data analysis, and it is a set of variables describing various user data. The so-called user profile is the user image constructed through data tags. Personalized recommendation, advertising system, event marketing, content recommendation, and interest preference are all applications based on user profiles. By analyzing a large amount of data information, the data is abstracted into tags, and then these tags are used to concretize the user image, and finally a user profile is formed for the above applications. When we want to select the educational user group in the campus environment for refined information push, we first screen out specific groups according to the user profile. The user profile is a complex system, and as the profile model gradually matures, different tags will be designed according to different business scenarios.
[0003] The knowledge graph is mainly used to describe various entities and concepts existing in the real world, and the relationships between them, so it can be considered as a semantic network. From the development process, the knowledge graph is developed on the basis of NLP (Natural Language Processing). The knowledge graph is closely related to NLP. The knowledge graph can be used to query complex associated information more efficiently, understand the user's intention from the semantic level, and improve the search quality. The knowledge graph effectively processes, processes, and integrates the data of complex documents, transforms them into simple and clear triples of "entity-relationship-entity" or "entity-attribute-attribute value", and finally aggregates a large amount of knowledge, so as to achieve rapid response and reasoning of knowledge.
[0004] Currently, the user profile technology is still in the state of "tagging" based on manual operations, with insufficient scalability in user behavior analysis and a lack of an efficient method for constructing user profiles for massive user behavior data. Summary of the Invention
[0005] The present invention provides a method and system for constructing a user profile based on a knowledge graph, aiming to solve the problem of the lack of an efficient method for profiling massive user data currently.
[0006] To solve the above technical problems, the present invention proposes a method and system for constructing a user profile based on a knowledge graph, including the following steps:
[0007] S1: Collect user information, extract a number of keywords, and output a description text about the user. The keywords include the user's registration information, search keywords, and post content.
[0008] S2: Construct a knowledge graph for the user profile, perform information extraction on the description text, classify and then map correspondingly to form triples in the form of entity-relationship-entity or entity-attribute-attribute value. The information extraction includes entity recognition, relationship recognition, and attribute recognition. The entity recognition uses a method that combines an improved convolutional neural network with an improved conditional random field model. The description text is input into the improved convolutional neural network to obtain a feature map, and the feature map is input into the improved conditional random field model for sequence labeling to identify the entities of interest in the description text.
[0009] S3: Perform knowledge fusion on the triples, perform coreference resolution on different triples about the same entity from multiple sources to map to the correct single entity, and use an unsupervised clustering method based on encyclopedic knowledge to disambiguate the synonymous triples representing different entities to solve the ambiguity caused by synonymous triples.
[0010] S4: Perform knowledge processing on the fused triples to form a structured and networked knowledge system. The knowledge processing includes ontology construction, knowledge reasoning, and quality assessment. The ontology construction is based on the common understanding of a domain to extract a general term. The knowledge reasoning obtains new knowledge or conclusions through various methods. The quality assessment quantifies the credibility of the knowledge, and ensures the quality of the knowledge base by discarding the knowledge with lower confidence.
[0011] S5: Use a graph database to store the triples after knowledge processing to form a knowledge graph.
[0012] S6: Mine user features based on the knowledge graph and construct a user profile according to the user features.
[0013] Preferably, the quality assessment quantifies the credibility, which is divided into data in four dimensions: the accuracy, coverage, consistency, and simplicity of the knowledge. The data is input into a preset quantization model to calculate the quantization values respectively. According to the quantization values, marker points are added in the four directions of the x and y axes in a plane rectangular coordinate system, and the area value of the quadrilateral formed by connecting the four marker points is used as the confidence.
[0014] Preferably, the improved convolutional neural network uses dilated convolution to replace the convolution operation of the traditional convolutional neural network and removes the pooling layer in the traditional convolutional neural network.
[0015] Preferably, the improved conditional random field model is based on the set label transfer rules, adds a mask to the transition matrix in advance, and assigns an extremely small value to the illegal label transfer score.
[0016] Preferably, the attribute recognition uses a crawler to crawl the relationship keywords of the identified entity on the network based on the identified entity.
[0017] Preferably, the unsupervised clustering method constructs a large-scale semantic network from Wikipedia and disambiguates according to the encyclopedic semantic knowledge in the semantic network.
[0018] Preferably, the knowledge reasoning adopts a knowledge graph reasoning method based on representation learning, maps entities and the relationships between entities to a vector space, and then establishes logical relationships through operations in the vector space.
[0019] Preferably, the graph database is a Neo4j database.
[0020] A user portrait construction system based on a knowledge graph, characterized by comprising: a data collection module and a knowledge graph module, the knowledge graph module further comprising: an information extraction module, a knowledge fusion module and a knowledge processing module, the data collection module and the knowledge graph module being configured to execute the above-mentioned user portrait construction method based on a knowledge graph.
[0021] The data collection module is used to collect user information and organize it into a description text of the user;
[0022] The information extraction module performs entity recognition, relationship recognition and attribute recognition on the description text to form triples in the form of entity-relationship-entity or entity-attribute-attribute value, and adopts a method combining an improved convolutional neural network and an improved conditional random field model to perform entity recognition on the description text;
[0023] The knowledge fusion module fuses the triples, specifically: performs coreference resolution on different triples about the same entity from multiple sources to map to the correct entity, and disambiguates the synonymous triples representing different entities to solve the ambiguity generated by the synonymous triples;
[0024] The knowledge processing module processes the fused triples, adopts an ontology construction method to extract a general term based on the common understanding of a domain, adopts a knowledge reasoning method to obtain new knowledge or conclusions through various methods, adopts a quality assessment method to quantify the credibility of the knowledge, and ensures the quality of the knowledge base by discarding knowledge with lower confidence.
[0025] Compared with the prior art, the present invention has the following technical effects:
[0026] 1. The user portrait construction method proposed by the present invention uses a knowledge graph for user portrait. The search mode of the knowledge graph is to search for the required content from triples. For multi-hop search, the connection and reasoning ability of the knowledge graph is superior to the Join operation of the relational database, greatly improving the data query efficiency and making the efficiency of using the knowledge graph for massive user portraits greatly improved.
[0027] 2. When the method for constructing a user portrait proposed by the present invention performs entity recognition operations, it uses an extended convolutional neural layer to replace the convolutional operation of the traditional convolutional neural network. It retains the advantage that the convolutional neural network can make full use of the parallel use of GPU resources, eliminates the influence of the subsequent pooling layer in the convolutional layer on the accuracy, increases the receptive field in the convolutional operation, and avoids overfitting during the neural network training process.
[0028] 3. When the method for constructing a user portrait proposed by the present invention performs entity recognition operations, it uses an improved conditional random field, adds a mask to the transition matrix in the conditional random field, and assigns an extremely small value to the illegal label transition score, so that the path containing the illegal path will necessarily have a very low score, which can effectively improve the efficiency of entity recognition.
[0029] 4. The method for constructing a user portrait proposed by the present invention uses a graph database to store the knowledge graph, applies graph theory to store the relationship information between entities, makes full use of the high performance of the graph database in data association relationship queries, makes the speed of data query and analysis faster, and improves the efficiency of portrait construction; the flexible data model of the graph database can adapt to changing business needs, and it can easily add or delete vertices and edges, expand or shrink the graph model, making the construction of the knowledge graph more flexible. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flowchart of the method for constructing a user portrait based on a knowledge graph according to the present invention;
[0031] Figure 2 is a flowchart of constructing a knowledge graph of the method for constructing a user portrait based on a knowledge graph according to the present invention;
[0032] Figure 3 is a schematic diagram of an extended convolution of the method for constructing a user portrait based on a knowledge graph according to the present invention;
[0033] Figure 4 is a schematic diagram of adding a mask of an improved conditional random field model of the method for constructing a user portrait based on a knowledge graph according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present application and with reference to the accompanying drawings.
[0035] Please refer to Figure 1 , the method and system for constructing a user portrait based on a knowledge graph in this embodiment include the following steps:
[0036] S1: Collect user information and organize it into a description text of the user.
[0037] The sources of user information include: the text in the user registration materials within the campus network environment, including basic information such as name, personal signature, major, grade, etc.; the content generated by the user during the use of the application, including text content such as comments, dynamics, diaries, and chat records published; the text that has a connection relationship with the user, such as the content read by the user.
[0038] S2: Perform information extraction on the described text to form triples in the form of entity-relationship-entity or entity-attribute-attribute value. The information extraction includes entity recognition, relationship recognition, and attribute recognition. The entity recognition uses the method of an improved convolutional neural network combined with an improved conditional random field model. The described text is input into the improved convolutional neural network to obtain a feature map, and the feature map is input into the improved conditional random field model for sequence labeling to identify the entities of interest in the described text.
[0039] Information extraction first performs word segmentation and part-of-speech tagging on natural language. For example, for the natural language content in the news browsed: "The Shenzhou-14 astronaut crew will return to the Dongfeng Landing Site in the near future. On the evening of December 1st, the Shenzhou-14 search and rescue recovery mission organized the last full-system comprehensive drill, and the Dongfeng Landing Site has made all preparations to welcome the return of the spacecraft." Entity recognition will identify entities such as Shenzhou-14, astronauts, Dongfeng Landing Site, and December 1st from it. Then, through relationship recognition and attribute recognition operations, several triples are generated according to the corresponding entities. For example, for the entity Shenzhou-14, after crawling network content, the generated triples include: Shenzhou-14 - Abbreviation - Shen-14, Shenzhou-14 - Launch Time - 10:44 on June 5, 2022, Shenzhou-14 - Launch Location - Jiuquan Satellite Launch Center, Shenzhou-14 - Return Location - Dongfeng Landing Site, Shenzhou-14 - Country - China, Shenzhou-14 - Height - 9 meters, Shenzhou-14 - Takeoff Weight - 8 tons, etc. These triples are sent to the knowledge fusion module for the next step of processing.
[0040] S3: Perform knowledge fusion on the triples, perform coreference resolution on different triples about the same entity from multiple sources to map to the correct single entity, and use an unsupervised clustering method based on encyclopedic knowledge to disambiguate homonymous triples representing different entities to resolve the ambiguity generated by homonymous triples.
[0041] Encyclopedic websites usually assign a separate page to each entity, which includes hyperlinks pointing to other entity pages. The encyclopedic knowledge model precisely uses this link relationship to calculate the similarity between entities. A large-scale semantic network can be constructed from Wikipedia, and disambiguation is performed according to the encyclopedic semantic knowledge in the semantic network.
[0042] S4: Process the fused triples to form a structured and networked knowledge system. The knowledge processing includes ontology construction, knowledge reasoning, and quality assessment. The ontology construction is based on the common understanding of a domain to extract a general term. The knowledge reasoning obtains new knowledge or conclusions through various methods. The quality assessment quantifies the credibility of knowledge and ensures the quality of the knowledge base by discarding knowledge with low confidence.
[0043] The quality assessment quantifies the credibility, which is divided into four dimensions of data: the accuracy, coverage rate, consistency, and simplicity of knowledge. The data is input into a preset quantization model to calculate the quantization values respectively. According to the quantization values, marked points are added in the four directions of the x and y axes in a plane rectangular coordinate system. The area value of the quadrilateral formed by connecting the four marked points is used as the confidence, and the confidence of the knowledge is positively correlated with the area value of the quadrilateral.
[0044] Among them, for the accuracy evaluation, the relevant knowledge is compared with the knowledge stored in the standard database or knowledge graph, and the offset percentage from the standard knowledge is calculated. The accuracy data a is 1 - offset percentage, and the final accuracy score is 5×a, so that the value range of the accuracy score is 0 < ≤5 to simplify the calculation of the quadrilateral area. The coverage rate score is evaluated through expert sampling inspection. By comparing whether a specific entity has the common attributes and relationships of its similar entities, it is determined whether the relevant knowledge of the entity is complete, and the coverage rate score is assigned according to the evaluation opinion. The value range of the coverage rate score is (0, 5]. The consistency examines whether the expression of knowledge is consistent. Among the existing knowledge, there are knowledge groups with contradictions. The consistency score of the knowledge group with contradictions is assigned 1, and the consistency score of the knowledge without contradictions is assigned 5. The quantization model does not change its score value. The simplicity evaluation of knowledge is measured by the coincidence ratio of attributes, entities, and relationships related to the knowledge field to reduce the weight of data elements irrelevant to the field of the knowledge graph. The greater the coincidence ratio of the relevant knowledge with the attributes, entities, and relationships related to this field, the higher its simplicity score. Similar to the accuracy evaluation method, multiplying the coincidence ratio value by 5 is the simplicity score.
[0045] S5: Use a graph database to store the triples after knowledge processing to form a knowledge graph.
[0046] S6: Mine user features based on the knowledge graph and construct a user portrait according to the user features.
[0047] The implementation methods of steps S2 to S5 are as Figure 2As shown in the figure, when constructing a knowledge graph, information extraction must be carried out first. Information extraction includes entity recognition, relationship recognition, and attribute recognition, and triples of "entity-relationship-entity" or "entity-attribute-attribute value" are sorted out. The purpose of entity recognition is to identify the named entities in the descriptive text, and to identify the entity boundaries and determine the entity types. The named entity recognition method is adopted to search for entities with describable meanings in a sentence.
[0048] In this embodiment, an improved CNN convolutional layer is used as the encoder and combined with an improved conditional random field for named entity recognition. Among them, dilated convolution is used in the improved CNN convolutional layer to replace the convolution operation of the traditional convolutional neural network, and the pooling layer in the traditional convolutional neural network is removed. Please refer to Figure 3 , which is a schematic diagram of dilated convolution with a dilation factor of 2 in this embodiment. If the dilation factor is 2, the number of inserted holes is 2-1. By inserting holes between consecutive elements to expand the input, the area covered by the input image can be expanded without pooling, and a wider field of view can be provided with the same computational cost.
[0049] In the conditional random field, a weight λ k is assigned to each element f of the transition matrix k . Given a sentence s, s can correspond to several tag sequences l. Therefore, by adding the weights of all keywords, each tag sequence l can be scored, and the formula is as follows:
[0050]
[0051] In the formula, n is the length of the descriptive text sentence s, and m is the number of feature functions.
[0052] Figure 4 is a schematic diagram of adding a mask to the transition matrix in this embodiment. The improved conditional random field model, based on the set label transition rules, pre-adds masks to the elements in the transition matrix whose scores are lower than a set value, and assigns the transition scores of the elements lower than this set value to a minimum value, so that the conditional random field model will not select the paths containing these elements when calculating the path scores. In Figure 4 , the elements with low scores include a13, a15, a17, a23, a25, a27, a32, a35, a37, a42, a43, a47, a53, a57, a65, a75. Then masks are added to the above elements, so that the conditional random field model no longer considers the above elements when calculating the path scores, avoiding the influence of illegal paths on the efficiency of the conditional random field model.
[0053] The attribute recognition is based on the recognized entities. A crawler is used to crawl the relationship keywords of the recognized entities on the network, which is used to construct an attribute list for the entities and attach attribute values to the entities. The goal of attribute recognition is to collect the attribute information of specific entities from different information sources. For example, for a public figure, information such as their nickname, birthday, nationality, educational background, etc. can be obtained from publicly available information on the network. Attribute recognition technology can gather this information from multiple data sources to achieve a complete delineation of the entity's attributes.
[0054] After entity recognition, the resulting text is a series of discrete named entities. To obtain semantic information, it is also necessary to extract the association relationships between entities from relevant corpora. By linking entities (concepts) through association relationships, a networked knowledge structure can be formed. The goal of relationship recognition is to solve the problem of entity semantic linking. The basic information of a relationship includes the parameter type and the tuple pattern that satisfies this relationship. Entity relationship recognition based on joint reasoning is adopted. The relationship recognition method of joint reasoning is the Markov logic network, which combines the Markov network with first-order logic for statistical relational learning, and at the same time incorporates reasoning into OIE (Open Information Extraction).
[0055] After passing through the information extraction module, the descriptive text is organized into a number of triples. The relationships between these triples are flat, lacking hierarchy and logic, and there is also a large amount of redundant and incorrect information. To solve these problems, it is necessary to perform knowledge fusion on the triples formed by information extraction. First, according to the given entity referential terms, a set of candidate entity objects are selected from the knowledge base, and then the referential terms are linked to the correct entity objects through similarity calculation. The specific method is as follows: entity referential terms are obtained through entity recognition from the text; entity disambiguation and coreference resolution are performed to determine whether the same-named entities in the knowledge base have different meanings from it and whether there are other named entities in the knowledge base that have the same meaning as it; after confirming the correct entity object corresponding in the knowledge base, the entity referential term is linked to the corresponding entity in the knowledge base.
[0056] Among them, entity disambiguation is a technology used to solve the problem of ambiguity caused by the same-named entities. Through entity disambiguation, entity links can be accurately established according to the current context. An unsupervised clustering method is adopted to achieve word sense disambiguation. The unsupervised clustering method constructs a large-scale semantic network from Wikipedia and performs disambiguation based on the encyclopedic semantic knowledge in the semantic network. A disambiguation model is established using word sense annotated corpora, and this disambiguation model is used for entity disambiguation. Coreference resolution is used to solve the problem of multiple referents corresponding to the same entity object. In a conversation, multiple referents may point to the same entity object. Using coreference resolution technology, these referential terms can be associated (merged) to the correct entity object.
[0057] When constructing a knowledge graph, knowledge input can be obtained from third-party knowledge base products or existing structured data. During the construction process of the knowledge graph, an important source of high-quality knowledge is the relational database of an enterprise or institution. To integrate this structured historical data into the knowledge graph, the Resource Description Framework can be used as a data model to convert the data in the relational database into triple data.
[0058] During the construction of the knowledge graph, through information extraction, knowledge elements such as entities, relationships, and attributes are extracted from the original corpus, and through knowledge fusion, the ambiguity between entity reference terms and entity objects is eliminated to obtain a series of basic factual expressions. To obtain a structured and networked knowledge system, knowledge processing is still required. Knowledge processing includes operations such as ontology construction, knowledge reasoning, and quality assessment.
[0059] The ontology construction is used to acquire, describe, and represent the knowledge of a related field, provide a common understanding of the knowledge in this field, determine the commonly recognized vocabulary in the field, provide the specific concept definitions and relationships between concepts in this field, provide the activities that occur in this field, as well as the main theories and basic principles in this field, so as to achieve the effect of human-computer communication. In this embodiment, the ontology construction adopts a seven-step method, mainly used for the construction of domain ontology. It includes the following steps: determining the professional field and scope of the ontology; examining the possibility of reusing existing ontologies; listing the important terms in the ontology; defining classes and the hierarchical system of classes. Feasible methods to improve the hierarchical system include: top-down method, bottom-up method, and comprehensive method; defining the attributes of classes; defining the facets of attributes; creating instances.
[0060] The knowledge reasoning adopts a knowledge graph reasoning method based on representation learning, maps entities and the relationships between entities into a vector space, and then establishes logical relationships through operations in the vector space. By inferring the associated relationships between entities, new knowledge is automatically generated to supplement the missing facts and improve the knowledge graph. The objects of knowledge reasoning are not limited to the relationships between entities, but can also be the attribute values of entities and the conceptual hierarchical relationships of ontologies. For example, attribute value reasoning can infer the age attribute of an object entity based on the birthday attribute of the object entity, and conceptual hierarchical relationship reasoning can infer new triple data based on the hierarchical progressive relationships of two or more triples.
[0061] The constructed knowledge graph may have some errors, mainly concentrated in the hypernym and hyponym problems, attribute problems, and logical problems of triples. The hypernym and hyponym problems refer to the emergence of circular structures in the graph. Generally speaking, a knowledge graph is a tree-like structure. If a circular structure appears, it needs to be accessed and corrected. The attribute problems refer to the deviation of entity attributes from the normal range of the attribute values. The logical problems refer to the logic between relationships not conforming to objective facts. Therefore, it is necessary to evaluate the quality of the knowledge graph, including knowledge graph completion and knowledge graph error detection.
[0062] The methods adopted in this embodiment for quality evaluation include the consistency check method and the comparison check method based on external knowledge. The consistency check method detects knowledge conflicts in the knowledge graph through detection rules pre-established by experts to discover knowledge quality problems. The comparison check method based on external knowledge uses a high-quality external knowledge source with a high overlap with the target knowledge graph as the reference data to perform quality detection on the target knowledge graph.
[0063] During the application process of the knowledge graph, the knowledge graph can be updated as it is used. After adding or updating entities, relationships, and attribute values, it is necessary to update the data layer. When performing update operations, issues such as the reliability of data sources and data consistency need to be considered, and facts and attributes with high frequencies in each data source are selected for quality evaluation and then added to the knowledge graph.
[0064] Finally, the triple information after quality evaluation is saved using a graph database to complete the construction of the knowledge graph. The graph database is the Neo4j database. The query performance of the Neo4j database is fully utilized, and the query performance will not decline as the data volume increases. The Neo4j database also has flexibility in design, with the natural stretching characteristics of the graph data structure and its unstructured data format, enabling the database to more flexibly adapt to changes in business requirements.
[0065] A user portrait construction system based on a knowledge graph includes: a data collection module and a knowledge graph module. The knowledge graph module further includes: an information extraction module, a knowledge fusion module, and a knowledge processing module. The data collection module and the knowledge graph module are configured to perform the above-mentioned user portrait construction method based on the knowledge graph.
[0066] The data collection module is used to collect user information and organize it into a description text of the user.
[0067] The information extraction module performs entity recognition, relationship recognition, and attribute recognition on the description text to form triples in the form of entity-relationship-entity or entity-attribute-attribute value. The method of combining an improved convolutional neural network and an improved conditional random field model is used to perform entity recognition on the description text.
[0068] The knowledge fusion module fuses the triples, specifically: performing coreference resolution on different triples about the same entity from multiple sources to map them to the correct single entity, and resolving the ambiguity generated by the homonymous triples representing different entities to disambiguate the homonymous triples.
[0069] The knowledge processing module processes the fused triples. It uses ontology construction methods to extract a general term based on the common understanding of a field, uses knowledge reasoning methods to obtain new knowledge or conclusions through various methods, uses quality assessment methods to quantify the credibility of the knowledge, and ensures the quality of the knowledge base by discarding the knowledge with lower confidence.
[0070] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.
Claims
1. Method and system for constructing user portrait based on knowledge graph, characterized in that, it includes the following steps: S1: Collect user information, extract a number of keywords, and output a description text about the user. The keywords include the user's registration information, search keywords, and post content; S2: Construct a knowledge graph of the user portrait, perform information extraction on the description text, classify and then map correspondingly to form triples in the form of entity-relationship-entity or entity-attribute-attribute value. The information extraction includes entity recognition, relationship recognition, and attribute recognition. The entity recognition uses a method of combining an improved convolutional neural network with an improved conditional random field model. Input the description text into the improved convolutional neural network to obtain a feature map, and input the feature map into the improved conditional random field model for sequence labeling to identify the entities of interest in the description text. Among them, the improved convolutional neural network uses dilated convolution to replace the convolution operation of the traditional convolutional neural network and removes the pooling layer in the traditional convolutional neural network; the improved conditional random field model adds a mask to the transition matrix in advance based on the set label transfer rules and assigns an extremely small value to the illegal label transfer score; S3: Perform knowledge fusion on the triples, perform coreference resolution on different triples about the same entity from multiple sources to map to the correct entity, and use an unsupervised clustering method based on encyclopedic knowledge to disambiguate the triples with the same name representing different entities to solve the ambiguity caused by the triples with the same name; S4: Perform knowledge processing on the fused triples to form a structured and networked knowledge system. The knowledge processing includes ontology construction, knowledge reasoning, and quality assessment. The ontology construction is based on the common understanding of a field and extracts a general term. The knowledge reasoning obtains new knowledge or conclusions through various methods. The quality assessment quantifies the credibility of the knowledge and ensures the quality of the knowledge base by discarding the knowledge with low confidence; S5: Use a graph database to store the triples after knowledge processing to form a knowledge graph; S6: Mine user features based on the knowledge graph and construct a user portrait according to the user features.
2. The method for constructing a user portrait based on a knowledge graph according to claim 1, characterized in that, the quality assessment quantifies the credibility, divides it into data in four dimensions of the accuracy, coverage, consistency, and simplicity of the knowledge, inputs them into a preset quantization model to calculate the quantization values respectively, and adds marker points in the four directions of the x and y axes in a plane rectangular coordinate system according to the quantization values. The area value of the quadrilateral formed by connecting the four marker points is used as the confidence.
3. The method for constructing a user portrait based on a knowledge graph according to claim 1, characterized in that, the attribute recognition crawls the relationship keywords of the recognized entity on the network using a crawler based on the recognized entity.
4. The method for constructing a user portrait based on a knowledge graph according to claim 1, characterized in that, The unsupervised clustering method constructs a large-scale semantic network from Wikipedia and performs disambiguation based on the semantic knowledge in the semantic network.
5. The method for constructing a user profile based on a knowledge graph according to claim 1, wherein, the knowledge reasoning adopts a knowledge graph reasoning method based on representation learning, maps entities and the relationships between entities to a vector space, and then establishes logical relationships through operations in the vector space.
6. The method for constructing a user profile based on a knowledge graph according to claim 1, wherein, the graph database is a Neo4j database.
7. A system for constructing a user profile based on a knowledge graph, wherein, comprising: a data collection module and a knowledge graph module, the knowledge graph module further includes: an information extraction module, a knowledge fusion module, and a knowledge processing module, and the data collection module and the knowledge graph module are configured to execute the method for constructing a user profile based on a knowledge graph according to any one of claims 1-6.
Citation Information
Patent Citations
A method and system for constructing a health knowledge graph
CN109669994A
User portrait prediction method based on multi-source transboundary data fusion
CN114238758A