A knowledge graph construction and information retrieval method, device, equipment and medium
Patent Information
- Application Number
- CN202211607298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-12-14
AI Technical Summary
[0031]由以上可见,应用本公开实施例提供的方案构建知识图谱时,首先获得文本描述的人物关联事件的事件类别,对文本进行实体识别,得到针对人物的目标实体以及目标实体的实体属性,然后基于事件类别确定实体间关系的候选关系模式,这样可以基于候选关系模式、目标实体、实体属性以及目标实体所属的目标文本单元,确定目标实体之间的实体关系,最终可以根据目标实体以及实体关系构建知识图谱,从而能够基于知识图谱进行快速、准确的信息检索。
Smart Images

Figure CN115905573B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and more particularly to the field of information retrieval technology. Background Technology
[0002] Currently, the amount of information on the internet is growing exponentially, including a massive amount of information related to people, such as their job titles, event attendance, appointments, and dismissals. Users often want to know various information about specific individuals. Clearly, searching for information about people is a common need in daily life.
[0003] In existing technologies, information retrieval for individuals is generally based on text-based search. For example, a database of information about individuals is built based on text related to them. When a user has a search request, the database is searched based on the user's input text to obtain the final search results. To ensure the real-time nature and accuracy of the search results, this database needs to be continuously updated manually. Summary of the Invention
[0004] This disclosure provides a knowledge graph construction, information retrieval method, apparatus, device, and medium.
[0005] According to one aspect of this disclosure, a method for constructing a knowledge graph is provided, comprising:
[0006] Perform semantic analysis on the text to obtain the event categories of the events associated with the person described in the text;
[0007] Entity recognition is performed on the text to obtain the target entity for the person and the entity attributes of the target entity;
[0008] Based on the event categories, candidate relationship patterns between entities are determined;
[0009] Based on the candidate relationship patterns, target entities, entity attributes, and target text units to which the target entities belong, the entity relationships between target entities are determined.
[0010] Construct a knowledge graph based on the target entity and entity relationships.
[0011] According to another aspect of this disclosure, an information retrieval method is provided, comprising:
[0012] Perform word segmentation on the searched text;
[0013] From the obtained word segments, determine the searched individuals included in the search text and the search intent targeting those individuals;
[0014] Based on the search target and search intent, a graph walk is used to search in the knowledge graph to obtain search results for the search target. The knowledge graph is a knowledge graph constructed according to the aforementioned knowledge graph method.
[0015] According to another aspect of this disclosure, a knowledge graph construction apparatus is provided, comprising:
[0016] The event category acquisition module is used to perform semantic analysis on the text to obtain the event categories of events associated with the person described in the text;
[0017] An entity recognition module is used to perform entity recognition on the text to obtain the target entity for the person and the entity attributes of the target entity;
[0018] The candidate relationship pattern determination module is used to determine candidate relationship patterns between entities based on the event category.
[0019] The entity relationship determination module is used to determine the entity relationship between target entities based on the candidate relationship pattern, target entity, entity attribute, and target text unit to which the target entity belongs;
[0020] The knowledge graph construction module is used to build a knowledge graph based on the target entity and entity relationships.
[0021] According to another aspect of this disclosure, an information retrieval device is provided, comprising:
[0022] The word segmentation module is used to segment the retrieved text into words.
[0023] The information determination module is used to determine, from the obtained word segmentation, the search person included in the search text and the search intent for the search person;
[0024] The retrieval module is used to perform a retrieval in the knowledge graph based on the retrieval person and the retrieval intent, and to obtain retrieval results for the retrieval person. The knowledge graph is a knowledge graph constructed according to the aforementioned knowledge graph construction method.
[0025] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0026] At least one processor; and
[0027] A memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the aforementioned knowledge graph method or information retrieval method.
[0029] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the aforementioned knowledge graph method or information retrieval method.
[0030] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, is the aforementioned knowledge graph method or information retrieval method.
[0031] As can be seen from the above, when constructing a knowledge graph using the solution provided in the embodiments of this disclosure, the event categories of the events associated with the person described in the text are first obtained, entity recognition is performed on the text to obtain the target entity for the person and the entity attributes of the target entity, and then candidate relationship patterns between entities are determined based on the event categories. In this way, the entity relationships between target entities can be determined based on the candidate relationship patterns, target entities, entity attributes, and the target text units to which the target entities belong. Finally, a knowledge graph can be constructed based on the target entities and entity relationships, thereby enabling fast and accurate information retrieval based on the knowledge graph.
[0032] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0033] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0034] Figure 1 A flowchart illustrating the first knowledge graph construction method provided in this embodiment of the disclosure;
[0035] Figure 2 A schematic diagram of a knowledge graph provided in an embodiment of this disclosure;
[0036] Figure 3 A flowchart illustrating the second knowledge graph construction method provided in this embodiment of the disclosure;
[0037] Figure 4 A schematic diagram illustrating an entity and attribute determination process provided in an embodiment of this disclosure;
[0038] Figure 5 A flowchart illustrating the third knowledge graph construction method provided in this embodiment of the disclosure;
[0039] Figure 6 A flowchart illustrating the first information retrieval method provided in this embodiment of the disclosure;
[0040] Figure 7 A flowchart illustrating the second information retrieval method provided in this embodiment of the disclosure;
[0041] Figure 8 A schematic diagram of the structure of a knowledge graph construction device provided in an embodiment of this disclosure;
[0042] Figure 9 This is a schematic diagram of the structure of an information retrieval device provided in an embodiment of the present disclosure;
[0043] Figure 10 This is a block diagram of an electronic device used to implement the knowledge graph construction method or information retrieval method of the embodiments of this disclosure. Detailed Implementation
[0044] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0045] First, the subject responsible for implementing the solution provided in the embodiments of this disclosure will be described.
[0046] The implementation subject of the solution provided in this disclosure is any electronic device with data processing, data storage and other functions.
[0047] The knowledge graph construction scheme provided in the embodiments of this disclosure will be described in detail below.
[0048] See Figure 1 This is a flowchart illustrating the first knowledge graph construction method provided in this embodiment of the present disclosure. The method includes the following steps S101-S105.
[0049] Step S101: Perform semantic analysis on the text to obtain the event categories of the events associated with the people described in the text.
[0050] In one implementation, the event category of a person-related event described in a text can be obtained based on a pre-trained event category prediction model.
[0051] Specifically, by inputting text into the model, the model outputs a text description of the events associated with the people in the text.
[0052] The embodiments disclosed herein do not limit the specific type or architecture of the above model. For example, the above model may be a BilSTM (Bi-directional Long Short-Term Memory networks) model, an LSA (Latent semantic analysis) model, an RNN (Rcurrent Neural Networks) model, or a convolutional neural network model based on the ResNet-50 (Residual Neural Network-50) architecture.
[0053] In one scenario, the above model could be an ERNIE (Enhanced Language Representation with Informative Entities) model.
[0054] Specifically, the above model can be a miniaturized ERNIE tiny model. ERNIE models often adopt an architecture based on 12 Transformer layers. By distilling and extracting the Transformer layer weights of the above ERNIE model, an ERNIE tiny model based on a 3-Transformer layer architecture can be obtained.
[0055] This simplifies the model's architecture and can significantly improve its performance with almost no impact on the model's effectiveness.
[0056] In another implementation, word segmentation vectors can be obtained based on the text. Then, based on the obtained word segmentation vectors and preset categories of related events, the event categories of the related events described in the text can be obtained. Detailed implementation methods are described in subsequent steps A-C, and will not be elaborated here.
[0057] Step S102: Perform entity recognition on the text to obtain the target entity for the person and the entity attributes of the target entity.
[0058] The target entity mentioned above can be a person, place, time, organization, activity, data, festival or anniversary, etc. The attributes of the target entity can be any information about the target entity, such as the person's name, the location of the organization, or the time of the activity.
[0059] In one implementation, the target entity for a person and the entity attributes of the target entity can be obtained based on a pre-trained entity discrimination model.
[0060] Specifically, by inputting text into the model, you can obtain the target entity for the person and the entity attributes of the target entity as output by the model.
[0061] The embodiments disclosed herein do not limit the specific type or architecture of the above model. For example, the above model may be a CNN model, a BERT model, etc.
[0062] In one scenario, the aforementioned models may include models based on the ERNIE+crf (conditional random field) architecture and models based on the GRU (Gated Recurrent Unit)+crf architecture. Detailed descriptions of these models will follow later. Figure 3 Steps S302 and S303 in the illustrated embodiment will not be described in detail here.
[0063] In another implementation, the text can be segmented to obtain text units that are suspected to contain entities. These text units are then subjected to word segmentation, and the target entity and its attributes are identified based on the segmented words. Detailed implementation methods will be provided later. Figure 3 Step S303 in the illustrated embodiment will not be described in detail here.
[0064] In one embodiment of this disclosure, after obtaining the target entity, normalization processing of the target entity can be performed based on algorithms such as NLPC (Natural Language Processing and Chinese Computing) and CPCA. This can achieve target entity disambiguation and improve the quality of the extracted target entities.
[0065] The normalization process mentioned above may include: time entity normalization, address entity normalization, and name entity completion normalization, etc.
[0066] Step S103: Based on the event category, determine the candidate relationship patterns between entities.
[0067] In this step, the relationship patterns corresponding to event categories can be determined as candidate relationship patterns between entities according to a preset correspondence. These preset correspondences can be pre-set by staff.
[0068] The above candidate relation schema can be understood as a relation format containing content to be determined.
[0069] Specifically, the aforementioned candidate relationship pattern can consist of a relationship subject, a relationship object, and a candidate relationship. It can be either "relationship subject" + "candidate relationship" + "relationship object" or "relationship object" + "candidate relationship" + "relationship subject". Among them, the relationship subject and the relationship object are the aforementioned content to be determined.
[0070] For example, the above candidate relationship pattern could be "relationship subject" + "attendance" + "relationship object". In this case, if the relationship subject is determined to be person P1 and the relationship object is activity E1, the relationship obtained based on the above candidate relationship pattern can be understood as: person P1 attended activity E1.
[0071] For example, the above candidate relationship pattern could be "relationship object" + "CEO" + "relationship object". In this case, if the relationship subject is determined to be company C1 and the relationship object is person P2, the relationship obtained based on the above candidate relationship pattern can be understood as: the CEO of company C1 is person P2.
[0072] Step S104: Determine the entity relationships between target entities based on candidate relationship patterns, target entities, entity attributes, and the target text units to which the target entities belong.
[0073] Text can be split into text units. After splitting, text can have one text unit or multiple text units.
[0074] Specifically, based on the target text unit to which the target entity belongs, the target candidate relationship pattern corresponding to the target text unit can be determined, and then the entity relationship between the target entities can be determined according to the target candidate relationship pattern.
[0075] For example, if the target text unit to which the target entities E1, E2, and E3 belong is T1, then based on the position of T1 in the text, the target candidate relation pattern corresponding to T1 can be determined as: "relation subject" + "visit" + "relation object". Thus, the relation subject E1 and the relation object E2 can be determined from the target entities E1, E2, and E3. Therefore, the relationship between the target entities can be obtained according to the target candidate relation pattern: E1 visited E2.
[0076] In another implementation, the relation subject and relation object in the target entity within the candidate relation pattern can be determined based on the target characters (excluding the target entity and its attributes) in the target text unit to which the target entity belongs. Then, based on the relation subject, relation object, and candidate relation pattern, entity relationships between the target entities are generated. For detailed implementation methods, see [link to implementation details]. Figure 3 Steps S302 and S303 in the illustrated embodiment will not be described in detail here.
[0077] Step S105: Construct a knowledge graph based on the target entity and entity relationships.
[0078] In one implementation, a new knowledge graph can be constructed based on the target entity and entity relationships.
[0079] Specifically, a knowledge graph can be obtained by using target entities as nodes and establishing directed edges between nodes based on the entity relationships between target entities.
[0080] The following is combined Figure 2 To provide a more intuitive introduction.
[0081] See Figure 2 This is a schematic diagram of a knowledge graph provided in an embodiment of this disclosure.
[0082] In the knowledge graph, nodes such as Person P2, Company C1, Location L1, and Activity E1 all correspond to target entities, and directed edges such as CEO, Visit, and Attendance between nodes correspond to relationships between entities.
[0083] For example, if node P2 has a directed edge pointing to location L1, and the relationship corresponding to the directed edge is "visit", then it can be seen that the entity relationship between person P2 and location L1 is: person P2 visited location L1.
[0084] In another implementation, if a knowledge graph already exists locally, the existing knowledge graph can be updated based on the target entity and entity relationships.
[0085] Specifically, it can be determined whether there are nodes representing the target entity in the aforementioned knowledge graph, and then the existing knowledge graph is updated according to the determination result in the following manner:
[0086] For the first target entity that does not have a node representation in the knowledge graph, the first target entity and its corresponding entity relationship can be added to the knowledge graph according to the above implementation method, thereby updating the existing knowledge graph.
[0087] For the second target entity represented by nodes that already exist in the knowledge graph, new directed edges between the second target entity and other entities can be added based on the entity relationships between the second target entities, or existing directed edges can be updated based on the aforementioned entity relationships, thereby updating the existing knowledge graph.
[0088] As can be seen from the above, when constructing a knowledge graph using the solution provided in the embodiments of this disclosure, the event categories of the events associated with the person described in the text are first obtained, entity recognition is performed on the text to obtain the target entity for the person and the entity attributes of the target entity, and then candidate relationship patterns between entities are determined based on the event categories. In this way, the entity relationships between target entities can be determined based on the candidate relationship patterns, target entities, entity attributes, and the target text units to which the target entities belong. Finally, a knowledge graph can be constructed based on the target entities and entity relationships, thereby enabling fast and accurate information retrieval based on the knowledge graph.
[0089] Furthermore, it can be seen that the solution provided in this embodiment of the present disclosure, when constructing a knowledge graph, first obtains the event categories of events associated with people described in the text, and then determines candidate relationship patterns between entities based on the event categories, thereby determining the entity relationships between target entities based on the candidate relationship patterns. That is, candidate relationship patterns are determined before determining entity relationships, and entity relationships are determined based on the previously obtained candidate relationship patterns. This allows for the purposeful determination of entity relationships using candidate relationship patterns as a benchmark, reducing the randomness in determining entity relationships compared to directly determining entity relationships from the entire text, and improving the efficiency of entity relationship determination.
[0090] The differences between the solutions provided in this disclosure and the prior art will be compared and explained below.
[0091] In existing technologies, information retrieval for people is generally based on text-based search. For example, a database of information about people is built based on text information related to them. When a user has a search request, the database is searched based on the user's input text to obtain the final search results.
[0092] To ensure the timeliness and accuracy of the search results, the aforementioned database of personal information needs to be continuously updated manually. For example, staff need to add new personal information, modify changed personal information, and delete incorrect personal information.
[0093] In the solutions provided in this disclosure, as long as the text as raw material is obtained, the electronic device can use the above solution to construct a knowledge graph or conveniently update an existing knowledge graph without the need for manual updating of the knowledge graph.
[0094] As can be seen from the above, the solution provided by the present disclosure embodiments can conveniently update the knowledge graph compared with the prior art, thereby saving a lot of manual maintenance costs.
[0095] The following steps A-C will describe another implementation of the aforementioned method for obtaining the event category of the person-related event described in the text.
[0096] Step A: Perform word segmentation on the text and vectorize the segmented words to obtain word vectors.
[0097] This disclosure does not limit the methods of word segmentation and vectorization described above; the following examples illustrate these methods.
[0098] Word segmentation: For example, word segmentation algorithms such as maximum matching, shortest path, and bidirectional matching can be used to segment text.
[0099] Vectorization: For example, the segmented words can be encoded using encoding algorithms such as One-Hot encoding and Word2vec to obtain segmented word vectors.
[0100] Step B: Based on the obtained word segmentation vectors, predict the probability that the person-related events described in the text belong to the preset category of person-related events.
[0101] In one implementation, for each preset category of person-related events, the probability that the person-related events described in the text belong to that category can be predicted based on the obtained word segmentation vector.
[0102] Specifically, the probabilities mentioned above can be predicted based on the similarity between the word segmentation vectors and the category vectors corresponding to the categories of events associated with the individuals. The category vectors corresponding to the categories can be set by staff based on experience.
[0103] In another implementation, candidate categories can be selected from preset categories of person-related events based on the obtained word segmentation vectors, and then for each candidate category, the probability that the person-related event described in the text belongs to that candidate category can be predicted.
[0104] Among them, candidate categories can be selected from the preset categories of person-related events based on the semantics of the word segmentation vectors and the preset correspondence between the length of the word segmentation vectors and the candidate categories.
[0105] The method for predicting the probability that a person-related event described in a text belongs to a candidate category can be obtained based on the previous implementation method, the only difference being the category, which will not be elaborated here.
[0106] Step C: Based on the predicted probabilities, obtain the event categories of the person-related events described in the text from the preset categories of task-related events.
[0107] Specifically, the above event categories can be obtained based on the predicted probabilities in the following ways.
[0108] In one implementation, a preset number of categories with the highest probability can be determined as the event categories of the person-related events described in the text.
[0109] In another implementation, the category whose probability is greater than a preset threshold can be determined as the event category of the person-related event described in the text.
[0110] As can be seen, after segmenting the text and vectorizing the segmented words to obtain segmented vectors, the probability of the events related to the people described in the text belonging to the category of preset events can be predicted based on the obtained segmented vectors. Thus, the event category of the events related to the people described in the text can be quickly and accurately determined based on the predicted probability.
[0111] In one embodiment of this disclosure, after the knowledge graph is constructed, it can be copied through data snapshots, thereby enabling rapid expansion of the knowledge graph and meeting users' incremental information retrieval needs.
[0112] exist Figure 1 Based on the illustrated embodiment, when determining the target entity and its entity attributes, the text can be segmented to obtain text units that may contain the entity. Then, the obtained text units are segmented into words, and the target entity and its entity attributes are identified based on the segmented words. In view of the above, this disclosure provides a second method for constructing a knowledge graph.
[0113] See Figure 3 This is a flowchart illustrating the second knowledge graph construction method provided in this embodiment of the present disclosure. The method includes the following steps S301-S306.
[0114] Step S301: Perform semantic analysis on the text to obtain the event categories of the events associated with the people described in the text.
[0115] The above step S301 is the same as the aforementioned Figure 1 Step S101 is the same in the illustrated embodiment, and will not be repeated here.
[0116] Step S302: Divide the text to obtain text units that are suspected to contain entities.
[0117] The text units mentioned above are text units that are suspected to include entities, and can also be referred to as coarse-grained entities.
[0118] In one implementation, the text can be segmented into words, and the resulting words can be vectorized to obtain segmentation vectors. Then, the similarity between the segmentation vectors and entities in a preset entity set can be used as the basis for the calculation.
[0119] Specifically, text units to which the word segmentation vectors corresponding to a similarity greater than a preset threshold belong can be identified as text units suspected of containing entities. The word segmentation and vectorization methods described above have been previously discussed. Figure 1 The embodiments shown are illustrated and will not be repeated here.
[0120] In another implementation, the text can be segmented to obtain candidate text units, and then non-entity text units that do not contain entities can be identified among the candidate text units. The text units that are not non-entity text units among the candidate text units are identified as text units that are suspected to contain entities.
[0121] When dividing text, it can be done according to punctuation marks or based on the semantic analysis algorithm mentioned above, which will not be described in detail here.
[0122] In one scenario, this implementation can be based on a model using the ERNIE+crf (conditional random field) architecture.
[0123] The above model can be trained by pre-labeling sample text units containing entities such as people, places, organizations, activities, data, and holidays and anniversaries, so that the model learns the features of text units containing entities, thereby training a model that can output text units that are suspected to contain entities.
[0124] As can be seen, by identifying text units other than entity text units in the candidate text units as text units that are suspected to contain entities, text units that do not contain entities are filtered out, thereby accurately identifying text units that are suspected to contain entities.
[0125] Step S303: Perform word segmentation on the obtained text units, and identify the target entity and its entity attributes based on the obtained word segmentation.
[0126] The following describes how to identify target entities and their attributes.
[0127] In one implementation, each obtained text unit can be segmented into words, and the target entity and its entity attributes can be identified based on the segmented words.
[0128] Specifically, based on the part-of-speech and contextual semantics of the obtained word segmentation, the boundaries of target entities and their attributes within the text unit can be identified. Therefore, based on these identified boundaries, the target entities and their attributes contained within the text unit can be determined. The contextual semantics of the word segmentation within the text unit can be calculated using Bayesian posterior probabilities.
[0129] The following is a simple example illustrating the identification of target entities and entity attributes.
[0130] For example, for the text unit “Xiaoming, who lives in community q1, is the CEO of company C1”, the word segments can be “Xiaoming”, “company C1”, and “CEO”. Through part-of-speech analysis, we can determine that the word segment “Xiaoming” represents a person, “CEO” represents a position, and “community q1” represents an address. Then, through the contextual semantics of the word segments in the text unit, we can determine that the word segment “Xiaoming” represents a person entity, “CEO” represents a position entity, and “community q1” represents an attribute of the person entity “Xiaoming”.
[0131] This allows us to determine the boundaries of target entities and their attributes within a text unit based on word segmentation, part-of-speech, and contextual semantics of the segmented words. Consequently, we can accurately and quickly identify the target entities and their attributes within the text unit based on the identified boundaries of the target entities and their attributes.
[0132] In another implementation, the obtained text units can be merged to obtain merged text units, and the word segmentation recognition target entity and the entity attributes of the target entity contained in each merged text unit can be identified.
[0133] The obtained text units can be merged in the following ways:
[0134] The first method involves segmenting the obtained text units and merging them based on the semantic information of the segmented words.
[0135] The second method involves merging adjacent text units containing fewer than a preset number of characters to obtain a new text unit.
[0136] The method of identifying the target entity and its entity attributes from merged text units can be obtained based on the previous implementation method, the only difference being the different text units, which will not be elaborated here.
[0137] In another implementation, the target entity and its attributes can be identified based on a model with a GRU+crf architecture.
[0138] In one scenario, the aforementioned model and an encyclopedia dictionary can be combined to identify the target entity and its attributes. The encyclopedia dictionary is used to normalize and disambiguate the identified target entity and its attributes.
[0139] Step S304: Based on the event category, determine the candidate relationship patterns between entities.
[0140] Step S305: Determine the entity relationships between target entities based on candidate relationship patterns, target entities, entity attributes, and the target text units to which the target entities belong.
[0141] Step S306: Construct a knowledge graph based on the target entity and entity relationships.
[0142] Steps S304-S306 above are the same as those mentioned above. Figure 1 In the illustrated embodiment, steps S103-S105 are the same and will not be repeated here.
[0143] As can be seen from the above, when performing entity recognition on text using the scheme provided in this embodiment, the text is first divided to obtain text units that are suspected of containing entities. Then, the target entity and its attributes are obtained based on the obtained text units. It is evident that by first identifying text units that are suspected of containing entities through coarse-grained filtering, and then obtaining the target entity and its attributes from the filtered text units, the scope of text processed when obtaining the target entity and its attributes can be effectively narrowed. This helps reduce the computational load consumed in processing, thereby improving the efficiency of obtaining the target entity and its attributes.
[0144] The following is combined Figure 4 This provides a more intuitive explanation of the target entity and entity attribute determination process shown in steps S302-S303 above.
[0145] See Figure 4 This is a schematic diagram of an entity and attribute determination process provided in an embodiment of this disclosure.
[0146] Depend on Figure 4 As can be seen, the process in the diagram can be roughly divided into the following steps 1-3:
[0147] Step 1: First, divide the text to obtain text units that are suspected to contain entities.
[0148] For example, the text in the image can be divided into text unit 1, which is suspected to contain entities, “Person P1 lives in community q1 and works for company C1”, and text unit 2, “Today I visited location L1 in district C of county B in city A”. The text unit “had a lot of fun” is filtered out because it does not contain entities.
[0149] Step 2: Perform word segmentation on each obtained text unit to obtain individual words.
[0150] For example, in the diagram, word segmentation 1-word segmentation 7: personnel P1, community q1, company C1, city A, county B, district C, and location L1 are all word segments obtained after segmenting the text units.
[0151] Step 3: Based on the part-of-speech and contextual semantics of the obtained word segmentation, the word segmentation can be divided into target entity and entity attributes.
[0152] For example, in the diagram, segment 1 is identified as entity 1, segment 3 as entity 2, segment 7 as entity 3, while segment 2 is identified as an attribute of entity 1, and segments 4-6 are identified as attributes of entity 3.
[0153] exist Figure 1 Based on the illustrated embodiment, when determining entity relationships, the relation subject and relation object in the target entity under the candidate relation pattern can be determined based on the target characters in the target text unit to which the target entity belongs, excluding the target entity and entity attributes. Then, based on the relation subject, relation object, and candidate relation pattern, entity relationships between target entities are generated. In view of the above, this disclosure provides a third method for constructing knowledge graphs.
[0154] See Figure 5 This is a flowchart illustrating the third knowledge graph construction method provided in this embodiment of the present disclosure. The method includes the following steps S501-S507.
[0155] Step S501: Perform semantic analysis on the text to obtain the event categories of the events associated with the people described in the text.
[0156] Step S502: Perform entity recognition on the text to obtain the target entity for the person and the entity attributes of the target entity.
[0157] Step S503: Based on the event category, determine the candidate relationship patterns between entities.
[0158] Steps S501-S503 above are the same as those described above. Figure 1 In the illustrated embodiment, steps S101-S103 are the same and will not be repeated here.
[0159] Step S504: Obtain the target characters in the target text unit to which the target entity belongs, excluding the target entity and entity attributes.
[0160] Since the target entity and entity attributes have been determined, the target characters other than the target entity and entity attributes can be obtained from the target text unit.
[0161] Step S505: Based on the obtained target characters, determine the relation subject and relation object in the target entity under the candidate relation pattern.
[0162] Since the target characters are characters other than the target entity and entity attributes, they often represent relationships such as attribution and indication between entities. Therefore, the relation subjects and relation objects in the target entities in the candidate relation pattern can be determined based on the obtained target characters.
[0163] Specifically, the relation subject and relation object in the target entity can be determined based on the auxiliary words contained in the target characters.
[0164] For example, for the text unit "Personnel P1 is Company C1's...", if the entities "Personnel P1" and "Company C1" have been identified, then based on the fact that the target characters contain the two particles "is" and "of", the entity after the particle "is" can be identified as the relation object, and the entity before the particle "is" can be identified as the relation subject; or, the entity between the particles "is" and "of" can be identified as the relation object, and the other entity can be identified as the relation subject, etc.
[0165] Step S506: Generate entity relationships between target entities based on the relation subject, relation object, and candidate relation schema.
[0166] Step S507: Construct a knowledge graph based on the target entity and entity relationships.
[0167] As can be seen from the above, when constructing a knowledge graph using the solution provided in the embodiments of this disclosure, the entity relationship between target entities can be determined based on the target characters in the target text unit to which the target entity belongs, excluding the target entity and entity attributes. This allows us to focus on the target characters rather than all the text, reducing the amount of data processing when determining entity relationships and improving the efficiency of entity relationship determination.
[0168] Furthermore, based on the target characters, the relation subjects and relation objects in the target entities under the candidate relation patterns can be determined. Thus, entity relationships between target entities can be generated conveniently and quickly based on the relation subjects, relation objects, and candidate relation patterns, further improving the efficiency of entity relationship determination.
[0169] In one embodiment of this disclosure, steps S504-S505 can be implemented based on a pre-trained relation extraction model.
[0170] The aforementioned relation extraction model is a sequence labeling model. The event category, target entity, entity attributes, text unit containing the entity, and word segmentation of the text are input into the relation extraction model. The model determines and outputs the entity relationships between target entities based on the LAC (Lexical Analysis of Chinese) algorithm.
[0171] In one scenario, the entity relationships between target entities can be obtained by combining the above model with dictionary-enhanced relational data.
[0172] The aforementioned dictionary-enhanced relation data is an entity relation table obtained through extensive text annotation and training. It contains all preset candidate relation patterns, the subject and object of each candidate relation pattern, the original string of the relation, etc. Based on the aforementioned dictionary-enhanced relation data, it is possible to quickly analyze whether there are relations between target entities in the text, and the types of relations that exist.
[0173] Corresponding to the knowledge graph construction method described above, this disclosure also provides an information retrieval method.
[0174] See Figure 6 This is a flowchart illustrating the first information retrieval method provided in this embodiment of the present disclosure. The method includes the following steps S601-S603.
[0175] Step S601: Perform word segmentation on the retrieved text.
[0176] The word segmentation process described above has been discussed in the preceding text. Figure 1 The embodiments shown are illustrated and will not be repeated here.
[0177] Step S602: From the obtained word segmentation, determine the search subjects included in the search text and the search intent targeting the search subjects.
[0178] The method for determining the searched individuals and their search intent based on word segmentation can be obtained using the aforementioned semantic analysis algorithms, which will not be described in detail here.
[0179] Step S603: Based on the search target and the search intent, a graph walk is used to search in the knowledge graph to obtain search results for the search target.
[0180] The aforementioned knowledge graph is a knowledge graph constructed according to any embodiment of the aforementioned knowledge graph construction method.
[0181] Specifically, the following methods can be used to obtain search results for the searched person.
[0182] In one implementation, graph walking algorithms such as DeepWalk and RandomWalk can be used to search the knowledge graph based on the searched person and the search intent, thereby obtaining search results for the searched person.
[0183] In another implementation, a search starting node with the searched person as the entity can be determined from the knowledge graph. Then, based on the entity relationships represented by the determined outgoing edges, a target outgoing edge corresponding to the search intent can be selected from the outgoing edges of the search starting node. Thus, search results for the searched person can be obtained based on the entity represented by the target node pointed to by the target outgoing edge. Detailed implementation methods will be provided later. Figure 7 Steps S703-S706 in the illustrated embodiment will not be repeated here.
[0184] As can be seen from the above, when using the solution provided in this embodiment for information retrieval, the searcher and the search intent for the searcher can be determined from the word segments obtained by word segmentation of the search text. Thus, based on the searcher and the search intent, a targeted search can be performed in the knowledge graph using a graph walking approach, thereby quickly obtaining search results for the searcher and improving the efficiency of information retrieval.
[0185] Furthermore, when existing technologies retrieve information based on a database of personal information, they often require further manual filtering of the search results before the final results can be obtained. Therefore, the solution provided in this disclosure saves human resources in the information retrieval process compared to existing technologies, and further improves the efficiency of information retrieval.
[0186] exist Figure 6 Based on the illustrated embodiment, when retrieving information from a knowledge graph, a starting node for retrieval with the retrieval target person as the entity can be determined from the knowledge graph. Then, based on the entity relationships represented by the determined outgoing edges, a target outgoing edge corresponding to the retrieval intent can be selected from the outgoing edges of the starting node. Thus, retrieval results for the retrieval target person can be obtained based on the entity represented by the target node pointed to by the target outgoing edge. In view of the above, this disclosure provides a second information retrieval method.
[0187] See Figure 7 This is a flowchart illustrating the second information retrieval method provided in this embodiment of the present disclosure. The method includes the following steps S701-S706.
[0188] Step S701: Perform word segmentation on the searched text.
[0189] Step S702: From the obtained word segmentation, determine the search subjects included in the search text and the search intent targeting the search subjects.
[0190] Steps S701-S702 above are the same as those described above. Figure 6 In the illustrated embodiment, steps S601-S602 are the same and will not be repeated here.
[0191] Step S703: Determine the starting node for retrieval with the retrieval person as the entity from the knowledge graph.
[0192] Specifically, the node in the knowledge graph that corresponds to the person being searched can be determined as the starting node for the search.
[0193] Step S704: Determine the outgoing edges of the starting node for retrieval.
[0194] In one implementation, all outgoing edges from the starting node of the retrieval can be determined.
[0195] In another implementation, outgoing edges with a relationship strength greater than a preset strength can be determined from among the outgoing edges in the starting node of the retrieval. Here, the relationship strength of the outgoing edge can be understood as the confidence level of the relationship represented by the outgoing edge.
[0196] Step S705: Based on the entity relationships represented by the determined outgoing edges, select the target outgoing edge corresponding to the retrieval intent from the determined outgoing edges.
[0197] Specifically, the represented entity relationships can be compared with the above-mentioned search intent. Figure 1 The outgoing edge that is determined is the target outgoing edge.
[0198] Step S706: Based on the entity represented by the target node pointed to by the target outgoing edge, obtain the search results for the searched person.
[0199] Specifically, the entity represented by the target node pointed to by the outgoing edge of the target can be determined as the search result for the person being searched.
[0200] The following specific examples, shown in steps D-G, illustrate steps S702-S706.
[0201] Step D: From the obtained word segmentation, determine that the person to be searched in the search text is "person P3", and the search intent for the person to be searched is "position".
[0202] Step E: Determine the starting node for retrieval of the corresponding entity "Personnel P3" from the knowledge graph.
[0203] Step F: From the outgoing edges of the starting node, determine the target outgoing edge representing the entity relationship "job title".
[0204] Step G: Obtain the target entity "Company C2" represented by the target node pointed to by the above target outgoing edge, and determine the target entity "Company C2" as the search result for the searched person.
[0205] As can be seen from the above, firstly, the starting node for retrieval with the searched person as the entity is determined from the knowledge graph. Then, based on the entity relationship represented by the outgoing edges of the determined starting node, the target outgoing edge corresponding to the search intent can be selected from the determined outgoing edges. Finally, based on the entity represented by the target node pointed to by the target outgoing edge, the search results for the searched person can be accurately obtained.
[0206] Corresponding to the above-described knowledge graph construction method, this disclosure also provides a knowledge graph construction apparatus.
[0207] See Figure 8 This is a schematic diagram of a knowledge graph construction device provided in an embodiment of the present disclosure. The device includes the following modules 801-805.
[0208] The event category acquisition module 801 is used to perform semantic analysis on the text to obtain the event category of the events associated with the person described in the text;
[0209] The entity recognition module 802 is used to perform entity recognition on the text to obtain the target entity for the person and the entity attributes of the target entity;
[0210] The candidate relationship pattern determination module 803 is used to determine candidate relationship patterns between entities based on the event category;
[0211] The entity relationship determination module 804 is used to determine the entity relationship between target entities based on the candidate relationship pattern, target entity, entity attribute and target text unit to which the target entity belongs;
[0212] The knowledge graph construction module 805 is used to construct a knowledge graph based on the target entity and entity relationships.
[0213] As can be seen from the above, when constructing a knowledge graph using the solution provided in the embodiments of this disclosure, the event categories of the events associated with the person described in the text are first obtained, entity recognition is performed on the text to obtain the target entity for the person and the entity attributes of the target entity, and then candidate relationship patterns between entities are determined based on the event categories. In this way, the entity relationships between target entities can be determined based on the candidate relationship patterns, target entities, entity attributes, and the target text units to which the target entities belong. Finally, a knowledge graph can be constructed based on the target entities and entity relationships, thereby enabling fast and accurate information retrieval based on the knowledge graph.
[0214] Furthermore, it can be seen that the solution provided in this embodiment of the present disclosure, when constructing a knowledge graph, first obtains the event categories of events associated with people described in the text, and then determines candidate relationship patterns between entities based on the event categories, thereby determining the entity relationships between target entities based on the candidate relationship patterns. That is, candidate relationship patterns are determined before determining entity relationships, and entity relationships are determined based on the previously obtained candidate relationship patterns. This allows for the purposeful determination of entity relationships using candidate relationship patterns as a benchmark, reducing the randomness in determining entity relationships compared to directly determining entity relationships from the entire text, and improving the efficiency of entity relationship determination.
[0215] In one embodiment of this disclosure, the entity relationship determination module 804 is specifically used to obtain target characters in the target text unit to which the target entity belongs, excluding the target entity and entity attributes; based on the obtained target characters, determine the relationship subject and relationship object in the target entity under the candidate relationship pattern; and generate entity relationships between target entities based on the relationship subject, relationship object and the candidate relationship pattern.
[0216] As can be seen from the above, when constructing a knowledge graph using the solution provided in the embodiments of this disclosure, the entity relationship between target entities can be determined based on the target characters in the target text unit to which the target entity belongs, excluding the target entity and entity attributes. This allows us to focus on the target characters rather than all the text, reducing the amount of data processing when determining entity relationships and improving the efficiency of entity relationship determination.
[0217] Furthermore, based on the target characters, the relation subjects and relation objects in the target entities under the candidate relation patterns can be determined. Thus, entity relationships between target entities can be generated conveniently and quickly based on the relation subjects, relation objects, and candidate relation patterns, further improving the efficiency of entity relationship determination.
[0218] In one embodiment of this disclosure, the entity recognition module 802 includes:
[0219] The text segmentation submodule is used to segment the text to obtain text units that are suspected to include entities;
[0220] The word segmentation processing submodule is used to segment the obtained text units and identify target entities and their attributes based on the segmented words.
[0221] As can be seen from the above, when performing entity recognition on text using the scheme provided in this embodiment, the text is first divided to obtain text units that are suspected of containing entities. Then, the target entity and its attributes are obtained based on the obtained text units. It is evident that by first identifying text units that are suspected of containing entities through coarse-grained filtering, and then obtaining the target entity and its attributes from the filtered text units, the scope of text processed when obtaining the target entity and its attributes can be effectively narrowed. This helps reduce the computational load consumed in processing, thereby improving the efficiency of obtaining the target entity and its attributes.
[0222] In one embodiment of this disclosure, the text segmentation submodule is specifically used to segment the text to obtain candidate text units included in the text; determine non-entity text units that do not contain entities among the candidate text units; and determine the text units other than non-entity text units among the candidate text units as text units that are suspected to contain entities.
[0223] As can be seen, by identifying text units other than entity text units in the candidate text units as text units that are suspected to contain entities, text units that do not contain entities are filtered out, thereby accurately identifying text units that are suspected to contain entities.
[0224] In one embodiment of this disclosure, the word segmentation processing submodule is specifically used to perform word segmentation processing on the obtained text unit; based on the part-of-speech of the obtained word segmentation and the contextual semantics of the obtained word segmentation in the obtained text unit, identify the boundary of the target entity and the boundary of the entity attribute of the target entity in the obtained text unit; based on the identified boundary of the target entity and the boundary of the entity attribute, determine the target entity and the entity attribute of the target entity contained in the obtained text unit.
[0225] This allows us to determine the boundaries of target entities and their attributes within a text unit based on word segmentation, part-of-speech, and contextual semantics of the segmented words. Consequently, we can accurately and quickly identify the target entities and their attributes within the text unit based on the identified boundaries of the target entities and their attributes.
[0226] In one embodiment of this disclosure, the event category acquisition module 801 is specifically used to perform word segmentation on the text and vectorize the obtained word segments to obtain word segmentation vectors; based on the obtained word segmentation vectors, predict the probability that the person-related event described in the text belongs to a preset category of person-related events; based on the predicted probability, obtain the event category of the person-related event described in the text from the preset categories of task-related events.
[0227] As can be seen, after segmenting the text and vectorizing the segmented words to obtain segmented vectors, the probability of the events related to the people described in the text belonging to the category of preset events can be predicted based on the obtained segmented vectors. Thus, the event category of the events related to the people described in the text can be quickly and accurately determined based on the predicted probability.
[0228] Corresponding to the above information retrieval method, this disclosure also provides an information retrieval device.
[0229] See Figure 9 This is a schematic diagram of a knowledge graph construction device provided in an embodiment of the present disclosure. The device includes the following modules 901-903.
[0230] The word segmentation processing module 901 is used to segment the retrieved text into words.
[0231] The information determination module 902 is used to determine, from the obtained word segmentation, the search person included in the search text and the search intent for the search person;
[0232] The retrieval module 903 is used to perform a retrieval in the knowledge graph based on the retrieval person and the retrieval intent using a graph walking method to obtain retrieval results for the retrieval person, wherein the knowledge graph is a knowledge graph constructed according to the aforementioned knowledge graph construction method.
[0233] As can be seen from the above, when using the solution provided in this embodiment for information retrieval, the searcher and the search intent for the searcher can be determined from the word segments obtained by word segmentation of the search text. Thus, based on the searcher and the search intent, a targeted search can be performed in the knowledge graph using a graph walking approach, thereby quickly obtaining search results for the searcher and improving the efficiency of information retrieval.
[0234] Furthermore, when existing technologies retrieve information based on a database of personal information, they often require further manual filtering of the search results before the final results can be obtained. Therefore, the solution provided in this disclosure saves human resources in the information retrieval process compared to existing technologies, and further improves the efficiency of information retrieval.
[0235] In one embodiment of this disclosure, the retrieval module 903 is specifically configured to: determine a retrieval starting node with the retrieval person as the entity from the knowledge graph; determine the outgoing edges of the retrieval starting node; select the target outgoing edge corresponding to the retrieval intent from the determined outgoing edges based on the entity relationships represented by the determined outgoing edges; and obtain retrieval results for the retrieval person based on the entity represented by the target node pointed to by the target outgoing edge.
[0236] As can be seen from the above, firstly, the starting node for retrieval with the searched person as the entity is determined from the knowledge graph. Then, based on the entity relationship represented by the outgoing edges of the determined starting node, the target outgoing edge corresponding to the search intent can be selected from the determined outgoing edges. Finally, based on the entity represented by the target node pointed to by the target outgoing edge, the search results for the searched person can be accurately obtained.
[0237] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0238] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0239] In one embodiment of this disclosure, an electronic device is provided, comprising:
[0240] At least one processor; and
[0241] A memory communicatively connected to the at least one processor; wherein,
[0242] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the aforementioned knowledge graph construction or information retrieval methods.
[0243] In one embodiment of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the aforementioned knowledge graph construction or information retrieval method.
[0244] In one embodiment of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the aforementioned knowledge graph construction or information retrieval method.
[0245] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0246] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0247] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0248] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as knowledge graph construction or information retrieval methods. For example, in some embodiments, the knowledge graph construction or information retrieval method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the knowledge graph construction or information retrieval method described above may be performed. Alternatively, in other embodiments, computing unit 1001 may be configured to perform knowledge graph construction or information retrieval methods by any other suitable means (e.g., by means of firmware).
[0249] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0250] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0251] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0252] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0253] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0254] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0255] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0256] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A knowledge graph construction method, comprising: The text is segmented into words, and the segmented words are vectorized to obtain segmentation vectors. Based on the similarity between the word segmentation vector and the category vector corresponding to the preset category of person-related events, the probability that the person-related events described in the text belong to the preset category of person-related events is predicted; Based on the predicted probabilities, the event categories of the person-related events described in the text are obtained from the preset categories of person-related events; Entity recognition is performed on the text to obtain the target entity for the person and the entity attributes of the target entity; Based on the event category, candidate relationship patterns for relationships between entities are determined. The candidate relationship patterns are relationship formats that represent the relationship subject, the relationship object, and the candidate relationship. Obtain the target characters in the target text unit to which the target entity belongs, excluding the target entity and its attributes; Based on the obtained target characters, determine the relation subject and relation object in the target entity under the candidate relation pattern; Based on the relationship subject, relationship object, and candidate relationship pattern, generate entity relationships between target entities; Construct a knowledge graph based on the target entity and entity relationships.
2. The method according to claim 1, wherein, The entity recognition process for the text, which yields the target entity for the person and its attributes, includes: The text is segmented to obtain text units that are suspected to contain entities; The obtained text units are segmented into words, and the target entities and their attributes are identified based on the segmented words.
3. The method according to claim 2, wherein, The text segmentation, yielding text units that are suspected to contain entities, includes: The text is divided to obtain candidate text units. Identify non-entity text units among the candidate text units that do not contain entities; Text units other than entity text units among the candidate text units are identified as text units that are suspected to include entities.
4. The method according to claim 2, wherein, The method for identifying target entities and their attributes based on the obtained word segmentation includes: Based on the part-of-speech of the obtained word segmentation and the contextual semantics of the obtained word segmentation in the obtained text unit, the boundaries of the target entity and the boundaries of the entity attributes of the target entity in the obtained text unit are identified. Based on the boundaries of the identified target entities and the boundaries of their attributes, the target entities and their attributes contained in the obtained text units are determined.
5. An information retrieval method, comprising: Perform word segmentation on the searched text; From the obtained word segments, determine the searched individuals included in the search text and the search intent targeting those individuals; Based on the search target and search intent, a graph walk is used to search in the knowledge graph to obtain search results for the search target. The knowledge graph is a knowledge graph constructed according to any one of claims 1-4.
6. The method according to claim 5, wherein, Based on the search target and search intent, a graph walk is used to perform a search in the knowledge graph to obtain search results for the search target, including: Determine the starting node for retrieval from the knowledge graph, with the retrieved person as the entity; Determine the outgoing edges of the starting node of the retrieval; Based on the entity relationships represented by the determined outgoing edges, select the target outgoing edge corresponding to the search intent from the determined outgoing edges; Based on the entity represented by the target node pointed to by the outgoing edge of the target, the search results for the searched person are obtained.
7. A knowledge graph construction apparatus, comprising: The event category acquisition module is used to perform word segmentation on the text and vectorize the segmented words to obtain word vectors. Based on the similarity between the word segmentation vector and the category vector corresponding to the preset category of person-related events, the probability that the person-related event described in the text belongs to the preset category of person-related events is predicted; based on the predicted probability, the event category of the person-related event described in the text is obtained from the preset categories of person-related events. An entity recognition module is used to perform entity recognition on the text to obtain the target entity for the person and the entity attributes of the target entity; The candidate relationship pattern determination module is used to determine candidate relationship patterns between entities based on the event category, wherein the candidate relationship pattern is a relationship format that represents the relationship subject, the relationship object, and the candidate relationship; The entity relationship determination module is used to obtain the target characters in the target text unit to which the target entity belongs, excluding the target entity and entity attributes. Based on the obtained target characters, determine the relation subject and relation object in the target entity under the candidate relation pattern; Based on the relationship subject, relationship object, and candidate relationship pattern, generate entity relationships between target entities; The knowledge graph construction module is used to build a knowledge graph based on the target entity and entity relationships.
8. The apparatus according to claim 7, wherein, The entity recognition module includes: The text segmentation submodule is used to segment the text to obtain text units that are suspected to include entities; The word segmentation processing submodule is used to segment the obtained text units and identify target entities and their attributes based on the segmented words.
9. The apparatus according to claim 8, wherein, The text segmentation submodule is specifically used to segment the text to obtain candidate text units included in the text; determine non-entity text units that do not contain entities among the candidate text units; and determine the text units other than non-entity text units among the candidate text units as text units that are suspected to contain entities.
10. The apparatus according to claim 8, wherein, The word segmentation processing submodule is specifically used to perform word segmentation processing on the obtained text units; based on the part-of-speech of the obtained words and the contextual semantics of the obtained words in the obtained text units, identify the boundaries of target entities and the boundaries of entity attributes of target entities in the obtained text units; based on the identified boundaries of target entities and the boundaries of entity attributes, determine the target entities and entity attributes contained in the obtained text units.
11. An information retrieval device, comprising: The word segmentation module is used to segment the retrieved text into words. The information determination module is used to determine, from the obtained word segmentation, the search person included in the search text and the search intent for the search person; The retrieval module is used to perform a retrieval in a knowledge graph using a graph walking method based on the retrieval person and the retrieval intent, and to obtain retrieval results for the retrieval person, wherein the knowledge graph is a knowledge graph constructed according to any one of claims 1-4.
12. The apparatus according to claim 11, wherein, The retrieval module is specifically used to determine the retrieval starting node with the retrieval person as the entity from the knowledge graph; determine the outgoing edges of the retrieval starting node; select the target outgoing edge corresponding to the retrieval intent from the determined outgoing edges based on the entity relationship represented by the determined outgoing edges; and obtain the retrieval results for the retrieval person based on the entity represented by the target node pointed to by the target outgoing edge.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method and device for creating knowledge graph
CN107665252A
Medical field intention recognition method, device and equipment and storage medium
CN112035635A