Knowledge graph construction method based on sparse scaling network and multiple hypergraphs

By using a deep learning model of sparse scaling networks and multiple hypergraphs, the problems of entity recognition ignoring redundant information and low accuracy of nested entity recognition in existing technologies are solved, achieving more efficient and accurate knowledge graph construction.

CN120671787APending Publication Date: 2025-09-19BEIJING INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510582035.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-08
Filing Date
2025-05-07
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing knowledge graph construction methods ignore redundant information in text during entity recognition, resulting in decreased accuracy. In addition, the accuracy of nested entity recognition is not high, and modeling is difficult and computationally expensive.

Method used

A deep learning model based on sparse scaling networks and multiple hypergraphs is used to identify nested named entities through entity boundary recognition, multiple local hypergraph generation and decoding, and then perform relationship extraction on this basis to generate text triples and construct a knowledge graph.

Benefits of technology

It improves the accuracy and efficiency of knowledge graph construction, can better identify nested named entities and extract entity relationships, and reduces the impact of redundant information on recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671787A_ABST
    Figure CN120671787A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method based on a sparse scaling network and multiple hypergraphs, and belongs to the technical field of information extraction. Comprising the following steps: firstly, preprocessing texts in a corpus; secondly, performing boundary identification on entities possibly existing in the text, and generating boundary candidate marks of the entities; then, each word or character is labeled, and multiple local hypergraph representations are generated from the front direction, the back direction, the left direction and the right direction; then, decoding the multiple local hypergraphs, and identifying a nested named entity; finally, a multi-layer perceptron-based model is used to learn mapping from entity grammar features to entity pair relationship types. Through redundant information processing based on the sparse scaling network, redundant information can be reduced, and key semantic features in the text can be captured more accurately; through a nested named entity recognition method based on multiple hypergraphs, the performance of knowledge graph construction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a knowledge graph construction method based on a sparse scaling network and multiple hypergraphs, and belongs to the technical field of information retrieval and information extraction. Background Art

[0002] Knowledge graph construction is a key research topic in natural language processing. Knowledge graphs are typically composed of triples, consisting of a head entity, a relationship, and a tail entity. Knowledge graph construction involves constructing a knowledge graph from natural language text. Knowledge graphs enable structured organization and representation of textual information, providing a foundation and technical support for knowledge application systems. They also enrich semantic knowledge for downstream tasks, improving their performance.

[0003] Knowledge graph construction primarily involves text-based entity recognition and relationship extraction. Named entity recognition (NER) involves identifying entities with specific meanings from text. Early NER methods mostly employed expert-constructed rule templates, primarily matching patterns and strings. Later, entity recognition methods based on statistical machine learning emerged. These models treat NER as a sequence labeling problem. Currently, entity recognition models are primarily categorized into three approaches: sequence labeling-based, hypergraph-based, and span-based.

[0004] Sequence tagging-based methods adopt a hierarchical approach and stack flat entity layers according to the hierarchical nature of the structure in nested named entities. Jue Wang et al. proposed a pyramid model in the paper "Pyramid: A Layered Model for NestedNamedEntity Recognition" (Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020). The model obtains entity references through a forward pyramid, stacking each word in the sentence from bottom to top to generate candidate entities of different lengths; then, nested named entities are identified through hierarchical decoding using an inverse pyramid structure. Hypergraph-based methods construct sentences into a hypergraph based on the nested named entity structure to identify nested entities. In the paper "Nested Named Entity Recognition Revisited" (Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics, 2018), Arzoo Katiyar et al. used a recurrent neural network to encode directed hypergraph representations and designed a model based on a long short-term memory (LSTM) network to learn hypergraph representations of nested entities in input sentences. Span-based methods identify nested entities by generating all possible spans and classifying them, and confirming whether they are valid entities and types. For example, Makoto Miwa et al., in the paper "Deep Exhaustive Model for Nested Named Entity Recognition" (Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018), enumerated all regions that do not exceed a specific length and learned their representations through a bidirectional long short-term memory network Bi-LSTM (Bidirectional Long Short-Term Memory), and then classified the regions to identify nested entities.

[0005] The task of relation extraction involves identifying semantic relationships between entities from natural language text, thereby constructing a knowledge graph. The goal of relation extraction is to identify the relationship between two entities from text. Currently, relation extraction models are mainly divided into two categories: pipeline methods and joint extraction methods.

[0006] The pipeline-based relation extraction method refers to the process of extracting the relationship between entities and generating triples after completing entity recognition. Yatian Shen et al. proposed a convolutional neural network (CNN) model based on the attention mechanism in the document "Attention-Based Convolutional Neural Network for Semantic Relation Extraction" (26th International Conference on Computational Linguistics, Proceedings of the Conference, 2016). The model uses word embedding, part-of-speech tagging embedding and position embedding information to mine important features hidden in the sentence. Relation extraction based on joint extraction refers to the use of a unified modeling approach to extract entity relationships. This method can directly extract structured triples from unstructured text. Xiangrong Zeng et al. proposed a relation extraction model based on a copy mechanism in the paper "Extracting Relational Facts by an End-to-End Neural Model with Copy Mechanism" (Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018). The model uses a sequence generation framework to extract relations, head entities, and tail entities in sequence, and copies the extracted entities, continuously decoding them to obtain new triple information.

[0007] Existing knowledge graph construction methods have the following main problems: First, when identifying entities from text, they pay little attention to and utilize redundant information in the text, which affects the accuracy of the constructed knowledge graph. Second, for nested entity recognition, existing methods have relatively low accuracy, are difficult to implement, and have high computational costs. Summary of the Invention

[0008] The present invention aims to address the relatively low accuracy of nested entity recognition and the difficulty of modeling in existing knowledge graph construction. It provides a method for constructing a knowledge graph based on a sparse scaling network and multiple hypergraphs. This method uses a deep learning model based on a sparse scaling network and multiple hypergraphs to identify entities based on unstructured text input by the user. It then extracts relationships from the identified entities, generates text triples, and finally uses these triples to construct a knowledge graph.

[0009] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions.

[0010] The knowledge graph construction method based on sparse scaling network and multiple hypergraphs is designed based on the hypergraph method. The method includes four modules:

[0011] (1) The first module is the entity boundary recognition module, which uses the sequence labeling framework to decompose the input text into word or character sequences and generate candidate boundary markers for entities.

[0012] (2) The second module is the multiple hypergraph module, which uses the BIOE annotation scheme to annotate words or characters. This module is trained using a sparse scaling network, which maintains a sparse hidden state through soft thresholding and expands the input token sequence to generate a local hypergraph. The multiple hypergraph module generates local hypergraphs in four directions (forward, backward, left, and right), forming a set of multiple local hypergraphs.

[0013] (3) The third module is the entity prediction module, which decodes named entities from the local hypergraph sets in four directions and records the predicted probabilities of candidate entities. After all entities are decoded, the candidate entities in different directions are merged using the predicted probabilities to obtain the nested entity recognition results.

[0014] (4) The fourth module is the relationship extraction module. After entity recognition is completed, the semantic relationship between entities is extracted, triples are generated, and a knowledge graph is constructed.

[0015] The method for constructing a knowledge graph based on a sparse scaling network and multiple hypergraphs includes the following steps:

[0016] Step 1: Preprocess the text in the dataset, including word segmentation, entity boundary identification, and part-of-speech tagging.

[0017] The specific steps include:

[0018] Step 1.1: Segment and tag the text sentences in the dataset;

[0019] Step 1.1.1: Clean the dataset, extract sentences and perform word segmentation;

[0020] Step 1.1.2: Tag the parts of speech for each word in the sentence;

[0021] Use StanfordCoreNLP tool to perform part-of-speech tagging on word segmentation;

[0022] Step 1.2: Extract the boundary information of the identified entities in the dataset;

[0023] Step 1.3: Combine the results of steps 1.1 and 1.2 to generate structured data;

[0024] Step 2: Identify the boundaries of entities that may exist in the text and generate candidate boundary markers for the entities;

[0025] The specific steps include:

[0026] Step 2.1: Connect the pre-trained BERT (Bidirectional Encoder Representation from Transformers) encoder Word Embedding Part-of-speech embeddings Character Embedding Four components to construct the word segmentation feature vector z i , as shown in formula (1);

[0027]

[0028] Step 2.2: Transform the eigenvector z i Input to the token-level bidirectional long short-term memory network Bi-LSTM (Bidirectional Long Short-Term Memory) layer, and treat its hidden state as the token representation t i , as shown in formulas (2)-(4);

[0029]

[0030] Step 2.3: Use the feedforward layer to represent t with tokens i As input, calculate the probability that the i-th token constitutes the left or right boundary of the entity As shown in formula (5);

[0031]

[0032] Step 2.4: Use focal loss L l / r It is trained as the minimization objective function of the entity boundary recognition module, as shown in formula (6);

[0033]

[0034] in is the basic label of the i-th token, N is the number of samples, and r is a parameter.

[0035] Step 3: Use the BIOE annotation scheme to annotate each word or character and generate multiple local hypergraph representations in the four directions of forward / backward / left / right;

[0036] The specific steps include:

[0037] Step 3.1: Use left margin information To initialize the unit memory C i,0 and hidden state h i,0 , enhance boundary information, as shown in formula (7);

[0038]

[0039] Step 3.2: Take the token(token) subsequence as input and generate a label at each time step to construct a hypergraph. Use a sparse scaling network structure to capture the important semantic features in the input token sequence, use the feedforward layer to calculate the label prediction probability, obtain the label sequence, and use the hypergraph generation rules GraphRules to expand the forward local hypergraph, as shown in formulas (8)-(12);

[0040]

[0041] p i,k =softmax(FFN(h i,k ))(9)

[0042]

[0043] in, is the input token sequence, h i,k is the hidden state, p i,k is the label prediction probability of the kth token in the subsequence corresponding to the i-th left boundary, is its predicted label, is the label sequence, is The i-th forward local hypergraph generated by implementing the local hypergraph design rules on , where SSN represents sparse scaling network.

[0044] The sparse scaling network uses soft thresholding to perform differential connections on multi-layer convolutional LSTM and GRU (Gated Recurrent Unit). Taking the trained threshold λ as a reference, the elements less than the threshold are set to 0, the elements greater than or equal to the threshold are retained, and reasonable weights are retained according to the conditions of the elements. t , the hidden state h′ after soft thresholding t It is expressed as shown in formula (13);

[0045]

[0046] In the BIOE sequence tagging scheme, B represents the start node, I represents the middle node, E represents the end node, O represents other nodes, and the mark is irrelevant. x represents the predicted entity type x, Indicates the end node of the predicted entity type x.

[0047] The definition of the hypergraph generation rules GraphRules is as follows, specifically including:

[0048] Rule 1: Initialize the local hypergraph to a single B node;

[0049] Rule 2: In each time step, only the last B / I node is used as the head node;

[0050] Rule 3: If cls is predicted x Node and Nodes are connected to the head node via hyperedges;

[0051] Rule 4: When an I / O tag is predicted, the I / O node will be connected to the head node with a normal edge;

[0052] Rule 5: Once an O node is predicted, the local hypergraph stops growing;

[0053] Step 3.3: Similar to the method of generating the forward local hypergraph in step 3.2, the GRU model generates local hypergraphs from the left and right directions respectively, and the LSTM model generates local hypergraphs from the forward and backward directions respectively, thereby forming a set of multiple local hypergraphs;

[0054] Step 3.4: When training hypergraph generation, use cross entropy L g As the objective function to be minimized, as shown in formula (14);

[0055]

[0056] Among them, ε is the node type set, y i,kis the underlying truth type of the kth token in the subsequence. Starting from the i-th left boundary, is the predicted probability that the kth token in the corresponding token subsequence is labeled as type c, and M is the number of samples.

[0057] Step 4: Decode multiple local hypergraphs to identify nested named entities;

[0058] The specific steps include:

[0059] Step 4.1: Search the local hypergraph in each direction from node B to Node paths,decode nested named entities for multiple local hypergraphs;

[0060] Use multiple sets of quadruple E f / b / l / r =(idx l ,idx r ,cls,p f / b / l / r ) represents the predicted named entities in the four directions of forward, backward, left and right. l and idx r Indicates the index of the left and right boundaries, cls indicates the predicted entity type, and p indicates that the last node is E cls The predicted probability of .

[0061] Step 4.2: Combine the forward, backward, left, and right direction sets of each named entity to form the final nested entity prediction result, as shown in formula (15). Where θ is a hyperparameter used as a threshold. If p m ≥θ, the model retains the predicted candidate named entities;

[0062]

[0063] Step 5: Use a multi-layer perceptron based model to learn the mapping from entity grammar features to entity-to-relation types;

[0064] The specific steps include:

[0065] Step 5.1: Transform the eigenvector f(x i ) is input into a multilayer perceptron and passes through multiple hidden layers to extract the feature h, as shown in formula (16). Where W1 is the weight matrix, b1 is the bias vector, and σ is the ReLu activation function.

[0066] h=σ(W1f(x i )+b1) (16)

[0067] Step 5.2: Input the output h of the hidden layer into the Softmax layer to classify the possible relationship labels in the relationship type. Where W2 and b2 are the weight matrix and bias vector of the Softmax layer, as shown in formula (17);

[0068] o=softmax(W2h+b2) (17)

[0069] Step 5.3: Use the cross entropy loss function to measure the difference between the predicted result and the true label. The training part updates the model parameters by minimizing the cross entropy loss function, as shown in formula (18);

[0070]

[0071] Where C is the number of categories, y i Is the real label, o i is the output of the model.

[0072] Based on the identified named entities and the relationships between the extracted entities, triples are constructed to form a knowledge graph.

[0073] At this point, the entire process of this method is completed.

[0074] Beneficial effects

[0075] To address the problem of knowledge graph construction, this paper proposes a knowledge graph construction method based on sparse scaling networks and multiple hypergraphs. Compared with the existing technology, it has the following advantages:

[0076] (1) Compared with the named entity recognition method based on sequence annotation, the method can better identify nested named entities belonging to a specific category from unstructured text, extract the semantic relationship between entity pairs, generate text triples, construct a knowledge graph with high accuracy, and provide large-scale knowledge fusion services.

[0077] (2) In response to the problem that existing methods pay less attention to the relationship between redundant information in named entity recognition, the method designs a redundant information processing method based on a sparse scaling network. The method designs a multi-layer convolutional LSTM / GRU model, performs differential connection through soft thresholding, sets corresponding threshold sub-networks, and uses identity paths to alleviate the difficulty of network training. The hidden state of the LSTM / GRU unit is soft-thresholded at each time step. Specifically, the elements in the hidden state that are less than a preset threshold are set to zero, while other elements are retained with reasonable weights. This helps to force the hidden state of the LSTM / GRU to remain relatively sparse, effectively reducing the impact of redundant information and more accurately capturing the key semantic features in the text.

[0078] (3) In order to solve the problems of high computational cost and difficulty in modeling a single hypergraph in the hypergraph method for nested entity recognition, the method designs a method based on multiple hypergraphs to generate local hypergraphs to better identify nested named entities. For each token subsequence, a soft thresholding method is used to generate and merge local hypergraphs in four directions: forward / backward / left / right. The LSTM and GRU units trained by the sparse scaling network are used to perform soft thresholding on their outputs to obtain a sparse representation, and the token subsequences are generated into multiple local hypergraphs in four directions. The convolutional LSTM extracts features from the forward and backward directions to generate hypergraphs, while the gated recurrent unit GRU generates hypergraphs from the left and right directions. Each local hypergraph segment is decoded using depth-first search to obtain its internal entity structure, and then the decoded local hypergraph results in each direction are merged to select a more effective nested entity prediction result.

[0079] (4) The method was tested on a public dataset, and the experimental results demonstrated the effectiveness and superiority of the method. The method has broad application prospects in the fields of question-answering systems, information retrieval, and recommendation systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 This is a flow chart of a knowledge graph construction method based on sparse scaling network and multiple hypergraphs and an embodiment proposed by the present invention. DETAILED DESCRIPTION

[0081] The following describes in detail a preferred implementation of a knowledge graph construction method based on a sparse scaling network and multiple hypergraphs proposed by the present invention in conjunction with an embodiment.

[0082] Example

[0083] This embodiment describes the process of using a knowledge graph construction method based on sparse scaling network and multiple hypergraphs according to the present invention, such as Figure 1As shown in the figure, the named entity labels in the dataset are represented by elements in the set {0, 1, 2, 3, 4, 5, 6, 7}, which has a total of 8 types. 0 represents no entity, and 1, 2, 3, 4, 5, 6, and 7 represent seven types of named entities: VEH, PER, GPE, FAC, ORG, LOC, and WEA. The relationship types in this dataset are represented by elements in the set {11, 12, 13, 14, 15, 16}, which has a total of 6 types. 11, 12, 13, 14, 15, and 16 represent six relationship types: Physical, Part-whole, Personal-Social, ORG-Affiliation, Agent-Artifact, and General-Affiliation.

[0084] from Figure 1 It can be seen that the specific steps include:

[0085] Step 1: Preprocess the text in the dataset, including word segmentation, entity boundary identification, and part-of-speech tagging.

[0086] The specific steps include:

[0087] Step 1: Preprocess the text in the dataset, including word segmentation, entity boundary identification, and part-of-speech tagging.

[0088] The specific steps include:

[0089] Step 1.1: Segment and tag the text sentences in the dataset;

[0090] Step 1.1.1: Clean the dataset, extract sentences and perform word segmentation;

[0091] Step 1.1.2: Tag the parts of speech for each word in the sentence;

[0092] Use StanfordCoreNLP tool to perform part-of-speech tagging on word segmentation;

[0093] Step 1.2: Extract the boundary information of the identified entities in the dataset;

[0094] Step 1.3: Combine the results of steps 1.1 and 1.2 to generate structured data;

[0095] In this embodiment, the three types of information, namely, word segmentation, part of speech and entity boundary, are merged. For example, {"tokens":["A","report","from","CNN","'s","Kelly","Wallace","when","we","return","."],"pos":["DT","NN","IN","NNP","POS","NNP","NNP","WRB","PRP","VBP","."],"entities":[{"type":"ORG","start":3,"end":4},{"type":"PER","start":3,"end":7},{"type":"ORG","start":8,"end":9}]}

[0096] Step 2: Identify the boundaries of entities that may exist in the text and generate candidate boundary markers for the entities;

[0097] The specific steps include:

[0098] Step 2.1: Connect the pre-trained BERT (Bidirectional Encoder Representation from Transformers) encoder Word Embedding Part-of-speech embeddings Character Embedding Four components to construct the word segmentation feature vector z i , the dimension is 2860, as shown in formula (1);

[0099]

[0100] Step 2.2: Transform the eigenvector z i Input to the token-level bidirectional long short-term memory network Bi-LSTM (Bidirectional Long Short-Term Memory) layer, and treat its hidden state as the token representation t i , as shown in formulas (2)-(4);

[0101]

[0102] Step 2.3: Use the feedforward layer to represent t with tokens i As input, calculate the probability that the i-th token constitutes the left or right boundary of the entity As shown in formula (5);

[0103]

[0104] Step 2.4: Use focal loss L l / r It is trained as the minimization objective function of the entity boundary recognition module, as shown in formula (6);

[0105]

[0106] in is the base label of the i-th token. In the multi-hypergraph module, low-quality boundaries can be removed. Therefore, during training, the number of token candidates is scaled by a factor of λ to include some low-quality boundary candidates. N is the number of samples and r is a parameter.

[0107] Step 3: Use the BIOE annotation scheme to annotate each word or character and generate multiple local hypergraph representations in the four directions of forward / backward / left / right;

[0108] The specific steps include:

[0109] Step 3.1: Use left margin information To initialize the unit memory C i,0 and hidden state h i,0 , enhance boundary information, as shown in formula (7);

[0110]

[0111] Step 3.2: Take the token subsequence as input and generate a label at each time step to construct a hypergraph. Use a sparse scaling network structure to capture the important semantic features in the input token sequence, use the feedforward layer to calculate the label prediction probability, obtain the label sequence, and use the hypergraph generation rules GraphRules to expand the forward local hypergraph, as shown in formulas (8)-(12);

[0112]

[0113] p i,k =softmax(FFN(h i,k ))(9)

[0114]

[0115] in, is the input token sequence, h i,k is the hidden state, p i,k is the label prediction probability of the kth token in the subsequence corresponding to the i-th left boundary, is its predicted label, is the label sequence, is The i-th forward local hypergraph generated by implementing the local hypergraph design rules on , where SSN represents sparse scaling network.

[0116] The sparse scaling network uses soft thresholding to perform differential connections on multi-layer convolutional LSTM and GRU (Gated Recurrent Unit). Taking the trained threshold λ as a reference, the elements less than the threshold are set to 0, the elements greater than or equal to the threshold are retained, and reasonable weights are retained according to the conditions of the elements. t , the hidden state h′ after soft thresholding t It is expressed as shown in formula (13);

[0117]

[0118] In the BIOE sequence tagging scheme, B represents the start node, I represents the middle node, E represents the end node, O represents other nodes, and the mark is irrelevant. x represents the predicted entity type x, Indicates the end node of the predicted entity type x.

[0119] The definition of the hypergraph generation rules GraphRules is as follows, specifically including:

[0120] Rule 1: Initialize the local hypergraph to a single B node;

[0121] Rule 2: In each time step, only the last B / I node is used as the head node;

[0122] Rule 3: If cls is predicted x Node and Nodes are connected to the head node via hyperedges;

[0123] Rule 4: When an I / O tag is predicted, the I / O node will be connected to the head node with a normal edge;

[0124] Rule 5: Once an O node is predicted, the local hypergraph stops growing;

[0125] Step 3.3: Similar to the method of generating the forward local hypergraph in step 3.2, the GRU model generates local hypergraphs from the left and right directions respectively, and the LSTM model generates local hypergraphs from the forward and backward directions respectively, thereby forming a set of multiple local hypergraphs;

[0126] Step 3.4: When training hypergraph generation, use cross entropy L g As the objective function to be minimized, as shown in formula (14);

[0127]

[0128] Where ε is the node type set, y i,k is the underlying truth type of the kth token in the subsequence. Starting from the i-th left boundary, is the predicted probability that the kth token in the corresponding token subsequence is labeled as type c, and M is the number of samples.

[0129] Step 4: Decode multiple local hypergraphs to identify nested named entities;

[0130] The specific steps include:

[0131] Step 4.1: Search the local hypergraph sets in each direction from node B to Node paths,decode nested named entities for multiple local hypergraphs;

[0132] Use multiple sets of quadruple E f / b / l / r =(idx l ,idx r ,cls,p f / b / l / r ) represents the predicted named entities in the four directions of forward, backward, left and right. l and idx r Indicates the index of the left and right boundaries, cls indicates the predicted entity type, and p indicates that the last node is E cls The predicted probability of .

[0133] Step 4.2: Combine the forward, backward, left, and right direction sets of each named entity to form the final nested entity prediction result, as shown in formula (15). Where θ is a hyperparameter used as a threshold. If p m ≥θ, the model retains the predicted candidate named entities;

[0134]

[0135] Step 5: Use a multi-layer perceptron based model to learn the mapping from entity grammar features to entity-to-relation types;

[0136] The specific steps include:

[0137] Step 5.1: Transform the text feature vector f(x i ) is input into a multilayer perceptron with a dimension of 1024, and features h are extracted through multiple hidden layers, as shown in formula (16). Where W1 is the weight matrix, b1 is the bias vector, and σ is the ReLu activation function.

[0138] h=σ(W1f(x i)+b1) (16)

[0139] Step 5.2: Input the output h of the hidden layer into the Softmax layer to classify the possible relationship labels in the six relationship types. Where W2 and b2 are the weight matrix and bias vector of the Softmax layer, as shown in formula (17);

[0140] o=softmax(W2h+b2) (17)

[0141] Step 5.3: Use the cross entropy loss function to measure the difference between the predicted result and the true label. The training part updates the model parameters by minimizing the cross entropy loss function, as shown in formula (18);

[0142]

[0143] Where C is the number of categories, y i Is the real label, o i is the output of the model.

[0144] Step 6: Based on the identified named entities and the relationships between the extracted entities, text triples are constructed to form a knowledge graph.

[0145] For example, for the text "In the bustling city of New York, Emily Johnson, talented young architect, works tirelessly at her firm, Johnson & Associates. She has been involved in numerous high-profile projects across the city, including the redesign of Central Park and the construction of the iconic One World Trade Center. Emily'sdedication to her work has earned her recognition and respect in the industry. Despite the challenges she faces, she remains committed to creating innovative and sustainable designs that shape the city's skyline.", the constructed triplet is:

[0146] (Emily Johnson,works at,Johnson&Associates)

[0147] (Emily Johnson,has been involved in,high-profile projects)

[0148] (high-profile projects,including,Central Park)

[0149] (high-profile projects,including,One World Trade Center)

[0150] (Emily Johnson,dedication to,her work)

[0151] (Emily's dedication,earned,recognition and respect)

[0152] (Emily Johnson, committed to creating, innovative and sustainable designs)

[0153] (innovative and sustainable designs, shaping, the city's skyline)

[0154] To illustrate the beneficial effects of the present invention, this experiment uses the same dataset under the same conditions, with the same training set, validation set and test set, and adopts three methods to compare the nested named entity recognition effects.

[0155] The first method is a nested entity extraction model based on a hierarchical pyramid. The second method is a nested entity recognition method based on vocabulary tree parsing. The third method is a knowledge graph construction method based on sparse scaling network and multiple hypergraphs proposed in this invention.

[0156] The evaluation metrics used are: Precision, Recall, and F1-score. Precision measures the proportion of correct entities among recognized entities, while Recall measures the proportion of correctly recognized entities among all correct entities. Generally, Precision and Recall are not discussed separately. F1-score provides a comprehensive consideration of Precision and Recall and is also known as the harmonic mean of Precision and Recall. The definitions of the three evaluation metrics are shown in formulas (19)-(21).

[0157]

[0158] Among them, TP is a true positive entity, which means an entity recognized by NER and consistent with the ground truth; FP is a false positive entity, which means an entity recognized by NER but inconsistent with the ground truth; FN is a false negative entity, which means an entity marked in the ground truth cannot be recognized by NER.

[0159] The results for nested named entity recognition are as follows: the hierarchical pyramid-based nested entity extraction model achieved a precision of 83.95%, a recall of 85.39%, and an F1-score of 84.66%. The lexical tree-based nested entity recognition method achieved a precision of 85.97%, a recall of 87.87%, and an F1-score of 86.91%. The proposed method achieved a precision of 86.40%, a recall of 88.30%, and an F1-score of 87.34%. Experiments demonstrate the effectiveness of the proposed knowledge graph construction method based on sparse scaling networks and multiple hypergraphs.

[0160] The above is only a preferred embodiment of the present invention, and the present invention should not be limited to the contents disclosed in the embodiment and the drawings. Any equivalent or modification completed without departing from the spirit disclosed in the present invention shall fall within the scope of protection of the present invention.

Claims

1. A method for constructing a knowledge graph based on a sparse scaling network and multiple hypergraphs, characterized by: The knowledge graph construction method includes four modules: (1) The first module is the entity boundary recognition module, which uses the sequence labeling framework to decompose the input text into word or character sequences and generate candidate boundary markers for entities. (2) The second module is the multiple hypergraph module, which uses the BIOE annotation scheme to annotate words or characters. This module is trained using a sparse scaling network, which maintains a sparse hidden state through soft thresholding and expands the input token sequence to generate a local hypergraph. The multiple hypergraph module generates local hypergraphs in four directions (forward, backward, left, and right), forming a set of multiple local hypergraphs. (3) The third module is the entity prediction module, which decodes named entities from the local hypergraph sets in four directions and records the predicted probabilities of candidate entities. After all entities are decoded, the candidate entities in different directions are merged using the predicted probabilities to obtain the nested entity recognition results. (4) The fourth module is the relationship extraction module. After entity recognition is completed, the semantic relationship between entities is extracted, triples are generated, and a knowledge graph is constructed. The method for constructing a knowledge graph based on a sparse scaling network and multiple hypergraphs includes the following steps: Step 1: Preprocess the text in the dataset, including word segmentation, entity boundary identification, and part-of-speech tagging. The specific steps include: Step 1.1: Segment and tag the text sentences in the dataset; Step 1.1.1: Clean the dataset, extract sentences and perform word segmentation; Step 1.1.2: Tag the parts of speech for each word in the sentence; Use StanfordCoreNLP tool to perform part-of-speech tagging on word segmentation; Step 1.2: Extract the boundary information of the identified entities in the dataset; Step 1.3: Combine the results of steps 1.1 and 1.2 to generate structured data; Step 2: Identify the boundaries of entities that may exist in the text and generate candidate boundary markers for the entities; The specific steps include: Step 2.1: Connect the pre-trained BERT (Bidirectional Encoder Representation from Transformers) encoder Word Embedding Part-of-speech embeddings Character Embedding Four components to construct the word segmentation feature vector z i , as shown in formula (1); Step 2.2: Transform the eigenvector z i Input to the token-level bidirectional long short-term memory network Bi-LSTM (Bidirectional Long Short-Term Memory) layer, and treat its hidden state as the token representation t i , as shown in formulas (2)-(4); Step 2.3: Use the feedforward layer to represent t with tokens i As input, calculate the probability that the i-th token constitutes the left or right boundary of the entity As shown in formula (5); Step 2.4: Use focal loss L l / r It is trained as the minimization objective function of the entity boundary recognition module, as shown in formula (6); in is the basic label of the i-th token, N is the number of samples, and r is a parameter. Step 3: Use the BIOE annotation scheme to annotate each word or character and generate multiple local hypergraph representations in the four directions of forward / backward / left / right; Step 4: Decode multiple local hypergraphs to identify nested named entities; Step 5: Use a multi-layer perceptron based model to learn the mapping from entity grammar features to entity-to-relation types; The specific steps include: Step 5.1: Transform the eigenvector f(x i ) is input into a multilayer perceptron and passes through multiple hidden layers to extract the feature h, as shown in formula (16). Where W1 is the weight matrix, b1 is the bias vector, and σ is the ReLu activation function. h=σ(W1f(x i )+b1) (16) Step 5.2: Input the output h of the hidden layer into the Softmax layer to classify the possible relationship labels in the relationship type. Where W2 and b2 are the weight matrix and bias vector of the Softmax layer, as shown in formula (17); o=softmax(W2h+b2) (17) Step 5.3: Use the cross entropy loss function to measure the difference between the predicted result and the true label. The training part updates the model parameters by minimizing the cross entropy loss function, as shown in formula (18); Where C is the number of categories, y i Is the real label, o i is the output of the model. Based on the identified named entities and the relationships between the extracted entities, triples are constructed to form a knowledge graph.

2. The method for constructing a knowledge graph based on a sparse scaling network and multiple hypergraphs according to claim 1, characterized in that: Step 3 specifically includes: Step 3.1: Use left margin information To initialize the unit memory C i,0 and hidden state h i,0 , enhance boundary information, as shown in formula (7); Step 3.2: Take the token subsequence as input and generate a label at each time step to construct a hypergraph. Use a sparse scaling network structure to capture the important semantic features in the input token sequence, use the feedforward layer to calculate the label prediction probability, obtain the label sequence, and use the hypergraph generation rules GraphRules to expand the local hypergraph, as shown in formulas (8)-(12); in, is the input token sequence, h i,k is the hidden state, p i,k is the label prediction probability of the kth token in the subsequence corresponding to the i-th left boundary, is its predicted label, is the label sequence, is The i-th forward local hypergraph generated by implementing the local hypergraph design rules on , where SSN represents sparse scaling network. The sparse scaling network uses soft thresholding to perform differential connections on multi-layer convolutional LSTM and GRU (Gated Recurrent Unit). Taking the trained threshold λ as a reference, the elements less than the threshold are set to 0, the elements greater than or equal to the threshold are retained, and reasonable weights are retained according to the conditions of the elements. t , the hidden state h′ after soft thresholding t It is expressed as shown in formula (13); In the BIOE sequence tagging scheme, B represents the start node, I represents the middle node, E represents the end node, O represents other nodes, and the mark is irrelevant. x Indicates that the predicted entity type is x, Indicates the predicted end node of entity type x. The definition of the hypergraph generation rules GraphRules is as follows, specifically including: Rule 1: Initialize the local hypergraph to a single B node; Rule 2: In each time step, only the last B / I node is used as the head node; Rule 3: If cls is predicted x Node and Nodes are connected to the head node via hyperedges; Rule 4: When an I / O tag is predicted, the I / O node will be connected to the head node with a normal edge; Rule 5: Once an O node is predicted, the local hypergraph stops growing; Step 3.3: Similar to the method of generating the forward local hypergraph in step 3.2, the GRU model generates local hypergraphs from the left and right directions respectively, and the LSTM model generates local hypergraphs from the forward and backward directions respectively, thereby forming a set of multiple local hypergraphs; Step 3.4: When training hypergraph generation, use cross entropy L g As the objective function to be minimized, as shown in formula (14); Among them, ε is the node type set, y i,k is the underlying truth type of the kth token in the subsequence. Starting from the i-th left boundary, is the predicted probability that the kth token in the corresponding token subsequence is labeled as type c, and M is the number of samples.

3. The method for constructing a knowledge graph based on a sparse scaling network and multiple hypergraphs according to claim 1, characterized in that: Step 4 specifically includes: Step 4.1: Search from node B to Node paths,decode nested named entities for multiple local hypergraphs; Use multiple sets of quadruple E f / b / l / r =(idx l ,idx r ,cls,p f / b / l / r ) represents the predicted named entities in the four directions of forward, backward, left and right. l and idx r Indicates the index of the left and right boundaries, cls indicates the predicted entity type, and p indicates that the last node is E cls The predicted probability of . Step 4.2: Merge the forward, backward, left, and right direction sets of the named entity to form the final nested entity prediction result, as shown in formula (15). Where θ is a hyperparameter used as a threshold. If p m ≥θ, the model retains the predicted candidate named entities; 。

Citation Information

Cited By

  • Cross-document question and answer method and system based on sparse hypergraph

    CN121412280A