Cultural knowledge graph construction method based on W2ner model

By adopting the W2ner model and large language model in the construction of cultural knowledge graph, the problem of nested entities and non-continuous entities recognition in cultural texts is solved, the flexibility and coverage of relationship extraction are improved, high-quality cultural knowledge graph construction is achieved, and the deep application of cultural data is supported.

CN120196768AActive Publication Date: 2025-06-24NANJING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510673642.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

When constructing a cultural knowledge graph, it is difficult for the existing technology to effectively deal with nested entities and discontinuous entities commonly found in cultural texts, and traditional methods cannot adapt to the open and complex characteristics of cultural relationship types, resulting in insufficient accuracy and completeness of entity recognition and relationship extraction.

Method used

Using the cultural knowledge graph construction method based on the W2ner model, by modeling the named entity recognition task as a relationship classification problem between characters, the input sentence is processed using the pre-trained model BERT and the bidirectional LSTM network, character vector representations containing context information are generated, and interactive features of different distance character pairs are captured through multi-layer perceptron and multi-grained expansion convolution. At the same time, with the help of the powerful semantic understanding ability of the large language model, open domain relationship extraction is carried out, and effective merging of relational phrases is realized through the K-means clustering algorithm.

Benefits of technology

This method can extract physical information from cultural text more accurately and comprehensively, improve the flexibility and coverage of relationship extraction, reduce the redundancy of knowledge graphs, improve the structured quality of knowledge graphs, and support the deep application and digital protection of cultural data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196768A_ABST
    Figure CN120196768A_ABST
Patent Text Reader

Abstract

The invention discloses a culture knowledge graph construction method based on a W2ner model, and belongs to the technical field of culture and artificial intelligence. S2, performing complex entity identification on the collected data based on a W2ner model; s3, carrying out open domain relation extraction based on a large language model; s4, realizing effective combination of relation phrases by using a K-means clustering algorithm, and performing relation standardization; s5, storing the standardized knowledge triple obtained after identification, extraction and alignment into a database, and performing visual output; according to the method, special technical modules for complex entity recognition, open domain relation extraction and relation standardization are integrated, and a knowledge graph construction technical scheme which is more suitable for cultural text characteristics and higher in automation degree is formed in combination with graph database storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of culture and artificial intelligence, and particularly relates to a method for constructing a cultural knowledge graph based on the W2ner model. Background Art

[0002] Cultural data (such as historical documents, local chronicles, art reviews, folk records, etc.) contains rich information. However, most cultural data is unstructured data. Although it contains rich information, due to its messy format and lack of a clear organizational structure, it makes the retrieval, understanding, analysis, and in-depth application of the data extremely difficult. A knowledge graph is a data representation technology based on a semantic network that can transform unstructured text into structured knowledge. Its basic structure consists of entities, relations, and attributes. Through the triple, which is the basic representation unit, the objects, concepts, and their connections in the real world are organized into a networked knowledge system. By clearly expressing the relationships between entities, the knowledge graph can effectively improve the retrieval efficiency of cultural data and the depth of information understanding, thereby supporting the digital protection, research, and dissemination of cultural heritage.

[0003] In recent years, pre-trained models represented by large language models (LLMs) have achieved major breakthroughs in the field of artificial intelligence. Their powerful context understanding, semantic reasoning, and text generation capabilities endow them with great potential to automatically extract semantic knowledge, identify entities and associated relationships from cultural data with a very large domain span, providing a promising new technical paradigm for overcoming the current bottlenecks in knowledge graph construction and achieving more efficient and large-scale automatic construction.

[0004] Currently, knowledge graphs are divided into domain knowledge graphs and general knowledge graphs. The general knowledge graph is a knowledge graph oriented to the whole domain, aiming to cover a wide range of common sense knowledge. Due to its wide coverage, the general knowledge graph emphasizes the breadth of entities; however, the general knowledge graph lacks domain depth and fine-grained semantic representation capabilities, and the professional data is sparse and of low quality.

[0005] Two core steps in the automated process of constructing a cultural knowledge graph are named entity recognition (NER) and relation extraction (RE). Named entity recognition technology is used to automatically identify specific types of entities such as person names, place names, and organization names from text. This technology has evolved from early rule- and dictionary-based methods to methods based on statistical machine learning (such as HMM, SVM, CRF), and in recent years, deep learning-based models have become the mainstream. Usually, the NER task is treated as a sequence labeling problem, whose principle is to assign a predefined class label (such as B-PER, I-PER, O) to each basic unit (such as a character) in the input text sequence, and determine the boundaries and types of entities by parsing the label sequence. However, this sequence labeling method exposes inherent defects when dealing with complex cultural domain texts. Cultural texts often contain nested entities (one entity contains another entity) and discontinuous entities (text fragments of the same entity are not adjacent). Due to the mechanism of sequence labeling that assigns a unique label to each character, in principle, it cannot effectively represent that a character belongs to multiple nested entities, or associate and recognize non-adjacent character fragments as the same entity, thus resulting in poor recognition performance in these common scenarios.

[0006] Relation extraction technology is used to identify and classify the semantic relationships existing between the identified entity pairs. Its technology has also gone through a development path from pattern matching to machine learning (including supervised learning that requires a large amount of labeled data and unsupervised clustering that does not rely on labeling), and then to deep learning-based models. However, applying these technologies in the cultural domain also faces challenges. The relationships between cultural entities often have open types, complex semantics, and are closely related to the deep cultural and historical background, which makes traditional methods that rely on predefined relation categories and a large amount of labeled data difficult to apply. Due to the flexibility of natural language expression, the same semantic relationship may be expressed by different words or phrases (synonymous relationships). If these semantically equivalent but differently expressed relationships are not effectively merged, it will lead to a large amount of redundant information in the finally constructed knowledge graph, reducing the quality and usability of knowledge.

[0007] Moreover, when existing technologies attempt to automatically construct a knowledge graph from unstructured texts in the cultural domain, it is found that the mainstream named entity recognition technology, especially the method based on sequence labeling, has an internal mechanism that is difficult to effectively handle the common nested entity and discontinuous entity structures in cultural texts, resulting in insufficient accuracy and integrity of entity recognition. Traditional methods that rely on predefined relation patterns cannot adapt to the characteristics of open relation types, complex semantics, and high dependence on background knowledge in the cultural domain. If not addressed, it will seriously affect the quality and consistency of the final knowledge graph. Summary of the Invention

[0008] In view of the problems mentioned in the background art, the present invention proposes a method for constructing a cultural knowledge graph based on the W2ner model. The core purpose is to overcome the limitations of the prior art in dealing with cultural knowledge of specific regions. By constructing a structured, networked, and semantic domain knowledge graph, the systematic integration, refined management, and convenient application of the cultural heritage knowledge system are realized.

[0009] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0010] A method for constructing a cultural knowledge graph based on the W2ner model, comprising the following steps:

[0011] S1: Data collection;

[0012] S2: Perform complex entity recognition on the collected data based on the W2ner model;

[0013] S3: Perform open-domain relation extraction based on the large language model;

[0014] S4: Use the K-means clustering algorithm to effectively merge relation phrases and perform relation normalization;

[0015] S5: Store the normalized knowledge triples obtained after recognition, extraction, and alignment into the database and perform visual output.

[0016] Preferably, in S1, text data is obtained from historical documents, news reports, statistical yearbooks, academic papers, Baidu Encyclopedia, and cultural encyclopedias.

[0017] Preferably, in S2, the entity recognition task is re-modeled as a relation classification problem between characters. The specific implementation process is as follows:

[0018] S21: Define the relation categories between any two characters x i and x j in the text;

[0019] S22: Use the pre-trained model BERT and the bidirectional LSTM network to process the input sentence and generate a character vector representation containing context information;

[0020] S23: Use conditional layer normalization to convert the character vector representation into a two-dimensional character pair relation grid, and then capture the interaction features of character pairs at different distances on this grid through a multi-layer perceptron and a multi-granularity dilated convolution to form the final character pair feature representation Q;

[0021] S24: Based on the character pair feature representation Q, use a multi-layer perceptron and a bi-affine predictor to predict each pair of characters (x i , x jPredefined relationship categories; and decode the relationships to output the final entity recognition results.

[0022] Preferably, in S21, the relationship categories include adjacent relationship, head-tail relationship, and no relationship;

[0023] Adjacent relationship: x i and x j are adjacent and belong to the same entity instance;

[0024] Head-tail relationship: If x i is the last character of entity A, and x j is the first character of entity A, then there is a head-tail relationship between (x i , x j ).

[0025] Preferably, in S23, the specific implementation process is as follows:

[0026] First, construct an initial character pair grid and perform conditional layer normalization;

[0027] Then perform MLP mapping, and for each grid point perform dimensionality reduction and non-linear transformation through a multi-layer perceptron, specifically:

[0028] ;

[0029] ;

[0030] ;

[0031] where F i,j represents the dimensionality-reduced feature output by the multi-layer perceptron for the (i, j) grid point, d f represents the output dimension of the multi-layer perceptron, W (1) , b (1) both represent the parameters of the first layer of the multi-layer perceptron; W (2) , b (2) both represent the parameters of the second layer of the multi-layer perceptron; d m represents the hidden layer dimension; represents the input vector of the (i, j) grid point; U i,j represents the output of the first layer of the multi-layer perceptron for the (i, j) grid point; ReLU represents the activation function; x represents the input of the linear transformation; X represents the output value of the activation function;

[0032] Finally, perform multi-granularity dilated convolution.

[0033] Preferably, the specific content of the multi-granularity dilated convolution is:

[0034] Let the dilation rate of the k-th granularity be dk , the convolutional kernel Performs a convolution operation on the input downsampled feature F and the dilation convolutional granularity K, specifically:

[0035] ;

[0036] Among them, d k represents the k-th dilation rate, k h , k w respectively represent the height and width of the convolutional kernel, d g represents the number of channels output by each dilated convolution, represents the convolutional kernel weight of the k-th granularity, and u and v represent two indices inside the convolutional kernel;

[0037] Performs convolutions on multiple granularities d1, d2, d3 to obtain G1, G2, G3, specifically:

[0038] ;

[0039] Among them, K represents the number of convolutional kernels, and k represents the current convolutional kernel; represents the result vector of the convolution at the output grid point (i, j) under the k-th dilation rate; d g represents the number of channels output by each dilated convolution;

[0040] Finally, sum to obtain Q i,j , which is the final interaction feature representation of each character pair (i, j) in the model.

[0041] Preferably, in S24, the specific implementation process is:

[0042] First, project through a multi-layer perceptron to obtain head and tail representations. For each position pair (i, j) and the existing feature vector Q i,j , respectively use two independent single-layer MLPs to obtain head and tail representations;

[0043] Then perform a bi-affine scoring. For the r-th type of relationship, calculate the affine score using the parameters W and U, specifically:

[0044] ;

[0045] Among them, || represents vector concatenation, W (r) , U (r) respectively represent the bi-affine parameter and bias, b (r) represents the bias scalar; represents the original score that the position pair (i, j) belongs to the r-th type; u i,j represents the head representation vector; v i,j represents the tail representation vector; T represents the transpose operator;

[0046] Finally, for relation classification and training loss, apply Softmax to all R-class scores at the position pair (i, j) to obtain a probability distribution, and then obtain the predicted relation category for each position pair (i, j).

[0047] Preferably, in S3, the specific implementation process is as follows:

[0048] S31: Incorporation and expansion of entity information;

[0049] S32: Implementation of the three-step prompting strategy based on the incorporation and expansion of entity information;

[0050] S33: Provide the entity information identified in S2 as a reference input to the large language model for entity recognition correction.

[0051] Preferably, in S4, the specific content of applying the K-means clustering algorithm to the vector representations of all relation phrases to reduce the relation types is as follows:

[0052] S41: Randomly select k samples as the initial cluster centers;

[0053] S42: For each sample x i , find the nearest centroid and assign the label c;

[0054] S43: For each cluster, use the mean of all samples in the cluster as the new centroid.

[0055] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0056] (1) The present invention uses the W2ner model to transform the named entity recognition task from the traditional sequence labeling paradigm to the word-word relation classification paradigm. This method can directly model and effectively identify the nested entities and discontinuous entities commonly existing in cultural texts, overcoming the recognition bottleneck of traditional sequence labeling methods for these complex structures. Experimental results show that this method has better performance than the baseline sequence labeling model when processing cultural data sets containing such entities, and can extract entity information more accurately and comprehensively.

[0057] (2) By virtue of the powerful semantic understanding ability of the large language model (LLM) and the "three-step method" prompting strategy, the present invention can automatically extract diverse and deep semantic relationships from cultural texts without predefined relationship patterns. This greatly improves the flexibility and coverage of relationship extraction, and is particularly applicable to the cultural field where the relationship types are complex and difficult to enumerate in advance. At the same time, integrating the entities identified by W2ner as reference information into the large language model helps to guide the LLM and may correct errors in the NER stage, improving the robustness of end-to-end information extraction. The present invention introduces the K-means clustering method to effectively merge relationship phrases, significantly reducing the redundancy of the knowledge graph, improving the structured quality of the knowledge graph, and facilitating subsequent efficient query and knowledge management.

[0058] (3) The present invention integrates specialized technical modules for complex entity recognition, open-domain relationship extraction, and relationship normalization, and combines graph database storage to form a knowledge graph construction method that is more adaptable to the characteristics of cultural texts and has a higher degree of automation. This method provides a more effective and accurate technical approach for deeply mining and structurally utilizing massive unstructured cultural heritage data resources, effectively supporting the deep application and digital protection of cultural data. Brief Description of the Drawings

[0059] Figure 1 is a flowchart of the cultural knowledge graph construction method based on the W2ner model of the present invention;

[0060] Figure 2 is a schematic diagram of the present invention modeling entity recognition as a relationship between words;

[0061] Figure 3 is a schematic diagram of the overall architecture of W2ner of the present invention;

[0062] Figure 4 is a schematic diagram of the three-step method prompting strategy of the present invention;

[0063] Figure 5 is a schematic diagram of the performance comparison of different models of the present invention;

[0064] Figure 6 is a schematic diagram of the difference in the performance of simple entities and complex entities of W2ner of the present invention. Detailed Embodiments

[0065] The following further clarifies the present invention in conjunction with specific embodiments. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0066] The method for constructing a cultural knowledge graph based on the W2ner model provided in this embodiment solves the deficiencies in complex entity recognition, open-domain relation extraction, and relation synonymy processing when constructing a knowledge graph for cultural domain texts. The present invention proposes a knowledge graph construction method that integrates deep learning, large language models (LLMs), and clustering algorithms; aiming to automatically and efficiently extract entities and relationships from unstructured cultural texts and construct a high-quality cultural knowledge graph. Specifically, it mainly includes the following steps:

[0067] S1: Data collection;

[0068] As Figure 1 shown, obtain text data from historical documents, news reports, statistical yearbooks, academic papers, Baidu Encyclopedia, and cultural encyclopedias.

[0069] S2: Perform complex entity recognition on the collected data based on the W2ner model;

[0070] Construct a cultural ontology-instance classification concept structure hierarchy through the Protégé tool.

[0071] This step aims to solve the problem of identifying nested entities and discontinuous entities in cultural texts. Its core idea is to abandon the traditional sequence annotation method and re-model the entity recognition task as a relation classification problem between characters (Tokens).

[0072] As Figure 1 shown, in this embodiment, entities include literature and art, celebrities, geography, scenic spots and historical sites, products, academics and literature, science and technology and education, history, customs, and clans; among them, ancient towns include Dongpu Ancient Town, Doumen Ancient Town, Jin Dongmen Old Street, Luzhi, Mudu, Zhouzhuang, Wuzhen, Wanan, Tongli, Shimen Town, Shengze, Shagou Ancient Town, Qinglong Town, Pingwang, Yansi, Chengkan, Dangkou Town, and Daohe Ancient Street.

[0073] S21: Define the relationship categories between any two characters x i and x j in the text;

[0074] First, define the possible relationship categories related to entity boundaries and types between any two characters x i and x j in the text.

[0075] As Figure 2 shown, the following main types of relationships are included in this embodiment:

[0076] Adjacent relationship (Next-Neighboring-Word, NHW): If x i and x jIf they are adjacent and belong to the same entity instance, there is an adjacent relationship between them.

[0077] Head-tail relationship ( ):If x i is the last character (Tail) of entity A, and x j is the first character (Head) of entity A, then there is a head-tail relationship between (x i , x j ), where represents the type of entity A (such as PER, LOC, ORG, etc.). This relationship connects the head and tail of the entity, even if they are not adjacent (handling non-continuous entities).

[0078] No relationship (NONE): There is no above-mentioned relevant relationship between the character pair.

[0079] Among them, Ye Gui and Zi Tianshi are both entities. There is an adjacent relationship and a head-tail relationship between the character Ye and the character Gui, an adjacent relationship between the character Ye and the character Tian, an adjacent relationship between the character Tian and the character Shi, and a head-tail relationship between the character Shi and the character Ye.

[0080] S22: The W2ner model first uses the pre-trained model BERT and the bidirectional LSTM network to process the input sentence and generate a character vector representation containing context information.

[0081] The BERT encoding process mainly includes three parts, namely position encoding (Position embeddings), sentence encoding (Segmentation embeddings), and word embedding encoding (Token ebeddings). Among them, the position encoding is used to mark the position information of the character in the sentence, the sentence encoding is responsible for determining the sentence information where the character is located, and the word embedding encoding is the feature vector of the character. Finally, the three parts of the encoding are summed as the final feature vector, thus completing the high-dimensional representation of natural semantics.

[0082] The tokenized input sequence of length n will have three different representations, specifically:

[0083] Token embedding, shape (1, n, 768), which is the vector representation of the word.

[0084] Segment embedding, shape (1, n, 768), which is the vector representation of the sentence.

[0085] Position embedding, shape (1, n, 768), which is the vector representation of the position.

[0086] In this embodiment, the BiLSTM (Bidirectional Long Short-Term Memory Network) structure is used. BiLSTM adds a reverse operation on the basis of the unidirectional LSTM. The forward LSTM is used to capture valuable text information from the previous context and transfer it to the current moment. The backward LSTM is used to obtain the influence of the subsequent context on the current moment, so that the model can consider the context information. In LSTM, there are three gate structures in total, namely the forget gate, the input gate, and the output gate.

[0087] 1) Forget gate;

[0088] For the cell state in the previous moment of LSTM, some information may become "obsolete". To prevent excessive memory from affecting the neural network's processing of the current input, some components in the previous cell state are selectively forgotten.

[0089] ;

[0090] Among them, f t represents the output vector of the sigmoid neural layer; represents the activation function sigmoid; W f represents the weight vector of the current step; h t-1 represents the output value of the previous moment of LSTM; x t represents the input of the network at the current moment; b f represents the bias term of the forget gate.

[0091] 2) Input gate;

[0092] Through the input gate layer, combined with the sigmoid function, it is determined which values are used for update; the tanh layer is used to generate new candidate values to be added to obtain the candidate values; finally, the update of the cell state C is completed; the calculation formula of the input gate is as follows:

[0093]

[0094]

[0095]

[0096] Among them, i t represents the acceptance weight; b t represents the bias term of the acceptance weight; W f represents the weight vector of the current step; represents the candidate state; represents an activation function that outputs a real value in (-1, 1); W cdenotes the weight vector (Weight); b c denotes the bias term; C t denotes the current cell state (after update); f t denotes the output of the forget gate; x t denotes the input to the network at the current time step; denotes the activation function sigmoid; h t-1 denotes the output value of the LSTM at the previous time step.

[0097] 3) Output gate;

[0098] The forget gate and the input gate are processes that remove unnecessary information and add new information to finally update the cell state to achieve the output. The role of the output gate is to output the input of the LSTM at the current step.

[0099] ;

[0100] ;

[0101] Among them, O t denotes the information obtained by extracting the vector after integrating the current input value and the output value at the previous time step using the sigmoid function; h t denotes the output of the model; denotes the compression process of the previously learned information, which plays a role in stabilizing the numerical value; b o denotes the bias term; W o denotes the weight vector; denotes the activation function sigmoid; x t denotes the input to the network at the current time step; h t-1 denotes the output value of the LSTM at the previous time step.

[0102] Concatenate the features at two time steps as the output vector.

[0103] The forward LSTM inputs "Jiang", "Nan", "Wen", "Hua" in sequence to obtain four vectors {h L0 , h L1 , h L2 , h L3}. The backward LSTM inputs "Hua", "Wen", "Nan", "Jiang" in sequence to obtain four vectors {h R0 , h R1 , h R2 , h R3}. Finally, concatenate the forward and backward hidden vectors to get {[h L0 , h R3 , [h L1 , h R2 , [h L2 , hR1 , [h L3 , h R0}), that is, {h0, h1, h2, h3}. For the entity recognition task, the sentence representation adopted in this embodiment is [h L2 , h R2 , which contains forward and backward information.

[0104] S23: Subsequently, the character vector representation is converted into a two-dimensional character pair relationship grid through conditional layer normalization (CLN), which integrates character information, relative position information, and region information. Then, a multi-layer perceptron (MLP) and multi-grained dilated convolution are applied to capture the interaction features of character pairs at different distances on this grid to form the final character pair feature representation Q. Multi-grained dilated convolution is also known as multi-grained dilation convolution.

[0105] As Figure 3 shown, the specific steps are as follows:

[0106] Step 1: Construct an initial character pair grid and perform conditional layer normalization (CLN);

[0107] Assume the sequence length is n, and each character passes through BiLSTM / BERT to obtain a vector H. The specific calculation formula of H is:

[0108] ;

[0109] For any position (i, j), the following four vectors are directly concatenated along the channel dimension, specifically:

[0110] ;

[0111] Among them, H i and H j are character representations, is the relative position encoding, is the region information encoding; || represents vector concatenation, i, j represent the position indices of any pair of characters in the sequence, and n represents the sequence length. d x represents the total dimension of the vector X i,j ; d h represents the dimension of the vector obtained by each character after BERT + BiLSTM; d p represents the dimension of the relative position encoding; d r represents the dimension of the region encoding.

[0112] Then, conditional layer normalization is performed on each character pair (i, j), specifically:

[0113] ;

[0114] ;

[0115] where μ represents the mean calculated over all n 2 Xs i,j in the vector dimension, and σ represents the standard deviation (including a stability term, which is a small constant to prevent division by zero); i and j represent the position indices of any pair of characters in the sequence, and n represents the sequence length.

[0116] ;

[0117] where CLN represents the conditional layer normalization operation; represents element-wise multiplication; μ represents the mean calculated over all n 2 Xs i,j in the vector dimension, and σ represents the standard deviation; the scaling parameter γ and the offset parameter β are dynamically generated according to the actual conditional vector C. d x is the dimension of X i,j , and this operation is parallel over all (i, j), outputting a normalized grid, specifically:

[0118] ;

[0119] where represents the character pair feature grid obtained after (conditional) layer normalization; n represents the sequence length; d x represents the vector dimension at each grid point, i.e., the number of channels of the concatenated vector X i,j before normalization.

[0120] Step 2: MLP mapping: For each grid point perform dimensionality reduction and non-linear transformation through a two-layer multi-layer perceptron (MLP), specifically:

[0121] ;

[0122] ;

[0123] ;

[0124] where F i,j represents the dimensionality-reduced feature output by the MLP at the (i, j) grid point, d f represents the MLP output dimension, and W (1) , b (1) represent the parameters of the first layer of the MLP; W (2) , b (2) represent the parameters of the second layer of the MLP. dm denotes the dimension of the hidden layer, i.e., the length of the output vector of the first layer; U denotes the output of the first layer of the MLP; denotes the input vector of the (i, j) grid point; U i,j denotes the output of the first layer of the MLP at the (i, j) grid point; ReLU denotes the activation function; x denotes the input of the linear transformation of this layer. X denotes the output value of the activation function.

[0125] Step 3: Multi-granularity dilated convolution;

[0126] Perform multiple groups of different dilation 2D convolutions on F. Let the dilation rate of the k-th granularity be d k , and the convolution kernel K (k) Perform a convolution operation on the input F and K, specifically:

[0127] ;

[0128] where K represents the number of dilation convolution granularities, d k represents the k-th dilation rate, k h , k w respectively represent the height and width of the convolution kernel, d g represents the number of channels of each dilation convolution output, K (k) represents the convolution kernel weight of the k-th granularity, and u and v are two indices inside the convolution kernel used to traverse the convolution kernel of size k h ×k w .

[0129] Perform convolutions on all multi-granularities d1, d2, d3 to obtain G1, G2, G3, specifically:

[0130] ;

[0131] where K is the number of convolution kernels and k is the current convolution kernel. represents the result vector of the convolution at the output grid point (i, j) at the k-th dilation rate; d g represents the number of channels of each dilation convolution output.

[0132] Finally, sum to obtain Q i,j which is the final interaction feature representation of each character pair (i, j) in the model.

[0133] S24: Finally, based on the character pair feature representation Q, combine the use of MLP and the Biaffine Predictor to predict the predefined NER relationship categories (such as NHW, i , x j ) between each pair of characters (x , NONE, etc.).

[0134] The predicted relationship matrix output by the model can be converted into the final entity recognition result through decoding (identifying specific relationship paths).

[0135] Step 1: Project the MLP to the head and tail representations;

[0136] For each position pair (i,j) and the existing feature vector Q i,j, Use two independent single-layer MLPs (with activation) to obtain the "head" and "tail" representations respectively.

[0137] ;

[0138] ;

[0139] ;

[0140] where W (h) , b (h) are the weight matrix and bias vector for the head mapping; W (t) , b (t) are the weight matrix and bias vector for the tail mapping, u i,j is the head representation vector; v i,j is the tail representation vector. ReLU represents the activation function; x represents the input of the linear transformation of this layer. X represents the output value of the activation function. d u represents the dimension of the head and tail representation vectors.

[0141] Step 2: Biaffine scoring;

[0142] In this embodiment, 3 relationship categories are predefined. For the r-th type of relationship, the affine score is calculated using the parameters W and U. The specific values of W and U are obtained through training.

[0143] ;

[0144] where || represents vector concatenation, W (r) , U (r) are the biaffine parameters and biases, and b (r) is the bias scalar. is the raw score that the position pair (i,j) belongs to the r-th class. u i,j is the head representation vector; v i,j is the tail representation vector. T represents the transpose operator.

[0145] Step 3: Relationship classification and training loss;

[0146] Perform Softmax on all R-class scores at (i,j) to obtain the probability distribution, specifically:

[0147] ;

[0148] where exp is the exponential function e x , is the probability that (i, j) is predicted as relation r, and R is the number of relation categories. is the raw score that the position pair (i, j) belongs to the r-th class; represents the raw score that the position pair (i, j) belongs to the r-th class, and is used to represent traversing all possible relation categories.

[0149] In the prediction stage, for each (i, j), take:

[0150] ;

[0151] where is the probability that (i, j) is predicted as relation r; is the predicted relation category, argmax is an operator, taking the r that makes the largest and regarding all character pairs (i, j) of (NA represents the class label for no relation) as having an entity from character x i to x j with the entity type being .

[0152] S3: Conduct open-domain relation extraction based on large language models;

[0153] This step utilizes the powerful understanding ability of large language models (LLMs) to solve the problem of open and difficult-to-predefine relation types in the cultural field and attempts to mitigate the errors brought by the pipeline method.

[0154] S31: Incorporation and expansion of entity information;

[0155] As Figure 4 shown, it includes entity information incorporation, entity type supplementation, hidden information reasoning, and error propagation correction.

[0156] S32: Implementation of a three-step prompting strategy based on the incorporation and expansion of entity information;

[0157] To guide the LLM to perform open-domain relation extraction that meets the requirements, the method of the present invention adopts a "three-step" prompting (Prompt) strategy in Prompt design and construction.

[0158] Requirement prompting: Clearly define the task objective, output format, and entity representation specification.

[0159] Input format prompting, unit consistency, output format prompting.

[0160] Domain knowledge hint: Incorporate necessary background knowledge and expert experience in the cultural domain.

[0161] Case hint: Guide the model to learn how to extract high-quality and semantically diverse relationships through a small number of typical examples (instance extraction results).

[0162] S33: Provide the entity information identified in S2 as a reference input to the large language model for entity recognition correction;

[0163] Provide the entity information identified by the W2ner model as a reference input to the LLM. Emphasize that it is for reference, allowing the LLM to correct entity recognition errors based on the context, further improving the accuracy of relationship extraction and reducing error propagation.

[0164] S4: Use the K-means clustering algorithm to effectively merge relationship phrases for relationship normalization;

[0165] Apply the K-means clustering algorithm to the vector representations of all relationship phrases to reduce the number of relationship types.

[0166] S41: Randomly select k samples as the initial cluster centers ;

[0167] where K represents the number of clusters.

[0168] S42: For each sample x i , find the closest centroid and assign the label c, specifically:

[0169] ;

[0170] where represents the centroid of the j-th cluster after the t-th iteration, and || represents the vector norm; represents the cluster label assigned to the sample x i in the t-th iteration; K represents the number of clusters; represents the centroid of the j-th cluster after the (t - 1)-th iteration.

[0171] S43: For each cluster j, use the mean of all samples in the cluster as the new centroid, specifically:

[0172] ;

[0173] ;

[0174] where represents the set of sample indices included in cluster j in the t-th iteration. x i represents the feature vector of the i-th sample. It represents the centroid of the j-th cluster after the t-th iteration.

[0175] If the centroid or the sample labels do not change (or the change is less than the threshold) in this iteration, stop; otherwise, let Repeat Step 2 and Step 3.

[0176] In this embodiment, the K-means clustering algorithm is used to effectively merge relationship phrases. The results of relationship normalization can be seen in Table 1 below.

[0177] Table 1 Partial K-means Relationship Clustering Results

[0178]

[0179] Table 1 shows partial K-means relationship clustering results, where the extracted words with synonymous relationships are merged; for example, the four words "occupation", "position", "function", and "hold a position" are merged into "position".

[0180] S5: Finally, store the normalized knowledge triples (entity, relationship, entity) and entity attributes obtained after recognition, extraction, and alignment into the graph database to complete the construction of the knowledge graph in the cultural field; obtain the knowledge graph in the cultural field and visualize and output it.

[0181] The present invention uses three basic indicators, namely precision, recall, and F1 score, to evaluate the results of entity extraction and relationship extraction.

[0182] Precision measures how many of the results identified as positive examples by the model are truly positive examples.

[0183] Recall measures how many of all the truly positive examples are successfully identified by the model.

[0184] The F1 score is the harmonic mean of precision and recall, comprehensively measuring the accuracy and integrity of the model.

[0185] The experimental results are as Figure 5As shown, the named entity recognition model based on W2ner proposed by the present invention outperforms the sequence labeling models using BERT+CRF or BERT+BiLSTM+CRF encoders in the entity recognition task in the cultural field. Compared with the way of multi-label classification for each character of the named entity in sequence labeling, W2ner adopts a unique character-character relationship classification method, aiming to identify consecutive, nested and discontinuous entities in the text at the same time. It uses multiple two-dimensional dilated convolutions to capture the semantic relationships between character pairs at different distances, and the method of defining two main relationships, NHW and THW, to identify entity boundaries and entity types can more accurately determine entity boundaries and entity types.

[0186] The experimental results show that the method proposed by the present invention not only has higher performance than the commonly used sequence labeling models, but also can identify complex entities such as nested and irregular types. Generally speaking, W2ner solves well the recognition of various entities in cultural data.

[0187] As Figure 6 shown, for simple entities (label) and complex entities (nested, irregular entities) (Entity), the F1 score gradually increases with the increase of the number of training epochs (Epoch). This indicates that the model becomes gradually more accurate and complete in identifying flat entities and complex entities. Generally speaking, the method proposed by the present invention effectively identifies simple entities and complex entities in cultural data.

[0188] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for constructing a cultural knowledge graph based on the W2ner model, characterized in that: It includes the following steps: S1: Data collection; S2: Perform complex entity recognition on the collected data based on the W2ner model; S3: Conduct open-domain relation extraction based on the large language model; S4: Use the K-means clustering algorithm to effectively merge relation phrases and perform relation normalization; S5: Store the normalized knowledge triples obtained after recognition, extraction, and alignment into the database and perform visual output.

2. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 1, wherein: In S1, text data is obtained from historical documents, news reports, statistical yearbooks, academic papers, Baidu Encyclopedia, and cultural encyclopedias.

3. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 1, characterized in that: In S2, the entity recognition task is re-modeled as a relation classification problem between characters. The specific implementation process is as follows: S21: Define the relationship category between any two characters x i and x j in the text; S22: Use the pre-trained model BERT and the bidirectional LSTM network to process the input sentence and generate a character vector representation containing context information; S23: Use conditional layer normalization to convert the character vector representation into a two-dimensional character pair relation grid, and then capture the interaction features of character pairs at different distances on this grid through a multi-layer perceptron and multi-granularity dilated convolution to form the final character pair feature representation Q; S24: Based on the character pair feature representation Q, use a multi-layer perceptron and a bi-affine predictor to predict the predefined relationship categories between each pair of characters (x i , x j ); and decode the relationships to output the final entity recognition results.

4. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 3, characterized in that: In S21, the relation categories include adjacent relation, head-tail relation, and no relation; Adjacent relationship: x i and x j are adjacent and belong to the same entity instance; Head and tail relationship: If x i is the last character of entity A, and x j is the first character of entity A, then there is a head and tail relationship between (x i , x j ).

5. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 3, wherein: In S23, the specific implementation process is as follows: First, construct the initial character pair grid and perform conditional layer normalization; Then perform MLP mapping on each lattice point Perform dimensionality reduction and non-linear transformation through a multi-layer perceptron, specifically: ; ; ; Among them, F i,j represents the dimensionality-reduced feature output by the multi-layer perceptron for the (i, j) grid point, d f represents the output dimension of the multi-layer perceptron, W (1) , b (1) both represent the parameters of the first layer of the multi-layer perceptron; W (2) , b (2) both represent the parameters of the second layer of the multi-layer perceptron; d m represents the hidden layer dimension; represents the input vector for the (i, j) grid point; U i,j represents the output of the first layer of the multi-layer perceptron for the (i, j) grid point; ReLU represents the activation function; x represents the input of the linear transformation; X represents the output value of the activation function; Finally, perform multi-granularity dilated convolution.

6. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 5, wherein: The specific content of the multi-granularity dilated convolution is: Let the dilation rate of the k-th granularity be d k , and the convolutional kernel K (k) Perform a convolution operation on the input dimensionality-reduced feature F and the dilation convolutional granularity number K, specifically: ; Among them, d k represents the k-th dilation rate, where k h , k w respectively represent the height and width of the convolutional kernel, and d g represents the number of channels output by each dilated convolution, and K (k) represents the convolutional kernel weight of the k-th granularity, and u and v represent two indices inside the convolutional kernel; Perform convolution on multi-granularities d1, d2, d3 to obtain G1, G2, G3 respectively, specifically: ; Among them, K represents the number of convolutional kernels, and k represents the current convolutional kernel; represents the result vector of convolution at the output grid point (i, j) under the k-th dilation rate; d g represents the number of channels output by each dilated convolution; Finally, the sum is obtained to get Q i,j , which is the final interaction feature representation for each character pair (i, j) in the model.

7. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 3, wherein: In S24, the specific implementation process is as follows: First, project through a multi-layer perceptron to obtain head and tail representations. For each position pair (i, j) and the existing feature vector Q i,j , use two independent single-layer MLPs to obtain head and tail representations respectively; Then perform bi-affine scoring. For the r-th type of relation, calculate the affine score using parameters W and U, specifically: ; Among them, || represents vector concatenation, W (r) , U (r) respectively represent bi-affine parameters and biases, and b (r) represents a bias scalar; represents the raw score that the position pair (i, j) belongs to the r-th class; u i,j represents the head representation vector; v i,j represents the tail representation vector; T represents the transpose operator; Finally, perform relation classification and training loss. Do Softmax on all R-class scores at the position pair (i, j) to obtain the probability distribution, and then obtain the predicted relation category for each position pair (i, j).

8. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 1, wherein: In S3, the specific implementation process is as follows: S31: Incorporation and expansion of entity information; S32: Implement the three-step prompting strategy based on the incorporation and expansion of entity information; S33: Provide the entity information identified in S2 as a reference input to the large language model for entity recognition correction.

9. The method for constructing a cultural knowledge graph based on the W2ner model according to claim 1, wherein: In S4, the specific content of applying the K-means clustering algorithm to the vector representations of all relation phrases to reduce the relation types is: S41: Randomly select k samples as the initial clustering centers; S42: For each sample x i , find the closest centroid and assign the label c; S43: For each cluster, use the mean of all samples in the cluster as the new centroid.

Citation Information

Patent Citations

  • Knowledge graph construction method for integrating fragmented network security information

    CN118569372A

  • Large-model-assisted self-lifting multi-modal industrial equipment knowledge graph construction method

    CN119577159A

  • Knowledge graph construction method for ethylene oxide derivatives production process

    US20230169309A1