Standardized government affairs data construction method and device based on knowledge graph

By adopting a standardized data construction method based on knowledge graph in government affairs scenarios, using feature extraction models and information computing technology, the problem of unsatisfactory entity extraction results in government affairs scenarios is solved, and the precise extraction of complex text data and the construction of government affairs knowledge graphs are realized.

WO2025108197A1PCT designated stage expired Publication Date: 2025-05-30CETC BIGDATA RES INST CO LTD

Patent Information

Application Number
PCT/CN2024/132401
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is not ideal for entity extraction in government affairs scenarios, especially when dealing with phrase entities with complex text data and nested combinations of multiple words, the effect is not ideal.

Method used

The standardized government data construction method based on the knowledge graph is adopted, and the government affairs knowledge graph is preprocessed by pre-processing the corpus of government affairs scenarios, and the initial entity is identified using the feature extraction model, and the mutual information value and left and right entropy of adjacent entities are calculated, and a phrase entity is combined to finally construct a government affairs knowledge graph.

Benefits of technology

It realizes the precise extraction of phrase entities with nested combinations of multiple words in complex text data in government affairs scenarios, and builds a richer and more accurate government affairs knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132401_30052025_PF_FP_ABST
    Figure CN2024132401_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and provides a standardized government affairs data construction method and device based on a knowledge graph. The standardized government affairs data construction method based on a knowledge graph of the present invention comprises: on the basis of seed words of a government affairs scenario, using a feature extraction model to recognize a plurality of initial entities in the government affairs scenario; and then using a mutual information value between the adjacent initial entities to obtain a first phrase entity; on the basis of the mutual information value, obtaining a second phrase entity by calculating the left / right entropy, so that the range of the phrase entities is further expanded; and finally obtaining target entities. A phrase entity formed by embedding a plurality of words is extracted, and thus richer and more accurate entities are obtained to construct a knowledge graph in a government affairs scenario.
Need to check novelty before this filing date? Find Prior Art

Description

A method and device for constructing standardized government data based on knowledge graph Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for constructing standardized government data based on a knowledge graph. Background Art

[0002] Knowledge graphs formally describe real-world objects and their relationships. Knowledge graphs store knowledge in the form of entity, relationship, and entity triples, constructing knowledge networks with entities as nodes and relationships as edges. Currently, many well-known knowledge graph projects organize large amounts of data, extracting knowledge from it for organization and management, and providing users with high-quality intelligent services, such as understanding search semantics and providing more accurate search answers.

[0003] In government scenarios, knowledge graphs can play a very important role. The government sector involves a large number of entities and relationships, such as government agencies, departments, officials, regulations, policies, etc., and there are complex connections and dependencies between these entities and relationships. Knowledge graphs can model these entities and relationships in the form of graphs, thereby achieving effective management and application of government information. When constructing a knowledge graph for government scenarios, it is necessary to focus on the standardized representation and processing of data (knowledge) in the process of building a government service knowledge base. The knowledge graph construction process includes the acquisition, cleaning and processing of data in government scenarios, and the extraction of knowledge (entities, relationships, attributes), thereby achieving knowledge representation, ontology construction, knowledge storage, and knowledge fusion.

[0004] However, text data in government scenarios is usually more complex, containing a large amount of professional terms and context-related information, and the extraction of entities and relationships may face challenges. Existing entity extraction methods require word segmentation first, and then obtain entities. Most word segmentation methods use dictionary-based word segmentation methods. The dictionaries used are cross-domain general dictionaries, mostly common words, and lack proprietary words in government scenarios. Moreover, many text entities in government scenarios are nested combinations of multiple words. For nested combinations of multiple words, for example, for the term "house rental party", the entities extracted using conventional word segmentation methods will be divided into "house rental" and "party", and the effect of this method is not ideal. Summary of the Invention

[0005] The present invention provides a method and device for constructing standardized government data based on knowledge graphs, which are used to solve the defect that the entity extraction effect in the existing technology is not ideal, and to achieve the effect of accurately extracting entities in government scenarios.

[0006] The present invention provides a method for constructing standardized government data based on a knowledge graph, comprising:

[0007] Perform data preprocessing on the corpus of government affairs scenarios;

[0008] Based on the seed words of the government affairs scenario, the feature extraction model is used to identify the initial entities in the preprocessed corpus;

[0009] Determining mutual information values ​​between adjacent initial entities, and combining adjacent initial entities whose mutual information values ​​are greater than a first threshold into a first phrase entity;

[0010] Determining left-right entropy of initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and determining a second phrase entity based on initial entities whose left-right entropy is less than a third threshold; the second threshold is less than the first threshold;

[0011] Determine the initial entity, the first phrase entity, and the second phrase entity, whose mutual information values ​​with adjacent initial entities are less than or equal to the first threshold, as target entities;

[0012] Based on a corpus of government affairs scenarios and the target entities, extracting relationships between the target entities;

[0013] The relationships between the identified target entities are stored in a relational database to construct a government knowledge graph.

[0014] According to a method for constructing standardized government data based on a knowledge graph provided by the present invention, determining the mutual information value between adjacent initial entities includes:

[0015] Traversing the corpus of government affairs scenarios, and determining the number of occurrences of two adjacent initial entities and a phrase consisting of two adjacent initial entities in the corpus of government affairs scenarios;

[0016] Divide the number of occurrences of two adjacent initial entities by the total number of all words in the corpus to obtain the probability of occurrence of the two adjacent initial entity words respectively, and divide the number of occurrences of the phrase composed of the two adjacent initial entities in the corpus of government affairs scenarios by the total number of all words in the corpus to obtain the probability of occurrence of the phrase composed of the two adjacent initial entities;

[0017] Based on the occurrence probabilities of two adjacent initial entity words and the occurrence probabilities of phrases composed of two adjacent initial entities, the mutual information values ​​between adjacent initial entities are determined.

[0018] According to a method for constructing standardized government data based on a knowledge graph provided by the present invention, determining the left and right entropies of initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus includes:

[0019] Determining a word selection window size, wherein the word selection window size is used to limit the range of left and right neighboring words of the initial entity in the preprocessed corpus;

[0020] Based on the word selection window size, obtain neighboring words within a left and right range of the initial entity whose mutual information value of adjacent initial entities is greater than a second threshold in the corpus of the government affairs scenario;

[0021] Determine the number of times each neighboring word appears in all word selection windows, and obtain the frequency of each neighboring word in all left word selection windows and all right word selection windows respectively;

[0022] Based on the frequency of each neighboring word in each word selection window, the left and right entropy of the initial entity corresponding to each neighboring word in the preprocessed corpus is determined.

[0023] According to a method for constructing standardized government data based on a knowledge graph provided by the present invention, determining a second phrase entity based on an initial entity whose left and right entropies are less than a third threshold value includes:

[0024] Determine a left entropy value and a right entropy value in the left and right entropies of an initial entity whose left and right entropies are less than a third threshold value in the preprocessed corpus;

[0025] The adjacent word corresponding to the smaller one of the left entropy value and the right entropy value and the initial entity are combined into the second phrase entity.

[0026] According to a method for constructing standardized government data based on a knowledge graph provided by the present invention, a feature extraction model is used to identify initial entities in a preprocessed corpus, including:

[0027] Based on the seed words of the government scenario, a text feature template is constructed; the feature template includes the word vector representation, part-of-speech tag, seed word distance and context information;

[0028] Training the feature extraction model using the initial training corpus and the text feature template;

[0029] The corpus in the preprocessed corpus is input into the trained feature extraction model to obtain the features corresponding to the initial entity output by the feature extraction model, and the initial entity is determined.

[0030] According to a method for constructing standardized government data based on a knowledge graph provided by the present invention, the feature extraction model is a BiLSTM-CRF model, and the feature extraction model includes an input layer, an encoding layer, and a decoding layer;

[0031] The input layer is used to convert words or characters in the input corpus text into continuous vector representations;

[0032] The encoding layer is used to extract contextual features of the vector sequence of the input corpus text and generate an encoding vector sequence containing semantic information;

[0033] The decoding layer is used to use a conditional random field to label the encoding vector sequence output by the encoding layer, and predict the label of each labeled position to generate and output features corresponding to the initial entity.

[0034] According to a method for constructing standardized government data based on a knowledge graph provided by the present invention, the seed words of the government scenarios include words in a government scenario-specific dictionary and words manually annotated in the government scenarios.

[0035] The present invention also provides a device for constructing standardized government data based on a knowledge graph, comprising:

[0036] The first processing module is used to preprocess the data of the corpus of government affairs scenarios;

[0037] The second processing module is used to identify initial entities in the preprocessed corpus using a feature extraction model based on seed words in government scenarios;

[0038] a third processing module, configured to determine mutual information values ​​between adjacent initial entities, and combine adjacent initial entities having mutual information values ​​greater than a first threshold into a first phrase entity;

[0039] a fourth processing module, configured to determine the left-right entropy of initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and to determine a second phrase entity based on initial entities whose left-right entropy is less than a third threshold; the second threshold is less than the first threshold;

[0040] a fifth processing module, configured to determine an initial entity, the first phrase entity, and the second phrase entity, whose mutual information value with adjacent initial entities is less than or equal to the first threshold, as a target entity;

[0041] a sixth processing module, configured to extract relationships between the target entities based on a corpus of government affairs scenarios and the target entities;

[0042] The seventh processing module is used to store the relationships between the determined target entities in a relational database to construct a government knowledge graph.

[0043] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements any of the above-described methods for constructing standardized government data based on knowledge graphs.

[0044] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements any of the above-mentioned methods for constructing standardized government data based on knowledge graphs.

[0045] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for constructing standardized government data based on knowledge graphs.

[0046] The present invention provides a method and device for constructing standardized government data based on a knowledge graph. The method utilizes a feature extraction model based on seed words of the government scenario to identify multiple initial entities in the government scenario, and then utilizes the mutual information value between adjacent initial entities to obtain the first phrase entity. The second phrase entity is obtained by calculating the left and right entropies based on the mutual information value, further expanding the scope of the phrase entity and finally obtaining the target entity. The extraction of phrase entities of multiple nested word combinations is realized, thereby obtaining richer and more accurate entities to construct a knowledge graph in the government scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] FIG1 is a flow chart of a method for constructing standardized government data based on a knowledge graph provided by the present invention;

[0049] FIG2 is a schematic diagram of the structure of a device for constructing standardized government data based on a knowledge graph provided by the present invention;

[0050] FIG3 is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0051] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0052] The following describes a method and device for constructing standardized government data based on knowledge graphs of the present invention in conjunction with Figures 1 to 3.

[0053] As shown in FIG1 , a method for constructing standardized government data based on a knowledge graph according to an embodiment of the present invention mainly includes steps 110 , 120 , 130 , 140 , 150 , 160 and 170 .

[0054] Step 110: preprocess the data of the corpus of government affairs scenarios.

[0055] The corpus of government affairs refers to text datasets used in the fields of government, administrative agencies, public services, etc. These datasets usually contain texts related to policies and regulations, public services, and administrative offices.

[0056] Corpora for government affairs scenarios typically include texts of various policies and regulations, such as the Constitution, laws, administrative regulations, rules, and normative documents. Public services provided by government agencies cover all areas of social life, and corpora for government affairs scenarios typically include texts related to public services, such as bidding announcements, budget reports, financial statements, personnel files, and social security benefits. Corpora for government affairs scenarios may also include various administrative office-related texts, such as official documents, correspondence, reports, memoranda, meeting minutes, and work plans.

[0057] Data preprocessing for government affairs corpora can include text data cleaning. Government affairs data sources are typically fixed and may contain duplicate data. Therefore, during data cleaning, it is necessary to identify and remove duplicate data to prevent interference with subsequent analysis. Government affairs data sources may also contain some noise data, such as irrelevant information and incorrect annotations. During data cleaning, it is necessary to identify and remove this noise data to prevent it from affecting subsequent analysis.

[0058] Since data in government scenarios may come from different institutions or departments, their formats and naming methods may vary. After cleaning, the data can be normalized to ensure uniform format and naming methods.

[0059] In the case of missing data in government scenarios, such as when certain fields are not filled in or recorded, these missing data can be identified and processed, such as filling in default values, deleting records, etc.

[0060] On this basis, corpus text that is convenient for entity extraction can be obtained.

[0061] Step 120 : Based on the seed words of the government affairs scenario, the feature extraction model is used to identify the initial entities in the preprocessed corpus.

[0062] In government affairs scenarios, seed words can be words related to government, administration, public affairs, etc.

[0063] For example, seed words can be vocabulary corresponding to policy documents, laws and regulations (such as the Constitution and Criminal Law), policy measures, etc.

[0064] Seed words can also be words corresponding to public services, such as education, health care, social security, environmental protection, infrastructure construction, etc.

[0065] Seed words can also be words corresponding to government information systems, government activities and meetings, etc., which are not limited here.

[0066] The seed words for government scenarios can serve as a starting point for relevant information within government scenarios and can be used to search, classify, organize, and extract entities within government scenarios. Of course, the specific seed word list needs to be further customized and expanded based on actual application scenarios and needs.

[0067] It's understood that seed words for government scenarios can include terms from government-specific dictionaries and manually annotated terms in government scenarios. A government-specific dictionary is a tool used to store government-specific terms and nouns, facilitating tasks such as entity recognition and information extraction from government text. In other words, seed words for government scenarios can be obtained by acquiring existing government-specific terms or manually annotating them.

[0068] According to the task requirements of extracting entities, an appropriate feature extraction model can be selected, such as a machine learning-based model (such as conditional random field, maximum entropy model) or a deep learning-based model (such as recurrent neural network, Transformer), to extract entities from the corpus of government scenarios.

[0069] When performing feature extraction, appropriate features can be designed based on the government context and seed words. Based on this, the feature extraction model can be trained using labeled sample data. This sample data can be composed of manually labeled entity data and a corpus of government contexts, ensuring coverage of various entity types in government contexts.

[0070] It can be understood that the preprocessed corpus is subjected to entity recognition using a trained feature extraction model. The feature extraction model can classify each word according to the features obtained by feature extraction, determine whether it is an entity, and obtain each initial entity.

[0071] In some embodiments, based on seed words of government scenarios, a feature extraction model is used to identify initial entities in a preprocessed corpus, including the following process.

[0072] You can first build a text feature template based on the seed words in the government scenario. The feature template includes the word vector representation, part-of-speech tag, seed word distance, and context information.

[0073] On this basis, the feature extraction model is trained using the initial training corpus and text feature templates, and then the corpus in the preprocessed corpus is input into the trained feature extraction model to obtain the features corresponding to the initial entity output by the feature extraction model and determine the initial entity.

[0074] It is understandable that, based on the characteristics of government scenarios, such as policies and regulations, public services, etc., some key words can be determined as seed words through expert knowledge or domain analysis, and then a text feature template can be constructed by combining multiple features such as word vector representation, part-of-speech tags, seed word distance, and contextual information.

[0075] On this basis, the corpus and text feature templates are fed into a feature extraction model for training, aiming to learn how to use different features to extract entity information. By feeding the preprocessed corpus into the feature extraction model, the trained model can be used to extract the initial entities in the corpus and obtain the corresponding text features. Based on the features corresponding to the entities output by the feature extraction model, valid entities can be determined based on certain rules and thresholds, i.e., the initial entities.

[0076] In this implementation, using seed words to construct text feature templates can improve the recall and precision of the extraction algorithm. Seed words guide the algorithm to find relevant entities and relationships in the text, while text feature templates capture key information, helping the algorithm accurately identify and extract entities and relationships.

[0077] Knowledge in the government sector is constantly changing and expanding, so the knowledge graph needs to be constantly updated and expanded. In the subsequent process, the text feature template can support the dynamic update and expansion of the knowledge graph by adding and modifying seed words.

[0078] In some embodiments, the feature extraction model is a BiLSTM-CRF model. The BiLSTM-CRF model is a deep learning model commonly used in sequence labeling tasks, which combines the advantages of both BiLSTM and CRF models.

[0079] BiLSTM is a bidirectional recurrent neural network that processes the input sequence using forward LSTM and backward LSTM, respectively, to obtain contextual information and effectively handle long sequence dependencies. In sequence labeling tasks such as entity recognition in government scenarios, BiLSTM can combine word and context information to generate a hidden state representation corresponding to each word.

[0080] CRF is a conditional random field model that takes into account the dependencies between adjacent labels and can perform global optimal labeling of the entire sequence, thereby improving the accuracy of the model. In entity recognition tasks, CRF can use the hidden state representation generated by BiLSTM to accurately label entities in the sequence.

[0081] The feature extraction model consists of input layer, encoding layer and decoding layer.

[0082] The input layer converts the words or characters in the input text into continuous vector representations. The purpose of the input layer is to convert discrete text input into dense vector representations for subsequent processing. For example, you can use a pre-trained word embedding model (such as Word2Vec or GloVe) to obtain vector representations of words.

[0083] The encoding layer extracts contextual features from the input text vector sequence and generates an encoded vector sequence containing semantic information. The encoding layer uses a bidirectional long short-term memory (BiLSTM) network to model the input sequence. By simultaneously considering forward and backward contextual information, the BiLSTM effectively captures long-range dependencies in the text. The encoding layer extracts contextual features from the input sequence and generates an encoded vector sequence rich in semantic information.

[0084] The decoding layer is used to annotate the sequence of encoded vectors output by the encoding layer using a conditional random field, predict the label at each annotated position, and generate features corresponding to the output initial entity. The decoding layer considers the dependencies between adjacent labels and can annotate the entire sequence based on global probabilities, rather than simply classifying each position independently. The role of the decoding layer is to use the CRF model to predict the label at each position based on the feature sequence output by the encoding layer, and generate the final annotation sequence, thereby obtaining the initial entity.

[0085] In this embodiment, BiLSTM-CRF can automatically learn features and capture contextual information compared to traditional rule-based or feature engineering methods. Traditional rule-based or feature engineering methods require manual feature design, while the BiLSTM-CRF model can automatically learn feature representations of input sequences, including contextual information, word vector representations, etc. This makes the model more flexible and versatile. The BiLSTM-CRF model is a sequence-based model that can use contextual information for labeling. Using contextual information for labeling is very important for entity extraction tasks, because entities composed of nested words in government scenarios are usually associated with their context. Only by considering contextual information can entities be identified more accurately.

[0086] In addition, the BiLSTM-CRF model can uniformly process multiple entity types, including names of people, places, and organizations, making the model more versatile and generalizable.

[0087] Step 130 : determining mutual information values ​​between adjacent initial entities, and combining adjacent initial entities with mutual information values ​​greater than a first threshold into a first phrase entity.

[0088] It is understandable that after the initial entities are identified, the positions between the initial entities can be identified to further identify adjacent initial entities. For example, it can be identified that "housing rental" is adjacent to "party".

[0089] Mutual information is a metric used to measure the correlation between two events. Determining the mutual information between adjacent initial entities can include the following process.

[0090] The corpus of government affairs scenarios may be traversed first to determine the number of occurrences of two adjacent initial entities and a phrase composed of two adjacent initial entities in the corpus of government affairs scenarios.

[0091] The number of occurrences of two adjacent initial entities is divided by the total number of all words in the corpus to obtain the occurrence probability of the two adjacent initial entity words, and the number of occurrences of the phrase composed of two adjacent initial entities in the corpus of government scenarios is divided by the total number of all words in the corpus to obtain the occurrence probability of the phrase composed of two adjacent initial entities.

[0092] Based on the occurrence probabilities of two adjacent initial entity words and the occurrence probabilities of phrases composed of two adjacent initial entities, the mutual information values ​​between adjacent initial entities are determined.

[0093] For example, assuming that the preprocessed corpus contains two adjacent words A and B, the mutual information value MI(A,B) of words A and B can be expressed as: MI(A,B)=log2(P(A,B) / (P(A)*P(B)));

[0094] Among them, P(A,B) represents the probability of occurrence of a phrase consisting of word A and word B, and P(A) and P(B) represent the probability of occurrence of word A and word B separately.

[0095] It is understandable that the first threshold can be set according to actual conditions. On this basis, adjacent initial entities with mutual information values ​​greater than the first threshold can be combined into a first phrase entity.

[0096] In this embodiment, by calculating the mutual information value of adjacent initial entities, it is possible to determine whether there is correlation between two adjacent initial entities, and then form a first phrase entity, thereby constructing an entity composed of multiple nested words, and obtaining a richer and more accurate initial entity.

[0097] Step 140 , determining left-right entropy of initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and determining a second phrase entity based on initial entities whose left-right entropy is less than a third threshold.

[0098] It should be noted that the second threshold is smaller than the first threshold.

[0099] In this case, the initial entities whose mutual information values ​​of adjacent initial entities are greater than the second threshold include initial entities that can constitute the first phrase entity, and also include some initial entities that cannot constitute the first phrase entity.

[0100] In some embodiments, the left-right entropy of the initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus may be determined specifically in the following manner.

[0101] First, the word selection window size is determined. The word selection window size is used to limit the range of the left and right neighboring words of the initial entity in the preprocessed corpus.

[0102] For example, the window size can be set to 2, which means considering 2 neighboring words on the left and right of the target word.

[0103] On this basis, based on the word selection window size, the neighboring words within the left and right range of the initial entities whose mutual information values ​​are greater than the second threshold in the corpus of the government affairs scenario are obtained. For example, the neighboring words can be obtained by sliding based on a fixed window size.

[0104] Determine the number of times each neighboring word appears in all word selection windows, and obtain the frequency of each neighboring word in all left word selection windows and all right word selection windows respectively; based on the frequency of each neighboring word in each word selection window, determine the left and right entropy of the initial entity corresponding to each neighboring word in the preprocessed corpus.

[0105] To determine the number of times the neighboring words of each initial entity appear in all word selection windows, and obtain the frequency of each neighboring word appearing in all left word selection windows and all right word selection windows, the following process can be used.

[0106] Traverse the entire corpus. For each word selection window, count the number of occurrences of each neighboring word within the window and add the counts to the corresponding neighboring word counter. Based on this, count the frequencies of all neighboring words in the left word selection window. This is done by dividing the count of each neighboring word by the total number of words in the left word selection window. Similarly, count the frequencies of all neighboring words in the right word selection window. This is done by dividing the count of each neighboring word by the total number of words in the right word selection window.

[0107] For each neighboring word, calculate its frequency distribution across all left-side word selection windows. This frequency distribution is calculated by dividing the frequency of the neighboring word in each left-side word selection window by the total number of words in that window. Information entropy can be used to measure the uncertainty of the neighboring words in the left-side word selection window. Based on the frequency distribution, calculate the probability of each neighboring word and then calculate its left entropy value. Similarly, repeat the above process for each neighboring word, calculating its frequency distribution and right entropy value across all right-side word selection windows.

[0108] After obtaining the left entropy value and the right entropy value, the average of the two can be calculated as the left and right entropies of the initial entity corresponding to each neighboring word in the preprocessed corpus, or the smaller of the two values ​​can be taken as the left and right entropies of the initial entity corresponding to each neighboring word in the preprocessed corpus. There is no restriction on the specific calculation method of the left and right entropies here.

[0109] In this case, the second phrase entity may be determined based on the initial entity whose left and right entropies are less than the third threshold.

[0110] In this embodiment, by calculating the left and right entropies, it is possible to further identify second phrase entities with strong boundary sense after forming phrases with some initial entities, and further identify possible entities formed by multiple nested words.

[0111] In some embodiments, based on an initial entity whose left and right entropies are less than a third threshold, a second phrase entity is determined, specifically including: determining the left entropy value and the right entropy value in the left and right entropies of the initial entity whose left and right entropies are less than the third threshold in the preprocessed corpus; and combining the adjacent words corresponding to the smaller one of the left entropy value and the right entropy value with the initial entity to form a second phrase entity.

[0112] In this embodiment, adjacent words with strong boundaries can be identified and then combined with the initial entity to form a second phrase entity, thereby improving the accuracy of entity extraction.

[0113] Step 150 : determining the initial entity, the first phrase entity, and the second phrase entity whose mutual information values ​​with adjacent initial entities are less than or equal to a first threshold as target entities.

[0114] It is understandable that the second phrase entity may also include phrase entities that are repeated with the first phrase entity. When identifying the target entity, the repeated phrase entities in the second phrase entity may be screened out to obtain the target entity.

[0115] Step 160 : Based on the corpus of government affairs scenarios and target entities, extract the relationships between target entities.

[0116] After identifying the target entities in the government affairs scenario corpus, the relationships between the target entities can be extracted based on the government affairs scenario corpus and the target entities.

[0117] In some embodiments, the relationship between target entities can be found based on predefined rules or patterns. For example, in a sentence, if there are specific keywords or phrases between two target entities, the corresponding relationship can be defined based on these keywords or phrases.

[0118] In some embodiments, machine learning or deep learning techniques can be used to build a relationship classification model, taking the target entity and its context as input to predict the relationship category between them. In this case, a training dataset with labeled relationship categories is required, and a supervised learning algorithm is used for model training.

[0119] It is understandable that the appropriate method can be selected according to the specific task and data characteristics. In practical applications, a combination of multiple technologies and methods can also be used to improve the accuracy and effect of relationship extraction.

[0120] Step 170: Store the relationships between the determined target entities in a relational database to construct a government knowledge graph.

[0121] The relational database is used to store the relationships between target entities. The relational database also contains entities and entity relationships extracted using other entities, which are not limited here. It is understood that after extracting the relationships between target entities, the determined relationships between target entities are stored in the relational database to construct a government knowledge graph.

[0122] According to an embodiment of the present invention, a method for constructing standardized government data based on a knowledge graph is provided. A feature extraction model is used based on seed words of government scenarios to identify multiple initial entities in government scenarios. The mutual information value between adjacent initial entities is then used to obtain a first phrase entity. The second phrase entity is obtained by calculating the left and right entropies based on the mutual information value, further expanding the range of the phrase entity and finally obtaining the target entity. This realizes the extraction of phrase entities of nested combinations of multiple words, thereby obtaining richer and more accurate entities to construct a knowledge graph in government scenarios.

[0123] The following describes a standardized government data construction device based on a knowledge graph provided by the present invention. The standardized government data construction device based on a knowledge graph described below and the standardized government data construction method based on a knowledge graph described above can refer to each other.

[0124] As shown in Figure 2, a standardized government data construction device based on a knowledge graph according to an embodiment of the present invention mainly includes a first processing module 210, a second processing module 220, a third processing module 230, a fourth processing module 240, a fifth processing module 250, a sixth processing module 260 and a seventh processing module 270.

[0125] The first processing module 210 is used to perform data preprocessing on the corpus of government affairs scenarios;

[0126] The second processing module 220 is used to identify initial entities in the pre-processed corpus using a feature extraction model based on the seed words of the government affairs scenario;

[0127] The third processing module 230 is used to determine the mutual information values ​​between adjacent initial entities, and combine adjacent initial entities with mutual information values ​​greater than a first threshold into a first phrase entity;

[0128] The fourth processing module 240 is configured to determine the left-right entropy of the initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and determine the second phrase entity based on the initial entities whose left-right entropy is less than a third threshold; the second threshold is less than the first threshold;

[0129] The fifth processing module 250 is configured to determine the initial entity, the first phrase entity, and the second phrase entity whose mutual information value with the adjacent initial entity is less than or equal to the first threshold as a target entity;

[0130] The sixth processing module 260 is used to extract the relationship between target entities based on the corpus of government affairs scenarios and target entities;

[0131] The seventh processing module 270 is used to store the relationships between the determined target entities in a relational database to construct a government knowledge graph.

[0132] According to an embodiment of the present invention, a standardized government data construction device based on a knowledge graph is provided. A feature extraction model is used based on seed words of the government scenario to identify multiple initial entities in the government scenario, and then the mutual information value between adjacent initial entities is used to obtain the first phrase entity. The second phrase entity is obtained by calculating the left and right entropies based on the mutual information value, further expanding the range of the phrase entity, and finally obtaining the target entity, thereby realizing the extraction of phrase entities of multiple nested word combinations, and thus obtaining richer and more accurate entities to construct the knowledge graph in the government scenario.

[0133] In some embodiments, the third processing module 230 is also used to traverse the corpus of government affairs scenarios, and determine the number of occurrences of two adjacent initial entities and the phrase composed of two adjacent initial entities in the corpus of government affairs scenarios; divide the number of occurrences of the two adjacent initial entities by the total number of all words in the corpus to obtain the occurrence probability of the two adjacent initial entity words, and divide the number of occurrences of the phrase composed of two adjacent initial entities in the corpus of government affairs scenarios by the total number of all words in the corpus to obtain the occurrence probability of the phrase composed of two adjacent initial entities; based on the occurrence probability of the two adjacent initial entity words and the occurrence probability of the phrase composed of two adjacent initial entities, determine the mutual information value between adjacent initial entities.

[0134] In some embodiments, the fourth processing module 240 is also used to determine the size of the word selection window, which is used to limit the range of left and right neighboring words of the initial entity in the preprocessed corpus; based on the word selection window size, the neighboring words within the left and right range of the initial entity whose mutual information value of adjacent initial entities is greater than the second threshold are obtained in the corpus of the government scenario; the number of times each neighboring word appears in all word selection windows is determined, and the frequency of each neighboring word appearing in all left word selection windows and all right word selection windows is obtained respectively; based on the frequency of each neighboring word appearing in each word selection window, the left and right entropy of the initial entity corresponding to each neighboring word in the preprocessed corpus is determined.

[0135] In some embodiments, the fourth processing module 240 is also used to determine the left entropy value and the right entropy value in the left and right entropies of the initial entity whose left and right entropies are less than the third threshold in the preprocessed corpus; and the adjacent words corresponding to the smaller one of the left entropy value and the right entropy value are combined with the initial entity to form a second phrase entity.

[0136] In some embodiments, the second processing module 220 is also used to construct a text feature template based on seed words of government scenarios; the feature template includes word vector representation, part-of-speech tags, seed word distance and context information of the words; the feature extraction model is trained using the initial training corpus and the text feature template; the corpus in the preprocessed corpus is input into the trained feature extraction model to obtain the features corresponding to the initial entity output by the feature extraction model, and determine the initial entity.

[0137] In some embodiments, the feature extraction model is a BiLSTM-CRF model, and the feature extraction model includes an input layer, an encoding layer, and a decoding layer; the input layer is used to convert words or characters in the input corpus text into continuous vector representations; the encoding layer is used to extract contextual features of the vector sequence of the input corpus text, and generate an encoding vector sequence containing semantic information; the decoding layer is used to use a conditional random field to label the encoding vector sequence output by the encoding layer, and predict the label of each labeled position to generate features corresponding to the output initial entity.

[0138] In some embodiments, the seed words of the government affairs scenario include words in the government affairs scenario-specific dictionary and words manually marked in the government affairs scenario.

[0139] Figure 3 illustrates a schematic diagram of the physical structure of an electronic device. As shown in Figure 3, the electronic device may include: a processor (processor) 310, a communication interface (Communications Interface) 320, a memory (memory) 330 and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute a standardized government data construction method based on a knowledge graph, which includes: preprocessing data on the corpus of government scenarios; based on the seed words of the government scenarios, using a feature extraction model to identify the initial entities in the preprocessed corpus; determining the mutual information value between adjacent initial entities, and combining adjacent initial entities with mutual information values ​​greater than a first threshold into a first phrase entity; determining the left and right entropies of initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and determining the second phrase entity based on the initial entities whose left and right entropy are less than a third threshold; the second threshold is less than the first threshold; determining the initial entity, the first phrase entity and the second phrase entity whose mutual information values ​​with the adjacent initial entities are less than or equal to the first threshold as target entities; extracting the relationship between the target entities based on the corpus of government scenarios and the target entities; storing the determined relationship between the target entities in a relational database to construct a government knowledge graph.

[0140] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0141] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the standardized government data construction method based on the knowledge graph provided by the above methods, the method including: performing data preprocessing on the corpus of government affairs scenarios; based on the seed words of the government affairs scenarios, using a feature extraction model to identify the initial entities in the preprocessed corpus; determining the mutual information value between adjacent initial entities, and combining the adjacent initial entities with a mutual information value greater than a first threshold into a first phrase entity; determining the left and right entropies of the initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and determining the second phrase entity based on the initial entities whose left and right entropy is less than a third threshold; the second threshold is less than the first threshold; determining the initial entity, the first phrase entity and the second phrase entity whose mutual information value with the adjacent initial entity is less than or equal to the first threshold as the target entity; extracting the relationship between the target entities based on the corpus of government affairs scenarios and the target entity; storing the determined relationship between the target entities in a relational database to construct a government affairs knowledge graph.

[0142] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the standardized government data construction method based on the knowledge graph provided by the above-mentioned methods, the method comprising: performing data preprocessing on the corpus of government affairs scenarios; identifying initial entities in the preprocessed corpus using a feature extraction model based on seed words of government affairs scenarios; determining the mutual information value between adjacent initial entities, and combining adjacent initial entities with mutual information values ​​greater than a first threshold into a first phrase entity; determining the left and right entropies of initial entities in the preprocessed corpus whose mutual information values ​​of adjacent initial entities are greater than a second threshold, and determining a second phrase entity based on the initial entities with left and right entropy less than a third threshold; the second threshold is less than the first threshold; determining the initial entity, the first phrase entity and the second phrase entity whose mutual information values ​​with the adjacent initial entities are less than or equal to the first threshold as target entities; extracting the relationship between the target entities based on the corpus of government affairs scenarios and the target entities; storing the determined relationship between the target entities in a relational database to construct a government affairs knowledge graph.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for constructing standardized government data based on knowledge graph, characterized in that: include: Perform data preprocessing on the corpus of government affairs scenarios; Based on the seed words of the government affairs scenario, the feature extraction model is used to identify the initial entities in the preprocessed corpus; Determine mutual information values ​​between adjacent initial entities, and combine adjacent initial entities whose mutual information values ​​are greater than a first threshold into a first phrase entity; Determine the left-right entropy of the initial entities whose mutual information values ​​of adjacent initial entities are greater than the second threshold in the preprocessed corpus, and determine the second phrase entity based on the initial entities whose left-right entropy is less than the third threshold; The second threshold is less than the first threshold; Determine an initial entity, the first phrase entity, and the second phrase entity, whose mutual information value with adjacent initial entities is less than or equal to the first threshold, as a target entity; Based on the corpus of government affairs scenarios and the target entities, extracting the relationship between the target entities; The relationships between the identified target entities are stored in a relational database to construct a government knowledge graph.

2. The method for constructing standardized government data based on knowledge graph according to claim 1 is characterized in that: The determining of the mutual information value between adjacent initial entities comprises: Traversing the corpus of government affairs scenarios, and determining the number of occurrences of two adjacent initial entities and a phrase composed of two adjacent initial entities in the corpus of government affairs scenarios; The number of occurrences of two adjacent initial entities is divided by the total number of all words in the corpus to obtain the occurrence probabilities of the two adjacent initial entity words, and the number of occurrences of a phrase composed of two adjacent initial entities in the corpus of government affairs scenarios is divided by the total number of all words in the corpus to obtain the occurrence probability of a phrase composed of two adjacent initial entities; Based on the occurrence probabilities of two adjacent initial entity words and the occurrence probabilities of phrases composed of two adjacent initial entities, the mutual information values ​​between adjacent initial entities are determined.

3. The method for constructing standardized government data based on knowledge graph according to claim 1 is characterized in that: The determining of the left-right entropy of the initial entities whose mutual information values ​​of adjacent initial entities are greater than the second threshold in the preprocessed corpus includes: Determine a word selection window size, wherein the word selection window size is used to limit the range of left and right neighboring words of the initial entity in the preprocessed corpus; Based on the word selection window size, obtain neighboring words within a left and right range of the initial entity whose mutual information value of adjacent initial entities is greater than a second threshold in the corpus of the government affairs scenario; Determine the number of times each neighboring word appears in all word selection windows, and obtain the frequency of each neighboring word in all left word selection windows and all right word selection windows respectively; Based on the frequency of occurrence of each neighboring word in each word selection window, the left-right entropy of the initial entity corresponding to each neighboring word in the preprocessed corpus is determined.

4. The method for constructing standardized government data based on knowledge graph according to claim 3 is characterized in that: The determining of the second phrase entity based on the initial entity whose left-right entropy is less than a third threshold comprises: Determine a left entropy value and a right entropy value in the left and right entropies of an initial entity whose left and right entropies are less than a third threshold in the preprocessed corpus; The neighboring word corresponding to the smaller one of the left entropy value and the right entropy value and the initial entity are combined into the second phrase entity.

5. The method for constructing standardized government data based on knowledge graph according to claim 1 is characterized in that: The seed words based on the government affairs scenario use a feature extraction model to identify initial entities in the preprocessed corpus, including: Based on the seed words of the government affairs scenario, a text feature template is constructed; the feature template includes a word vector representation, a part-of-speech tag, a seed word distance, and context information; Training the feature extraction model using the initial training corpus and the text feature template; The corpus in the preprocessed corpus is input into the trained feature extraction model to obtain the features corresponding to the initial entity output by the feature extraction model, and determine the initial entity.

6. The method for constructing standardized government data based on knowledge graph according to claim 5 is characterized in that: The feature extraction model is a BiLSTM-CRF model, and the feature extraction model includes an input layer, a coding layer and a decoding layer; The input layer is used to convert words or characters in the input corpus text into continuous vector representations; The encoding layer is used to extract contextual features of the vector sequence of the input corpus text and generate an encoding vector sequence containing semantic information; The decoding layer is used to use a conditional random field to annotate the encoding vector sequence output by the encoding layer, and predict the label of each annotated position to generate and output features corresponding to the initial entity.

7. The method for constructing standardized government data based on knowledge graph according to claim 1 is characterized in that: The seed words of government affairs scenarios include words in the government affairs scenario-specific dictionary and words manually annotated in the government affairs scenarios.

8. A device for constructing standardized government data based on knowledge graph, characterized in that: The first processing module is used to perform data preprocessing on the corpus of government affairs scenarios; The second processing module is used to identify the initial entities in the preprocessed corpus using a feature extraction model based on the seed words of the government affairs scenario; A third processing module, configured to determine mutual information values ​​between adjacent initial entities, and combine adjacent initial entities whose mutual information values ​​are greater than a first threshold into a first phrase entity; A fourth processing module is used to determine the left-right entropy of the initial entities whose mutual information values ​​of adjacent initial entities are greater than a second threshold in the preprocessed corpus, and determine the second phrase entity based on the initial entities whose left-right entropy is less than a third threshold; the second threshold is less than the first threshold; A fifth processing module, configured to determine the initial entity, the first phrase entity, and the second phrase entity, whose mutual information value with adjacent initial entities is less than or equal to the first threshold, as target entities; A sixth processing module, for extracting the relationship between the target entities based on the corpus of government affairs scenarios and the target entities; The seventh processing module is used to store the relationship between the determined target entities in a relational database to construct a government knowledge graph.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the method for constructing standardized government data based on the knowledge graph as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for constructing standardized government data based on knowledge graph as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Power dispatching field named entity identification method and system based on BiLSTM-CRF model

    CN111553158A

  • Knowledge extraction method and device, electronic equipment and storage medium

    CN111639498A

  • Standardized government affair data construction method and device based on knowledge graph

    CN117251685A

  • Training method and training device for language representation model

    WO2023185082A1

Cited By

  • Government affair information recommendation method and device based on knowledge graph and multi-mode fusion

    CN120353924A

  • Dynamic knowledge graph driven cross-department government affair data collaboration method and system

    CN120912158A

  • Government affair system optimization method and device, nonvolatile storage medium and electronic equipment

    CN121119310A

  • Electronic bidding document generation method and system based on artificial intelligence

    CN121145813A

  • Data alignment method, device and equipment based on ontology modeling technology and medium

    CN121166934A