A Method for Identifying Core Governance Objective-Tool Combinations Based on Citation Association

Through the core governance objective-tool combination recognition method based on reference association, combined with semantic drive and self-spheric unit algorithm, the problem of difficult to analyze the relationship between governance objectives and tools in the existing technology is solved, and an in-depth understanding and identification of governance objective-tool combinations and evolutionary paths is achieved, providing more scientific governance decision support.

CN119848238BActive Publication Date: 2025-06-27NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510338485.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively analyze the complex relationship between governance goals and tools, track the evolution path of governance portfolios, and evaluate governance effects, and lacks a comprehensive analysis method for the evolution of governance goals-tool combination patterns and dynamic changes in core portfolios.

Method used

The core governance objective-tool combination recognition method based on citation association is adopted, and the model recognition of governance objective-tool combination is revealed through semantic-driven target and tool recognition, combined with citation network analysis and self-spheric unit algorithm (ESU), revealing the core objective-tool combination and its evolutionary trajectory.

Benefits of technology

It has achieved in-depth disclosure of the interactive mode between governance goals and tools, improved the accuracy and efficiency of goal and tool identification, provided scientific data support and decision-making reference, and provided stronger support for governance optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848238B_ABST
    Figure CN119848238B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying core governance objective - tool combinations based on citation association, comprising the following steps: (1) Selecting a target domain, collecting governance documents to construct a data pool and performing preprocessing; (2) Annotating and converting file data; (3) Constructing a W2NER model; (4) Training the W2NER model; (5) Extracting the W2NER model; (6) Constructing a governance objective - tool citation network; (7) Identifying core governance objective - tools; The present invention can reveal core governance objective - tool combinations and their interaction logic, contribute to comprehensively understanding the relationships and functions among governance document elements, provide scientific data support and decision-making basis for governance effect evaluation and optimization, and has important theoretical significance and practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data representation and computer systems, and particularly relates to a method for identifying core governance objective-tool combinations based on citation association. Background Art

[0002] With the continuous improvement of governance capabilities, the complexity and diversity of the governance system have become increasingly apparent. A single governance tool is difficult to cope with complex governance objectives. How to effectively analyze the complex relationship between governance objectives and governance tools, track the evolution path of governance combinations, and evaluate governance effects has become an important topic in the field of governance research. Governance objectives and tools are semantic structures extracted from texts, which can reflect the main objectives, implementation plans, and expected results of governance. The development and evolution process of governance often manifests as the dynamic combination of governance objectives and governance tools. This governance combination design is not only affected by the content of the governance documents themselves but also closely related to the citation and reference relationships between the documents. Therefore, constructing a network structure that can reveal the interaction pattern between governance objectives and tools has important value for improving the scientificity and systematicness of governance research.

[0003] Currently, there are the following deficiencies in the identification of governance objectives and tools: (1) The identification method relies on manual annotation or shallow semantic features, making it difficult to capture the complex associations between governance objectives and tools; (2) The citation network construction technology is simple and fails to reveal the deep interaction pattern between documents; (3) There is a lack of a comprehensive analysis method for the evolution of governance objective-tool combination patterns and the dynamic changes of core combinations. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to provide a method for identifying core governance objective-tool combinations based on citation association. Through semantic-driven objective and tool identification, combined with citation network analysis and motif identification of governance objective-tool combinations based on the self-sphere unit algorithm, the core objective-tool combinations and their evolution trajectories are revealed, providing scientific data support and decision-making references for governance decisions.

[0005] Technical Solution: A method for identifying core governance objective-tool combinations based on citation association according to the present invention includes the following steps:

[0006] (1) Select the target field, collect governance documents to construct a data pool and perform preprocessing;

[0007] (2) Sample data annotation and conversion: Collect sample data, and according to the cue words of objectives and tools, use Doccano software to annotate discontinuous governance objectives, tools, and entity relationships, and divide the training, validation, and test sets;

[0008] (3) Construct the W2NER model: Combine the bidirectional transformer model BERT, bidirectional long short-term memory LSTM, convolutional layer, and bi-affine Biaffine mechanism to construct a model for sequence annotation and relation extraction tasks, extract the semantic features of the text, and extract governance objectives, tool entities, and relationships.

[0009] (4) Train the W2NER model: A training model for joint extraction of text entities and relationships based on adversarial training and dynamic learning rate adjustment.

[0010] (5) Extract the W2NER model: Extract governance objectives, tool entities, and relationships between entities in the text using a deep learning-based method for joint extraction of text entities and relationships.

[0011] (6) Construct a governance objective-tool citation network: Extract citation and cited relationships between texts based on the book title marks and citation prompt words in the text content; identify, clean, and integrate governance objectives and tools in the text through the W2NER model, and construct a governance objective-tool citation network based on this information.

[0012] (7) Identify core governance objective-tools: Use the self-sphere unit algorithm ESU to perform a directed motif analysis on the governance objective-tool citation network, recursively identify the induced subgraph in the network, that is, the multiple relationships and interaction patterns between the target node and the tool node, and identify the core governance objective-tool combination based on this and analyze its evolution process.

[0013] Further, in step (1), the preprocessing is: perform a cleaning operation on the data.

[0014] Further, step (1) includes the following steps:

[0015] (11) Construct a keyword retrieval strategy for the target domain, and retrieve target file data with keywords in the title in the database.

[0016] (12) By reading the file data, identify and summarize the relevant keyword list for the target domain.

[0017] (13) Combine and match the file issue numbers and titles in the data pool to remove duplicate files, and only retain one of them.

[0018] (14) Automatically parse the text genre structure of the text based on natural language processing technology, and identify basic elements such as the file title, publishing agency, document type, and writing and effective dates.

[0019] Further, step (2) includes the following steps:

[0020] (21) Label the governance objectives. Based on the governance objective entity type and the verb phrase “realize / promote”, use the Doccano software to point the relationship pointer from one end of the discontinuous entity to the other end. If it is a continuous entity, directly label it as the entity type of governance objective.

[0021] (22) Labeling governance tools. Based on the entity type of governance tools and common verbs and nouns, including the semantic structure of "setting standards and establishing systems", Doccano software is used to use the relationship pointer from one end of the discontinuous entity to the other end. If it is a continuous entity, it is directly labeled as the entity type of governance tool;

[0022] (23) Mark the relationship between governance objectives and tools, combine the relationship between adopting governance tools to “achieve” governance objectives, and mark the relationship between governance objectives and tool models;

[0023] (24) Read the annotated data and export the annotated data set as a json file, which contains the annotated entity set and the annotated entity relationship set;

[0024] (25) Preprocessing of entity-relationship data: defining the target data format, removing isolated entities that are not involved in any relationship, and extracting the text index range of the starting entity and the target entity to generate relationship data;

[0025] (26) Dataset division: divide the data into a training set for model training, a validation set for model parameter tuning, and a test set for model performance testing.

[0026] Furthermore, step (3) includes the following steps:

[0027] (31) The model receives preprocessed data, including sentence word ID sequence, subword to word mapping matrix, two-dimensional mask matrix, entity distance matrix and sentence length, and verifies and masks the input data to make the data format consistent and mask invalid positions;

[0028] (32) Generate contextual representations based on multi-layer feature extraction. Use the pre-trained BERT model to contextually encode the input sequence and extract word-level or sub-word-level embeddings. Use the pre-trained BERT model to contextually encode the input sequence and extract word-level or sub-word-level embeddings. Input the word-level embeddings into a bidirectional LSTM to capture the global context information of the sentence.

[0029] (33) Conditional normalization processing, using the conditional layer normalization module to enhance the features of LSTM output;

[0030] (34) Integrate distance and relationship features, map the distance matrix between entities into a high-dimensional embedding representation, generate embeddings of relationship types according to the relationship matrix, and concatenate the distance embeddings, relationship embeddings, and context representations, i.e., the outputs of the LSTM processed by CLN, to form the feature input;

[0031] (35) Use a multi-dilation rate convolutional layer to perform local modeling on the feature input, and take the output of the convolutional layer as the feature input;

[0032] (36) Perform bi-affine interaction modeling, project the entity representations into a specific space, calculate the interaction scores of entity pairs, and predict the relationship categories between entities according to the interaction scores.

[0033] Further, step (4) includes the following steps:

[0034] (41) Data loading and preprocessing, load the training set, validation set, and test set, the data format includes text, entity annotations, and relationship annotations, and check and count the distribution of the loaded data;

[0035] (42) Model initialization, construct a joint extraction model according to the configuration file, the model structure includes using BERT for context feature extraction, bidirectional LSTM to further capture the global context, convolutional layers for local feature modeling, and bi-affine mechanisms for joint prediction of entities and relationships; set the loss function and group the parameters, set the learning rate scheduler to dynamically adjust the learning rate;

[0036] (43) Model training, set the model to training mode, perform forward propagation, calculate the loss, adversarial training, gradient clipping, and parameter update, and learning rate update steps for each batch of data; after each training epoch, calculate the average loss and record the prediction accuracy, recall rate, and F1 score;

[0037] (44) Model validation, set the model to evaluation mode, verify batch by batch, perform forward propagation and prediction, and count the performance, save the current model parameters whose F1 score on the validation set is higher than the historical best;

[0038] (45) Early stopping mechanism, when the performance on the validation set fails to improve for several consecutive rounds, record the number of stagnations; when the number of stagnations reaches the set threshold, end the training early to avoid overfitting;

[0039] (46) Model saving and loading, save the model parameter file at the specified path for subsequent loading or inference; load the saved model parameters to support continued training or direct evaluation.

[0040] Further, step (5) includes the following steps:

[0041] (51) Text tokenization: Perform character-level tokenization on the input text, replace consecutive spaces with a single space, and remove redundant characters.

[0042] (52) Data loading and construction: Use the tokenized text to construct a prediction dataset, including the text tokenization results and auxiliary features. After encapsulating the text into the model input format, load the data in batches.

[0043] (53) Model loading and initialization: Load the joint extraction model and load the trained model parameters.

[0044] (54) Model inference: Input the batch-loaded data into the model, perform forward propagation batch by batch, calculate the class scores at each position, and classify and output the entity class, i.e., the target or tool, and the entity relationship class.

[0045] (55) Decoding and result generation: Decode the indexes and classes output by the model, map the decoded entity and relationship results into a readable format, and finally generate a JSON file of the extraction results.

[0046] (56) Result storage: Output the extracted governance objectives, tools, and relationships, and store them in a local file according to the file ID and governance combination; where the governance combination includes: governance objective, implementation, and governance tool.

[0047] Further, step (6) includes the following steps:

[0048] (61) Identify identifiers: Identify the angle bracket identifiers embedded in the text content based on the KMP string matching algorithm.

[0049] (62) Complete the file name: Obtain the file publishing agency, and combine the publishing agency and the specific file within the angle brackets to construct the complete file name.

[0050] (63) Supplement the citation relationship: Identify the prompt words "according to", "implement in light of", "in accordance with", "provisions", "in combination with", "follow" accompanied before and after the angle brackets, and further identify and supplement the citation relationship.

[0051] (64) Citation relationship storage: Store the file citation relationship, i.e., the citing file ID and the cited file ID, according to the files cited in this article and the logic of the files cited in this article.

[0052] (65) Node construction and attribute mapping: Each governance objective or tool is defined as a node in the citation network. The attributes of the node include: node name, type, i.e., governance objective or tool, source text ID, and frequency of occurrence, etc.; the size of the node is proportional to the frequency of occurrence of the governance objective or tool in all texts, and the nodes with higher frequencies will be displayed as larger nodes in the network.

[0053] (66)Governance objective - tool combination reference network. Combine the governance objectives and tools in the reference network, and construct a "governance objective - tool" combination reference network according to the existing governance objective - tool combination patterns; each text data will contain multiple governance objective and tool combination patterns, representing the many - to - many relationship between governance objectives and tools.

[0054] Further, step (7) includes the following steps:

[0055] (71)Use the ESU algorithm for recursive mining: Apply the ESU (Ego - Sphere Unit) algorithm to recursively mine the sub - graphs in the governance objective - tool reference network. Each recursion will expand the SUB set, where the SUB set contains the nodes directly adjacent to the current node, and the EXT set contains the nodes adjacent to at least one node in the SUB set and with a larger numerical label. The ESU algorithm identifies all eligible induced sub - graphs through this recursive process;

[0056] (72)Motif analysis and sub - graph classification: Conduct motif analysis on each branch of the ESU tree, and identify the frequently occurring sub - graph structures in the network through a recursive method; during this process, the ESU algorithm performs an isomorphism test on each sub - graph through McKay's nauty algorithm to ensure the classification of sub - graphs with different structures and sort them according to their frequencies and concentrations;

[0057] (73)Identify the core governance objective - tool combinations: Identify the frequently occurring sub - graphs in the network, i.e., the core governance objective - tool combinations, through the ESU algorithm, and calculate the frequency F R (G′) of each sub - graph in the random network; at the same time, calculate the actual frequency F G (G′) of this sub - graph in the target network; use the P - value for statistical hypothesis testing to evaluate whether a sub - graph significantly appears in the target network; at the same time, set the N value to make the evaluation of sub - graph frequencies more stable and accurate by generating multiple random network samples; the formula is as follows:

[0058] Where, N is the number of random network samples, is the P - value, given with a probability of (as its null - hypothesis), represents the frequency of in the random network; when the P value is less than the threshold of 0.05, this sub - graph can be called a significant pattern.

[0059] (74) Core governance objective - tool motif visualization. Based on the identification of core governance objective - tool combinations, the motif network structure is displayed using visualization techniques. The governance objectives and tools are used as network nodes, and the size and color of the nodes represent the frequency and type of the nodes. The thickness of the edges reflects the strength of the citation relationship between the governance objective and the tool.

[0060] (75) Analyze the evolution process of the citation network based on the core governance objective - tool combination motif at different times or during the implementation stages of the documents. By tracking the trend of the core motif over time, identify the alternation and combination pattern changes between the governance objectives and tools.

[0061] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the methods for identifying core governance objective - tool combinations based on citation associations.

[0062] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages. The present invention combines the unified named entity recognition model (W2NER) as word - word classification and the ego - sphere unit (ESU) algorithm, which can deeply explore the interaction relationship between governance objectives and tools, and overcome the problem that the prior art does not fully consider the interaction between elements. Through the semantic - driven citation network and efficient motif analysis method, it can more accurately identify the objectives and tools in the text, improving the accuracy and efficiency of objective and tool recognition. Through the construction of citation relationships and motif analysis, the present invention can reveal the interaction patterns between governance objectives and tools at multiple levels, identify the induced sub - graphs in the network in a recursive manner, and analyze the multiple relationships and interaction patterns between the target nodes and tool nodes, which helps to deeply understand the complex associations between governance objectives and tools and provides more powerful support for governance optimization. Based on the perspectives of semantic structure and citation network evolution, the present invention extracts the combination patterns of governance objectives and tools, identifies the key nodes and restrictive relationships in the governance objective - tool network. Further, by analyzing the dynamic evolution path of the governance document citation network, it reveals the interaction and combination patterns between governance objectives and tools at different stages. Through the construction of a semantic - based citation network and efficient motif analysis, the present invention can reveal the core objective - tool combinations and their interaction logics in the governance system, comprehensively understand the relationships and functions between the elements of governance documents, and provide scientific data support and decision - making basis for governance effect evaluation and optimization, having important theoretical significance and practical value. Description of the Drawings

[0063] Figure 1 is a flowchart of the present invention;

[0064] Figure 2It is the two-dimensional matrix of word-word relationships of the present invention;

[0065] Figure 3 It is the two-dimensional matrix of THW and NHW relationships constructed from the relative position information of grid characters and annotation data entities of the present invention;

[0066] Figure 4 It is the flow chart of the method for extracting governance objectives and tool recognition based on semantic representation based on the W2NER model of the present invention;

[0067] Figure 5 It is the pseudocode of the ESU algorithm of the present invention;

[0068] Figure 6 It is the ESU subgraph search algorithm of the present invention; among them, Figure 6 in (a) is the target graph with a size of 5; Figure 6 in (b) is the ESU tree with a depth of k. Detailed implementation manners

[0069] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0070] As Figure 1 shown, the embodiment of the present invention provides a method for identifying core governance objective-tool combinations based on citation associations, including the following steps:

[0071] S1. Select a target domain, collect governance documents to build a data pool, and perform data cleaning operations; including the following steps:

[0072] S1.1. Construct a keyword retrieval strategy for the target domain, and retrieve files in the database whose titles contain the keyword;

[0073] S1.2. By reading the above files, identify and summarize the relevant keyword list for the target domain; based on the selected target domain keywords, the relevant texts are from the PKULAW database (https: / / www.pkulaw.com / ), and the retrieval formula is "title = ((keyword1) OR (keyword2) OR … OR (keyword 𝑛 ))), PY = (Start Time-Now);

[0074] S1.3. Remove duplicate texts according to the combination matching of the document number and title fields in the data pool, and only keep one of the texts;

[0075] S1.4. Automatically parse the text genre structure of the text based on natural language processing technology, and identify the basic elements such as the document title, publishing agency, document type, and the date of writing and effectiveness.

[0076] S2. Document Sample Data Annotation and Conversion: Collect a certain proportion of document samples. Combine the entity type definitions and prompt words, and use the Doccano software (an open-source multi-functional text annotation platform) to annotate discontinuous governance objectives, tools, and entity relationships. Then divide them into training, validation, and test sets according to the ratio of [0.7, 0.2, 0.1]. The steps are as follows:

[0077] S2.1. Governance Objective Annotation: Based on the governance objective entity type and combined with verb phrases such as "achieve / promote", use the Doccano software to use the pointer of the relationship to point from one end of the discontinuous entity to the other end. If it is a continuous entity, directly annotate the entity type as the governance objective;

[0078] S2.2. Governance Tool Annotation: Based on the governance tool entity type and combined with common verbs and nouns including semantic structures such as "formulate standards, establish systems, organize and carry out", use the Doccano software to use the pointer of the relationship to point from one end of the discontinuous entity to the other end. If it is a continuous entity, directly annotate the entity type as the governance tool;

[0079] S2.3. Annotation of the Governance Objective - Tool Relationship: Combine the relationship of adopting governance tools to achieve governance objectives, and annotate the relationship of the governance objective - tool mode;

[0080] S2.4. Annotation Data Reading: Assume that the input sentence X consists of N characters, , the model predicts the relationship category R of the two characters in each word pair ( ), where . Among them, None means there is no relationship between the two characters and they do not belong to the same entity; NNW, that is, Next - Neighboring - Word, means that the two characters are adjacent positions in the same entity; THW - * means that the two characters are in the same entity, and are respectively the end and the start of the entity, used to judge the category and boundary of the entity, where * represents the entity type. The annotation data reading is a word - word two - dimensional relationship matrix,

[0081] Such as Figure 2 shown. The annotated dataset is exported as a json file. The annotated entity set includes entity ID, category, start position, end position, and the entity relationship set contains start entity ID, target entity ID, and relationship category;

[0082] S2.5. Entity-relationship data preprocessing: Define the target data format, limit the length to the maximum sentence length, and store the index range and category information of entities and relationships as entities. If an entity is an isolated entity, treat it as a self-loop relationship and generate a virtual relationship of "the starting entity is equal to the target entity". Traverse the entities and remove the isolated entities that do not participate in any relationships, and extract the text index ranges of the starting entity and the target entity to generate relationship data.

[0083] S2.6. Dataset division: Divide it into a training set for model training, a validation set for model parameter tuning, and a test set for model performance testing according to a certain proportion.

[0084] S3. Construction of the W2NER model: Combine the Bidirectional Encoder Representation from Transformers (BERT), Bidirectional Long Short-Term Memory (LSTM), Convolutional Neural Network, and Biaffine mechanisms to construct a model for sequence labeling and relationship extraction tasks, extract the semantic features of the text, and extract governance objectives, tool entities, and relationships, as Figure 4 shown, including the following steps:

[0085] S3.1. Model receiving data preprocessing: Include the sentence word ID sequence, sub-word to word mapping matrix, two-dimensional mask matrix, entity distance matrix, and sentence length. Perform verification and masking processing on the input data to ensure consistent data format and mask invalid positions.

[0086] S3.2. Generation of context representation based on multi-layer feature extraction. Use the pre-trained BERT model to perform context encoding on the input sequence to extract word-level or sub-word-level embeddings; use the pre-trained BERT model to perform context encoding on the input sequence to extract word-level or sub-word-level embeddings; input the word-level embeddings into a bidirectional LSTM to capture the global context information of the sentence.

[0087] S3.3. Conditional normalization processing: Use the Conditional Layer Normalization (CLN) module to enhance the features of the LSTM output.

[0088] S3.4. Integrate distance and relation features, map the distance matrix between entities into a high-dimensional embedding representation to enhance the relation modeling ability, generate embeddings of relation types according to the relation matrix, further enrich the feature expression of the sentence, and concatenate the distance embedding, relation embedding, and context representation (the output of the LSTM processed by CLN) to form the feature input. As Figure 3 shown, where dist_inputs refers to the relative position information of grid characters, with dimension [B, L, L]; grid_labels refers to the two-dimensional matrix of THW and NHW relations constructed by labeled data entities, with dimension [B, L, L].

[0089] S3.5. Use a dilated convolution layer to perform local modeling on the feature input, and use the output of the convolution layer as the feature input for subsequent prediction modules.

[0090] S3.6. Perform bi-affine interaction modeling, project the entity representation into a specific space, calculate the interaction scores of entity pairs, and predict the relation categories between entities according to the interaction scores.

[0091] S4. Training of the W2NER model: A training model for joint extraction of text entities and relations based on adversarial training and dynamic learning rate adjustment; includes the following steps:

[0092] S4.1. Data loading and preprocessing, load the training set, validation set, and test set, the data format includes text, entity annotations, and relation annotations, and check and count the distribution of the loaded data to ensure the balance of training samples.

[0093] S4.2. Model initialization, construct a joint extraction model according to the configuration file, the model structure includes using BERT for context feature extraction, bidirectional LSTM to further capture the global context, a convolution layer for local feature modeling, and a bi-affine mechanism for joint prediction of entities and relations; set the loss function and group parameters, and set a learning rate scheduler to dynamically adjust the learning rate.

[0094] S4.3. Model training, set the model to training mode, perform forward propagation, calculate the loss, adversarial training, gradient clipping, and parameter update steps for each batch of data, and update the learning rate; after each training epoch, calculate the average loss and record the prediction accuracy, recall rate, and F1 score.

[0095] S4.4. Model validation, set the model to evaluation mode, verify batch by batch, perform forward propagation and prediction, and count the performance, and save the current model parameters with the F1 score of the validation set higher than the historical best.

[0096] S4.5, Early stopping mechanism: When the performance on the validation set fails to improve for several consecutive rounds, record the number of stagnations; when the number of stagnations reaches the set threshold, end the training early to avoid overfitting.

[0097] S4.6, Model saving and loading: Save the model parameter file at the specified path for subsequent loading or inference; load the saved model parameters to support continued training or direct evaluation.

[0098] S5, W2NER model extraction: Extract governance objectives, tool entities, and relationships between entities in the text based on a deep learning-based method for joint extraction of text entities and relationships; including the following steps:

[0099] S5.1, Text tokenization: Perform character-level tokenization on the input text, replace consecutive spaces with a single space, and remove extra characters.

[0100] S5.2, Data loading and construction: Use the tokenized text to construct a prediction dataset, including the text tokenization results and auxiliary features. After encapsulating the text into the model input format, load the data in batches.

[0101] S5.3, Model loading and initialization: Load the joint extraction model and load the trained model parameters.

[0102] S5.4, Model inference: Input the batch-loaded data into the model, perform forward propagation batch by batch, calculate the class scores at each position, and classify and output the entity class (objective or tool) and the entity relationship class.

[0103] S5.5, Decoding and result generation: Decode the indices and classes output by the model, map the decoded entity and relationship results to a readable format, and finally generate a JSON file of the extraction results.

[0104] S5.6, Result storage: Output the extracted governance objectives, tools, and relationships, and store them in a local file according to the file ID and governance combination (governance objective, realization, governance tool).

[0105] S6, Construction of governance objective-tool citation network: Extract the citation and cited relationships between texts based on the book title marks and citation prompt words in the text content; identify, clean, and integrate governance objectives and tools in the text through the W2NER model, and construct a governance objective-tool citation network based on this information; including the following steps:

[0106] S6.1, Identifier recognition: Efficiently identify the book title marks embedded in the text content based on the string matching algorithm (Knuth-Morris-Pratt, KMP).

[0107] S6.2. Improve the file name, obtain the file publishing agency, and construct the complete file name by combining the publishing agency and the specific file within the book title marks;

[0108] S6.3. Supplement the citation relationships, identify the prompt words "based on", "implement", "according to", "regulation", "combined with", "followed by" before and after the book title marks, and further identify and supplement the citation relationships;

[0109] S6.4. Store the citation relationships. According to the logic of the files cited in this article and the files citing this article, store the file citation relationships (citing file ID, cited file ID);

[0110] S6.5. Node construction and attribute mapping. Each governance goal or tool is defined as a node in the citation network. The attributes of the node include: node name, type (governance goal or tool), source text ID, and frequency of occurrence, etc. The size of the node is proportional to the frequency of occurrence of the governance goal or tool in all texts. The nodes with higher frequencies will be displayed as larger nodes in the network;

[0111] S6.6. Governance goal - tool combination citation network. Combine the governance goals and tools in the citation network, and construct a "governance goal - tool" combination citation network according to the existing governance goal - tool combination patterns. Each text data will contain multiple governance goal - tool combination patterns, indicating the many - to - many relationship between governance goals and tools.

[0112] S7. Identification of core governance goal - tool: Use the Ego - Sphere Unit (ESU) algorithm to perform a directed motif analysis on the governance goal - tool citation network. Identify the induced subgraphs in the network through a recursive method, especially the multiple relationships and interaction patterns between the target nodes and tool nodes, and based on this, identify the core governance goal - tool combinations and analyze their evolution processes. It includes the following steps:

[0113] S7.1. Use the ESU algorithm for recursive mining: Apply the ESU (Ego - Sphere Unit) algorithm to mine the subgraphs in the governance goal - tool citation network through a recursive method. Each recursion will expand the SUB set, which contains the nodes directly adjacent to the current node, and the EXT set, which contains the nodes adjacent to at least one node in the SUB set and with a larger numerical label. The ESU algorithm identifies all eligible induced subgraphs through this recursive process. The pseudocode of the ESU subgraph search algorithm is as Figure 5 shown;

[0114] S7.2, Motif Analysis and Subgraph Classification: Motif analysis is performed on each branch of the ESU-Tree (ESU-Tree), and frequently occurring subgraph structures (motifs) in the network are identified through a recursive method. During this process, the ESU algorithm performs an isomorphism test on each subgraph through McKay's nauty algorithm to ensure the classification of subgraphs with different structures and sorts them according to frequency and concentration;

[0115] A schematic diagram of the ESU-Tree tree structure is as Figure 6 shown, attached Figure 6 in (a) shows the target graph of size 5; attached Figure 6 in (b) shows the ESU-Tree of depth k, which is used to extract subgraphs of size 3 from the target graph. The leaf nodes correspond to the set S4 in the target attached Figure 6 in (a) or all induced subgraphs of size 4. The nodes in the ESU-Tree include two adjacent sets. The first set is SUB, which contains the nodes directly adjacent to the current node; the second set is EXT, which contains all nodes adjacent to at least one SUB node and whose digital labels are greater than the SUB nodes. The algorithm expands the SUB set by expanding the EXT set until the required subgraph size (here 3) is reached, and this subgraph size is at the bottom layer of the ESU-Tree (or its leaf nodes).

[0116] S7.3, Identifying Core Governance Objective-Tool Combinations: Frequent subgraphs (core governance objective-tool combinations) in the network are identified through the ESU algorithm, and the frequency F R (G′) of each subgraph in the random network is calculated. At the same time, the actual frequency F G (G′) of this subgraph in the target network is calculated. The P-value is used for statistical hypothesis testing to evaluate whether a subgraph appears significantly in the target network; at the same time, the N value is set to make the evaluation of subgraph frequency more stable and accurate by generating multiple random network samples;

[0117] ;

[0118] where, N is the number of random network samples, is the P-value, given with a probability of (as its null-hypothesis), represents the frequency of in the random network; when the P value is less than the threshold of 0.05, this subgraph can be called a significant pattern.

[0119] S7.5, Core Governance Objective - Visualization of Tool Motifs. Based on the identification of core governance objective - tool combinations, visualization technology is used to display the motif network structure. The governance objectives and tools are used as network nodes, and the size and color of the nodes represent the frequency and type of the nodes. The thickness of the edges reflects the strength of the citation relationship between the governance objective and the tool. During this visualization process, the interaction relationship of the core objective - tool combination is highlighted, and their importance and connection patterns in the overall system are displayed through dynamic visualization;

[0120] S7.6, Citation Network Based on Core Governance Objective - Tool Combination Motifs, Analyze its evolution process at different times or during the implementation stage of governance documents. By tracking the trend of the core motif over time, identify the alternation and changes in the combination patterns between governance objectives and tools. This analysis helps to reveal the evolution path of governance objectives and tools, predict future governance directions, and adjust existing governance combinations to improve their implementation effects.

[0121] The embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the methods for identifying a core governance objective - tool combination based on citation association.

Claims

1. A method for identifying core governance objectives and tool combinations based on reference associations, characterized in that: The following steps are involved: (1) Select the target area, collect governance documents, build a data pool and perform preprocessing; (2) Text data annotation and conversion: Collect sample data, and use Doccano software to annotate discontinuous governance goals, tools, and relationships between entities according to the prompt words of the goals and tools, and divide them into training, validation, and test sets; (3) Constructing the W2NER model: Combining the bidirectional transformer model BERT, bidirectional long short-term memory LSTM, convolutional layer and biaffine mechanism to build a model for sequence labeling and relationship extraction tasks, extracting the semantic features of the text, and extracting governance objectives, tool entities and relationships; (4) Training the W2NER model: a training model for the joint extraction of text entities and relations based on adversarial training and dynamic learning rate adjustment; (5) Extraction W2NER model: A text entity and relationship joint extraction method based on deep learning is used to extract governance objectives, tool entities, and entity relationships in the text; (6) Constructing a governance goal-tool citation network: Extracting the citation and cited relationships between texts based on the book title marks and citation prompts in the document content; using the W2NER model to identify, clean, and integrate the governance goals and tools in the text, and constructing a governance goal-tool citation network based on this information; (7) Identify core governance goals and tools: Use the self-sphere unit algorithm (ESU) to perform directed motif analysis on the governance goal-tool reference network, recursively identify the induced subgraph in the network, that is, the multiple relationships and interaction patterns between the goal nodes and the tool nodes, and analyze its evolution process based on the identification of the core governance goal-tool combination; the following steps are included: (71) Use the ESU algorithm for recursive mining: Apply the ESU algorithm to mine the subgraphs in the governance goal-tool reference network in a recursive manner; each recursion will expand the SUB set, which contains nodes directly adjacent to the current node, and the EXT set contains nodes that are adjacent to at least one node in the SUB set and have a larger numerical label; the ESU algorithm identifies all eligible induced subgraphs through this recursive process; (72) Motif analysis and subgraph classification: Motif analysis is performed on each branch of the ESU tree, and a recursive method is used to identify subgraph structures that frequently appear in the network. During this process, the ESU algorithm performs an isomorphism test on each subgraph using McKay's nauty algorithm to ensure that subgraphs of different structures are classified and sorted according to their frequency and concentration. (73) Identify the core governance goal-tool combination: Use the ESU algorithm to identify the frequent subgraphs in the network, that is, the core governance goal-tool combination, and calculate the frequency F of each subgraph in the random network. R (G′); at the same time, calculate the actual frequency F of the subgraph in the target network G (G′); use the P value to perform statistical hypothesis testing to evaluate whether a subgraph appears significantly in the target network; at the same time, set the N value to generate multiple random network samples to make the evaluation of subgraph frequency more stable and accurate; the formula is as follows: ; in, N is the number of random network samples, is the P value, The probability of is given, In a random network frequency; when P When the value is less than the threshold of 0.05, the sub-image is called a significant pattern; (74) Visualization of the core governance goal-tool motif. Based on the identification of the core governance goal-tool combination, the network structure of the motif is displayed using visualization technology. The governance goals and tools are used as network nodes, and the size and color of the nodes represent the frequency and type of the nodes. The thickness of the edges reflects the strength of the citation relationship between the governance goals and tools. (75) Based on the citation network of the core governance objective-tool combination motif, analyze its evolution process at different times or document implementation stages; by tracking the trend of core motifs over time, identify changes in the alternation and combination patterns between governance objectives and tools.

2. According to claim 1, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: In step (1), the preprocessing is to clean the data.

3. According to claim 2, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: Step (1) includes the following steps: (11) Construct a keyword search strategy for the target field and search the database for target document data whose titles contain keywords; (12) Identify and summarize a list of relevant keywords in the target field by reading document data; (13) Remove duplicate files by combining and matching the two fields of the document number and title in the data pool, and only keep one file; (14) Automatically parse the text structure of text data based on natural language processing technology to identify the basic elements such as document title, publishing agency, document type, and date of writing and effective date.

4. According to claim 1, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: Step (2) includes the following steps: (21) Label the governance objectives. Based on the governance objective entity type and the verb phrase "realize / promote", use the Doccano software to use the relationship pointer from one end of the discontinuous entity to the other end. If it is a continuous entity, directly label it as the entity type of governance objective. (22) Labeling governance tools. Based on the entity type of governance tools and common verbs and nouns, including the semantic structure of "setting standards and establishing systems", the Doccano software is used to use the relationship pointer from one end of the discontinuous entity to the other end. If it is a continuous entity, it is directly labeled as the entity type of governance tool; (23) Mark the relationship between governance objectives and tools, combine the relationship between adopting governance tools to "achieve" governance objectives, and mark the relationship between governance objectives and tool models; (24) Read the annotated data and export the annotated data set as a json file, which contains the annotated entity set and the annotated entity relationship set; (25) Preprocessing of entity-relationship data: defining the target data format, removing isolated entities that are not involved in any relationship, and extracting the text index range of the starting entity and the target entity to generate relationship data; (26) Dataset division: divide the data into a training set for model training, a validation set for model parameter tuning, and a test set for model performance testing.

5. According to claim 1, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: Step (3) includes the following steps: (31) The model receives preprocessed data, including sentence word ID sequence, subword to word mapping matrix, two-dimensional mask matrix, entity distance matrix and sentence length, and verifies and masks the input data to make the data format consistent and mask invalid positions; (32) Generate contextual representations based on multi-layer feature extraction. Use the pre-trained BERT model to contextually encode the input sequence and extract word-level or sub-word-level embeddings. Use the pre-trained BERT model to contextually encode the input sequence and extract word-level or sub-word-level embeddings. Input the word-level embeddings into a bidirectional LSTM to capture the global context information of the sentence. (33) Conditional normalization processing, using the conditional layer normalization module to enhance the features of LSTM output; (34) Fusion of distance and relation features, mapping the distance matrix between entities into a high-dimensional embedding representation, generating an embedding of the relation type based on the relation matrix, and concatenating the distance embedding, relation embedding, and context representation, i.e., the LSTM output processed by CLN, to form the feature input; (35) Use a multi-expansion rate convolutional layer to locally model the feature input and use the convolutional layer output as the feature input; (36) Bi-affine interaction modeling projects entity representations into a specific space, calculates the interaction scores of entity pairs, and predicts the relationship categories between entities based on the interaction scores.

6. According to claim 1, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: Step (4) includes the following steps: (41) Data loading and preprocessing: loading training sets, validation sets, and test sets. The data formats include text, entity annotations, and relationship annotations. The distribution of the loaded data is checked and counted. (42) Model initialization: construct a joint extraction model according to the configuration file. The model structure includes using BERT for context feature extraction, bidirectional LSTM to further capture global context, convolutional layer for local feature modeling, and dual affine mechanism for joint prediction of entities and relationships; setting the loss function and grouping parameters, and setting the learning rate scheduler to dynamically adjust the learning rate; (43) Model training: Set the model to training mode, perform forward propagation, loss calculation, adversarial training, gradient clipping and parameter update, and learning rate update steps for each batch of data; after each training round, calculate the average loss and record the predicted accuracy, recall rate, and F1 score; (44) Model verification: set the model to evaluation mode, verify batch by batch, perform forward propagation and prediction, and statistically analyze the performance, saving the current model parameters whose F1 score on the verification set is higher than the historical best; (45) Early stopping mechanism: when the performance of the validation set fails to improve after several consecutive rounds, the number of stagnations is recorded; when the number of stagnations reaches the set threshold, the training is terminated early to avoid overfitting; (46) Model saving and loading: save the model parameter file in the specified path for subsequent loading or inference; load the saved model parameters to support continued training or direct evaluation.

7. According to claim 1, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: Step (5) includes the following steps: (51) Text segmentation: perform character-level segmentation on the input text, replace consecutive spaces with single spaces, and remove redundant characters; (52) Data loading and construction: Use the segmented text to build a prediction dataset, including the text segmentation results and auxiliary features. After encapsulating the text into the model input format, load the data in batches; (53) Model loading and initialization, loading the joint extraction model, and loading the trained model parameters; (54) Model reasoning: input batch-loaded data into the model, perform forward propagation batch by batch, calculate the category score of each position, and classify and output the entity category, i.e., the target or tool and entity relationship category; (55) Decoding and result generation: decode the index and category of the model output, map the decoded entity and relationship results into a readable format, and finally generate a JSON file of the extraction results; (56) Result storage: output the extracted governance goals, tools and relationships, and store the governance combination in a local file according to the text ID; the governance combination includes: governance goals, implementation, and governance tools.

8. According to claim 1, a method for identifying core governance objectives and tool combinations based on reference associations is characterized in that: Step (6) includes the following steps: (61) Identification of identifiers: identifying book title marks embedded in text content based on the string matching algorithm KMP; (62) Complete the file name, obtain the issuing agency of the file, and construct the complete file name by combining the issuing agency and the specific file in the quotation marks; (63) Supplement the citation relationship, identify the prompt words "based on", "implemented", "according to", "regulated", "combined with", and "follow" before and after the quotation marks, and further identify and supplement the citation relationship; (64) Citation relationship storage: according to the logic of the files cited in this article and the files citing this article, the file citation relationship, i.e., the citing file ID and the cited file ID, is stored; (65) Node construction and attribute mapping. Each governance goal or tool is defined as a node in the reference network. The attributes of the node include: node name, type (i.e., governance goal or tool), source text ID, and frequency of occurrence. The size of the node is proportional to the frequency of occurrence of the governance goal or tool in all texts. Nodes with higher frequencies will appear as larger nodes in the network. (66) The governance goal-tool combination citation network is constructed by combining the governance goals and tools in the citation network and according to the existing governance goal-tool combination patterns. Each text data will contain multiple governance goal and tool combination patterns, indicating a many-to-many relationship between governance goals and tools.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, it implements a method for identifying a core governance objective-tool combination based on reference association according to claim 1.

Citation Information

Patent Citations

  • Graph classification method and system fusing high-order structure embedding and composite pooling

    CN114792384A

  • Event extraction method and device

    CN116306581A