Policy blood relationship construction method and system

By automating the processing of policy documents through web crawling, NLP, and data visualization technologies, accurate policy lineage can be constructed, solving the problems of low accuracy and efficiency in existing technologies, and making it suitable for policy research in multiple fields.

CN121880633APending Publication Date: 2026-04-17WUXI BASIC PARTICLE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUXI BASIC PARTICLE TECHNOLOGY CO LTD
Filing Date
2023-09-11
Publication Date
2026-04-17

Smart Images

  • Figure CN121880633A_ABST
    Figure CN121880633A_ABST
Patent Text Reader

Abstract

The invention discloses a policy blood relationship construction method and system. The method comprises the following steps: firstly, crawling and collecting an original policy document through a crawler, and preprocessing and cleaning the collected original data to obtain a target document; identifying field information of the target document through a document content identification technology, and performing association policy analysis processing on the target document by using an NLP technology; integrating the field information and the associated policy analysis processing result to obtain policy blood relationship data, and storing and managing the policy blood relationship data through a relational database or a graph database; and finally, displaying the policy blood relationship data in a chart or graph form by using a data visualization technology. And verifying the accuracy and credibility of the constructed policy blood relationship through a result verification mechanism. The policy consanguinity can be more accurately constructed, the policy can be comprehensively and accurately traced and tracked, and the policy research efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, and in particular to a method and system for constructing policy kinship relationships. Background Technology

[0002] When formulating new policies or adjusting existing ones, businesses need to make decisions based on an understanding of the existing policies. Analyzing policy relationships can help policymakers better grasp the relationship between new and existing policies, ensuring that new policies are consistent with current policies or addressing problems arising from existing policies.

[0003] In the process of policy research and decision-making, accurately understanding and tracing the evolution and lineage of policies is crucial. However, existing methods for constructing policy lineages have the following problems:

[0004] Lack of precision: Existing methods are prone to errors or omissions in the construction of policy lineage, leading to inaccurate results. This poses a challenge for researchers and policy practitioners.

[0005] Time-consuming and cumbersome: Traditional manual or semi-automatic methods require a lot of manual operation and time-consuming policy data collection process, resulting in low efficiency in policy research.

[0006] Limited to specific fields: Existing technologies are often limited by specific fields or types of policies during the policy lineage construction process, and cannot meet diverse application needs. Summary of the Invention

[0007] Based on this, embodiments of this application provide a method and system for constructing policy lineage relationships, which can effectively help understand and analyze the interrelationships and influences between different policies.

[0008] Firstly, a method for constructing policy kinship is provided, which includes:

[0009] Raw policy documents are collected by web crawling, and the collected raw data is preprocessed and cleaned to obtain the target documents.

[0010] The target document's field information is identified using document content recognition technology;

[0011] NLP technology is used to perform policy-related analysis on the target document;

[0012] The field information and the results of related policy analysis are integrated to obtain policy lineage data, which is then stored and managed through a relational database or graph database.

[0013] Use data visualization techniques to display policy lineage data in the form of charts or graphs; this can be achieved using chart libraries, graph libraries, or visualization tools.

[0014] The accuracy and credibility of the constructed policy lineage are verified through the results verification mechanism.

[0015] Optionally, the method further includes:

[0016] Continuously acquire new policy documents and update and expand the original policy lineage data.

[0017] Optionally, the verification mechanism for the accuracy and credibility of the constructed policy lineage specifically includes:

[0018] Collect a set of standard datasets with known policy lineages as a reference dataset;

[0019] For the collected reference dataset, the lineage relationships in the policy lineage data are labeled; specifically, the associations and evolutionary relationships between policies are marked.

[0020] A certain number of samples were selected from the constructed policy lineage data to participate in the verification;

[0021] The association results of the labeled policy lineage data are compared with the labels of the reference dataset. The accuracy of the results is evaluated by comparing the consistency and similarity between the two.

[0022] The verification results are statistically analyzed to generate a verification result report.

[0023] Optionally, the accuracy assessment of the results by comparing the consistency and similarity between the two includes:

[0024] Specifically, the accuracy of the results is evaluated by calculating precision and recall; where precision represents the proportion of correctly identified samples out of all samples judged as identical or similar; and recall represents the proportion of correctly identified samples out of all truly identical or similar samples.

[0025] Optionally, the collected raw data is preprocessed and cleaned to obtain the target document, including:

[0026] The collected raw data is cleaned, missing values ​​are handled, and outliers are removed.

[0027] Optionally, the step of using NLP technology to perform policy correlation analysis on the target document further includes:

[0028] Relationship extraction techniques are used to extract information about the relationships between policies from policy texts; this relationship information includes at least policy references, impacts, and similarities.

[0029] Optionally, after integrating the field information and the results of related policy analysis to obtain policy lineage data, the method further includes:

[0030] Constructing a knowledge graph in the policy domain involves structurally representing and linking policy documents and related knowledge; ontology modeling and semantic networks can be used to construct the knowledge graph.

[0031] Optionally, the field information includes at least the document title, document category, document number, issuing authority, policy document level, industry category, issuance time, and policy document region; the associated policy analysis and processing includes at least word frequency statistics, part-of-speech tagging, entity recognition, and syntactic analysis.

[0032] Secondly, a policy kinship construction system is provided, which includes:

[0033] The collection module is used to crawl and collect raw policy documents, and to preprocess and clean the collected raw data to obtain the target documents.

[0034] The fill recognition module is used to identify field information of the target document using document content recognition technology;

[0035] The correlation processing module is used to perform correlation policy analysis on target documents using NLP technology.

[0036] The data integration module is used to integrate the field information and the results of related policy analysis to obtain policy lineage data, and to store and manage the policy lineage data through a relational database or graph database.

[0037] The data visualization module is used to display policy lineage data in the form of charts or graphs using data visualization technology; this can be achieved using chart libraries, graph libraries, or visualization tools.

[0038] The results verification module is used to verify the accuracy and credibility of the constructed policy lineage through a results verification mechanism.

[0039] Optionally, the system further includes:

[0040] The update extension module is used to continuously acquire new policy documents and update and expand the original policy lineage data.

[0041] The beneficial effects of the technical solutions provided in this application include at least the following:

[0042] (1) Precise policy lineage: Compared with the prior art, the present invention can more accurately construct policy lineage, and realize comprehensive and accurate tracing and tracking of policies. This can ensure that the results of policy lineage construction are more accurate and reliable.

[0043] (2) Improved policy research efficiency: This invention employs a highly efficient technical means to construct policy lineages, which greatly improves the efficiency of policy research compared to traditional manual or semi-automatic methods. Operators can quickly and accurately obtain policy-related information, avoiding tedious manual operations.

[0044] (3) Wider range of applications: The technical means employed in this invention are applicable to policy lineage construction in various fields, not limited to specific industries or types of policies. This allows the policy lineage construction method to be applied to more diverse scenarios, expanding its scope of application. It can provide enterprises with comparative results of similar policies in different cities, enabling them to decide which city to register, settle, and develop in. Attached Figure Description

[0045] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0046] Figure 1 A flowchart for constructing policy kinship relationships is provided for an embodiment of this application. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0048] In the description of this invention, the terms “comprising,” “having,” and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may also include other steps or units that are not expressly listed but are inherent to these processes, methods, products, or apparatuses, or steps or units added based on further optimizations of the inventive concept.

[0049] Existing technologies lack an effective method to help policymakers and implementers comprehensively understand and analyze the policy lineage. Specific methods for constructing policy lineages require selection and combination based on specific needs and problems. This application invents a method for constructing policy lineages. Through policy lineage construction, policies, publicized lists, payment records, etc., can be linked to policy documents, and policy documents can be linked to guiding documents from higher levels, ensuring that every policy and every payment has a basis, vertically displaying the source, connections, and citations of each policy. By comparing the similarity of policy texts, policies with similar content can also be identified, and the quality of policies in different regions can be compared horizontally. For details, please refer to... Figure 1 The document illustrates a flowchart of a method for constructing policy kinship relationships according to an embodiment of this application. This method may include the following steps:

[0050] S1: Raw policy documents are collected by web crawler, and the collected raw data is preprocessed and cleaned to obtain the target documents.

[0051] In this application embodiment, the relevant policy documents mainly refer to the policy documents of the enterprise. This step involves document collection and preprocessing, specifically:

[0052] Policy documents, policies, and publicized lists are collected through web scraping, including publicly available information from companies, research reports, and news media reports, and stored in a database. Data preprocessing and cleaning techniques are then used to process and clean the collected policy data. This improves data quality and reduces interference during subsequent analysis.

[0053] S2 identifies the field information of the target document through document content recognition technology.

[0054] This step implements automatic filling, specifically by using document content automatic recognition technology to identify and automatically fill in the field information within the document. The fields include: document title, document category, document number, issuing authority, policy document level, industry category, issuance time, and policy document region.

[0055] S3 uses NLP technology to perform policy-related analysis on the target document.

[0056] Natural Language Processing (NLP): NLP technology can be used to process and analyze policy texts, extracting information such as keywords, themes, and entities to help understand the content and semantics of policy documents.

[0057] This step implements policy association by using NLP technology to process and analyze policy texts, including word frequency statistics, part-of-speech tagging, entity recognition, and syntactic analysis. Relationship extraction technology is used to extract information about the relationships between policies from the policy texts, such as policy citations, impacts, and similarities. Relationship extraction can be based on rules, machine learning, or deep learning methods. Furthermore, the consistency of publication time, issuing agency, policy region, and keywords is carefully compared to comprehensively identify the relationships between policies and documents, and between documents themselves.

[0058] Among them, Relation Extraction can extract information about the relationships between policies from policy texts, such as policy references, impacts, and similarities, providing support for policy lineage construction.

[0059] Data mining: Data mining techniques can utilize information and attributes in policy texts to perform pattern recognition, cluster analysis, association rule mining, etc., to help reveal the implicit relationships and characteristics between policies.

[0060] S4 integrates field information and related policy analysis results to obtain policy lineage data, and stores and manages the policy lineage data through a relational database or graph database.

[0061] This step involves data integration, specifically consolidating the relationships between policies and establishing their lineage. Relational databases or graph databases are used to store and manage this policy lineage data.

[0062] In optional embodiments of this application, knowledge graph construction may also be included. By constructing a knowledge graph in the policy domain, policy documents and related knowledge are represented and linked in a structured manner. Ontology modeling and semantic networks can be used to construct the knowledge graph. The knowledge graph of this application, by constructing a knowledge graph in the policy domain and representing and linking policy documents and related knowledge in a structured manner, can help with policy lineage analysis and relationship reasoning.

[0063] S5 uses data visualization technology to display policy lineage data in the form of charts or graphs.

[0064] This is achieved using chart libraries, graph libraries, or visualization tools. This step implements data visualization, using data visualization techniques to present the policy kinship network to users in the form of charts or graphs, allowing users to intuitively understand the relationships and impacts between policies. Chart libraries, graph libraries, or visualization tools can be used for this purpose.

[0065] S6 verifies the accuracy and credibility of the constructed policy lineage through a results verification mechanism.

[0066] An outcome verification mechanism is introduced to verify the accuracy and credibility of the constructed policy lineage.

[0067] The specific process of the result verification mechanism includes:

[0068] Reference Dataset Collection: A set of standard datasets with known policy lineages are collected as a reference. These datasets have been validated and approved.

[0069] Reference dataset annotation: For the collected reference dataset, lineage relationships are annotated, that is, the connections and evolutionary relationships between policies are indicated. These annotations can serve as reference standards for verification.

[0070] Validation Sample Selection: A certain number of samples are selected from the constructed policy lineage data to participate in the validation. These samples should have different types of policy associations to ensure the comprehensiveness and representativeness of the validation.

[0071] Verification result comparison: The association results obtained using the policy lineage construction method of this invention are compared with the annotations of the reference dataset. The accuracy of the results is evaluated by comparing the consistency and similarity between the two.

[0072] Validation metric evaluation: The validation results are quantitatively evaluated by defining appropriate validation metrics (such as precision, recall, etc.). These metrics will be used to measure the performance and accuracy of the policy lineage construction method.

[0073] Precision and recall are performance metrics used to evaluate a classification model. The specific calculation process is as follows:

[0074] True Positive (TP): The number of samples that the model correctly predicts as positive.

[0075] False positives (FP): The number of samples that the model incorrectly predicts as positive.

[0076] False Negative (FN): The number of samples that the model incorrectly predicts as negative examples.

[0077] Accuracy refers to the proportion of samples that are actually true positives out of all samples predicted as positive by the model. Accuracy = TP / (TP + FP).

[0078] Recall calculation: Recall is the proportion of samples that are actually positive that are predicted as positive by the model. Recall = TP / (TP + FN).

[0079] Validation Result Report: The validation results are statistically analyzed to generate a validation result report. This report will describe in detail the performance of the policy lineage construction method during the validation process, including numerical values ​​for metrics such as accuracy and recall.

[0080] The innovation of this invention lies in:

[0081] Automated processing: The use of natural language processing and text analysis technologies to automate the processing of policy documents greatly improves the efficiency of policy lineage construction.

[0082] Relevance Integration: By classifying and integrating key policy information, the relevance between policies can be accurately captured, thereby forming an accurate policy lineage network.

[0083] Visualization and Analysis: Provides user-friendly visualization tools and powerful analytical capabilities, enabling users to delve into policy lineages and discover potential influencing factors and evolutionary trends.

[0084] S7 continuously acquires new policy documents and updates and expands the original policy lineage data.

[0085] In this step, the kinship map needs to be continuously updated and expanded as new policies are formulated and implemented. Therefore, the method should be scalable and flexible to construct new policy kinship lines.

[0086] By comprehensively utilizing the above technical solutions, a complete policy lineage analysis system can be constructed, providing a user-friendly interface and functions to help users gain a deeper understanding of the correlation and mutual influence between policies, and to provide a scientific basis for policy formulation and decision-making.

[0087] In summary, the present invention employs the following technical means to achieve a more accurate construction of policy kinship relationships:

[0088] Automated text analysis and processing: Utilizing natural language processing and text mining techniques, this function automates the text analysis and processing of policy documents to extract key information such as policy titles, paragraphs, and keywords.

[0089] Among them, text mining: text mining technology can discover the correlation, co-occurrence patterns and trends between policy texts through processing and analysis, providing a basis for policy lineage analysis.

[0090] Semantic linking and correlation analysis: This technique compares and correlates key information in different policy documents to identify similarities and connections. This method can help determine the lineage between policies.

[0091] This invention employs the following technical means to improve the efficiency of policy research:

[0092] Automated policy data collection: By using web crawlers and data scraping technology, policy documents and related information are collected automatically, avoiding the tedious manual collection process.

[0093] Data preprocessing and cleaning: Data preprocessing and cleaning techniques are used to process and clean the collected policy data. This improves data quality and reduces interference in subsequent analysis processes.

[0094] Highly efficient correlation analysis algorithms: Employing efficient correlation analysis algorithms can accurately identify the connections and relationships between policy documents. This allows researchers to quickly obtain the necessary policy correlation information, improving the efficiency of policy research.

[0095] This application also provides a policy kinship construction system. The system includes:

[0096] The collection module is used to crawl and collect raw policy documents, and to preprocess and clean the collected raw data to obtain the target documents.

[0097] The fill recognition module is used to identify field information of the target document using document content recognition technology;

[0098] The correlation processing module is used to perform correlation policy analysis on target documents using NLP technology.

[0099] The data integration module is used to integrate field information and related policy analysis results to obtain policy lineage data, and to store and manage the policy lineage data through a relational database or graph database.

[0100] The data visualization module is used to display policy lineage data in the form of charts or graphs using data visualization technology; this can be achieved using chart libraries, graph libraries, or visualization tools.

[0101] The results verification module is used to verify the accuracy and credibility of the constructed policy lineage through a results verification mechanism.

[0102] In optional embodiments of this application, the system further includes:

[0103] The update extension module is used to continuously acquire new policy documents and update and expand the original policy lineage data.

[0104] The policy kinship construction system provided in this application embodiment is used to implement the above-described policy kinship construction method. Specific limitations of the policy kinship construction system can be found in the above-described limitations of the policy kinship construction method, and will not be repeated here. Each part of the above-described policy kinship construction system can be implemented wholly or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the device in hardware form, or stored in the memory of the device in software form, so that the processor can call and execute the operations corresponding to each module.

[0105] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0106] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A policy kinship construction method, characterized in that, The method includes: Raw policy documents are collected by web crawling, and the collected raw data is preprocessed and cleaned to obtain the target documents. The target document's field information is identified using document content recognition technology; NLP technology is used to perform policy-related analysis on the target document; The field information and the results of related policy analysis are integrated to obtain policy lineage data, which is then stored and managed through a relational database or graph database. Use data visualization techniques to display policy lineage data in the form of charts or graphs; this can be achieved using chart libraries, graph libraries, or visualization tools. The accuracy and credibility of the constructed policy lineage are verified through the results verification mechanism.

2. The policy kinship building method of claim 1, wherein, The method further includes: Continuously acquire new policy documents and update and expand the original policy lineage data.

3. The policy kinship building method of claim 1, wherein, The aforementioned result verification mechanism verifies the accuracy and credibility of the constructed policy lineage relationships, specifically including: Collect a set of standard datasets with known policy lineages as a reference dataset; For the collected reference dataset, the lineage relationships in the policy lineage data are labeled; specifically, the associations and evolutionary relationships between policies are marked. A certain number of samples were selected from the constructed policy lineage data to participate in the verification; The association results of the labeled policy lineage data are compared with the labels of the reference dataset. The accuracy of the results is evaluated by comparing the consistency and similarity between the two. The verification results are statistically analyzed to generate a verification result report.

4. The method for constructing policy kinship according to claim 3, characterized in that, The accuracy of the results is assessed by comparing the consistency and similarity between the two, including: Specifically, the accuracy of the results is evaluated by calculating precision and recall; where precision represents the proportion of correctly identified samples out of all samples judged as identical or similar; and recall represents the proportion of correctly identified samples out of all truly identical or similar samples.

5. The method for constructing policy kinship according to claim 1, characterized in that, The collected raw data is preprocessed and cleaned to obtain the target documents, including: The collected raw data is cleaned, missing values ​​are handled, and outliers are removed.

6. The method for constructing policy kinship according to claim 1, characterized in that, The process of using NLP technology to perform policy correlation analysis on target documents also includes: Relationship extraction techniques are used to extract information about the relationships between policies from policy texts; this relationship information includes at least policy references, impacts, and similarities.

7. The method for constructing policy kinship according to claim 1, characterized in that, After integrating the aforementioned field information and the results of related policy analysis to obtain policy lineage data, the data also includes: Constructing a knowledge graph in the policy domain involves structurally representing and linking policy documents and related knowledge; ontology modeling and semantic networks can be used to construct the knowledge graph.

8. The method for constructing policy kinship according to claim 1, characterized in that, The field information includes at least the document title, document category, document number, issuing authority, policy document level, industry category, issuance time, and policy document region; the associated policy analysis and processing includes at least word frequency statistics, part-of-speech tagging, entity recognition, and syntactic analysis.

9. A policy kinship construction system, characterized in that, The system includes: The collection module is used to crawl and collect raw policy documents, and to preprocess and clean the collected raw data to obtain the target documents. The fill recognition module is used to identify field information of the target document using document content recognition technology; The correlation processing module is used to perform correlation policy analysis on target documents using NLP technology. The data integration module is used to integrate the field information and the results of related policy analysis to obtain policy lineage data, and to store and manage the policy lineage data through a relational database or graph database. The data visualization module is used to display policy lineage data in the form of charts or graphs using data visualization technology; this can be achieved using chart libraries, graph libraries, or visualization tools. The results verification module is used to verify the accuracy and credibility of the constructed policy lineage through a results verification mechanism.

10. The policy kinship construction system according to claim 9, characterized in that, The system also includes: The update extension module is used to continuously acquire new policy documents and update and expand the original policy lineage data.