File analysis method and device, medium and product
By constructing a clause tree and conducting semantic vector analysis, the problem of inefficient interpretation of policy documents in the government system is solved, intelligent consistency analysis and conflict detection of policy documents are realized, and the accuracy and transparency of policy implementation are improved.
Patent Information
- Application Number
- CN202510778448.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The interpretation of policy documents in the existing government affairs system relies on manual analysis and judgment, which has strong subjectivity, low efficiency, and high missed detection rate. Due to the heterogeneity of data, it is difficult to quantify policy implementation indicators, making it difficult to accurately identify the implicit meaning and contextual relationships in the clause.
By converting policy documents into clause trees, using the contextual relationship of clause statements for vectorization, calculating cosine similarity and point mutual information between semantic vectors, building a structured clause tree for policy document analysis, combining reinforcement learning and adaptive update models, an intelligent policy consistency analysis is achieved.
It improves the accuracy and efficiency of policy document conflict detection, helps decision makers quickly identify and resolve cross-level and cross-departmental policy conflicts, improves the consistency and transparency of policy implementation, and supports policy optimization and dynamic adaptability.
Smart Images

Figure CN120295978A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to a method, device, medium and product for file analysis. Background Art
[0002] With the development of digital technology, more and more work scenarios have achieved digital office work. The digital application of government affairs systems is one such scenario.
[0003] In the prior art, the digital transformation of government affairs systems has provided many conveniences for staff, reduced work contents such as printing and signing of paper documents, and significantly improved work efficiency. However, at the same time, many problems have also arisen. For example, with the improvement of office efficiency, the update frequency and quantity of policy documents have both increased significantly. However, policy interpretation relies on manual judgment, which has problems such as strong subjectivity, low efficiency, and high omission detection rate, and the heterogeneity of government affairs data makes it difficult to quantify policy implementation indicators. Summary of the Invention
[0004] The present disclosure provides a method, device, medium and product for file analysis.
[0005] According to a first aspect of the present disclosure, a method for file analysis is provided. The method specifically includes: obtaining policy documents belonging to different effectiveness levels or different administrative departments; wherein, the policy documents include a first document and a second document; parsing the policy documents to obtain a clause tree; wherein, the nodes in the clause tree contain clause statements in the policy documents; the clause tree includes a first tree corresponding to the first document and a second tree corresponding to the second document; performing vectorization processing on the clause statements in the clause tree to obtain a first semantic vector corresponding to the first document and a second semantic vector corresponding to the second document; using the first tree and the first semantic vector, the second tree and the second semantic vector to calculate the policy document analysis result.
[0006] Based on the above content, it can be seen that when performing file analysis, due to the different levels, administrative departments, and effectiveness levels of policy document issuance. Therefore, when formulating policy documents, it is necessary to fully consider whether the policy document conflicts or contradicts other policies. Since policy documents often have a large amount of content and complex clause levels, in order to facilitate the analysis of policy documents, the policy documents can be converted into a clause tree, and then the clause tree can be converted into a semantic vector. Furthermore, the first document is analyzed using the clause tree and the semantic vector to obtain an analysis result. This analysis result is used to represent the degree of conflict between the first document and the second document. Through the above method, it is possible to comprehensively and detailedly analyze whether the first document has an obvious conflict with other second documents. And this analysis process is a comprehensive analysis based on clause statements and the context relationship of clause statements, rather than just analyzing through the similarity of words and phrases. Therefore, the analysis result is more comprehensive and accurate.
[0007] In at least one embodiment of the present disclosure, it further includes: according to the analysis result, the result identifier indicating the conflict degree is marked in different colors in the nodes of the first tree.
[0008] Based on the above, it can be seen that this visualization identification method based on the clause tree can not only clearly reflect the logical relationship within the policy, but also help decision-makers efficiently identify and resolve policy conflicts across levels and departments, improving the consistency and transparency of policy implementation. At the same time, it enables users to intuitively understand which clauses in the current clause are in conflict through the clause tree, effectively improving the inspection efficiency of policy documents.
[0009] In at least one embodiment of the present disclosure, the policy document is parsed to determine the format structure and / or semantic structure included in the policy document; the clause statements used as nodes in the clause tree are obtained by splitting based on the format structure and / or semantic structure; if no format structure or semantic structure is found, the policy document is segmented to obtain the clause statements used as nodes in the clause tree; the clause tree is constructed using the context relationship of the clause statements.
[0010] Based on the above, it can be seen that each node in the constructed clause tree has a clear and accurate context relationship, as well as the logical relationship implied in the context. It provides a data basis for subsequent clause analysis, thus facilitating the improvement of the accuracy of clause conflict analysis.
[0011] In at least one embodiment of the present disclosure, the clause statements in the nodes of the clause tree are obtained; the clause statements are vectorized to obtain sentence vectors with a specified dimension.
[0012] Based on the above, it can be seen that the finally generated sentence vectors will serve as the core input for subsequent modules such as correlation calculation and conflict assessment, supporting the system to achieve intelligent policy consistency analysis. And it provides a data basis for subsequent similarity calculation, conflict analysis, etc.
[0013] In at least one embodiment of the present disclosure, according to the release information of the policy document, the effective time and the validity level corresponding to each clause statement are determined; the clause statements carrying the effective time and the validity level are used as nodes to construct the clause tree.
[0014] Based on the above, it can be seen that this structured clause tree not only intuitively presents the logical framework of the policy document, but also provides an accurate data basis for subsequent semantic comparison, conflict location and dynamic monitoring, thus effectively supporting the intelligent analysis ability of the policy consistency warning system. Different from the prior art that only analyzes based on the split words and phrases, the clause tree provides more logical relationships and hidden meanings between clauses.
[0015] In at least one embodiment of the present disclosure, calculate the cosine similarity and point mutual information between a first semantic vector and a second semantic vector; calculate the correlation degree between the first semantic vector and the second semantic vector by using the cosine similarity and the point mutual information; when the correlation degree is greater than the correlation degree threshold, calculate the clause analysis result.
[0016] Based on the above, it can be seen that this process can not only effectively identify the explicit and implicit associations between policy clauses, but also significantly improve the accuracy and efficiency of conflict detection through a dual comparison mechanism of structuring and semanticization, providing a basis for policy optimization.
[0017] In at least one embodiment of the present disclosure, analyze the structural path similarity between a first node corresponding to a first semantic vector in a first tree and a second node corresponding to a second semantic vector in a second tree; when the structural path similarity is greater than a first threshold, calculate the cosine similarity and point mutual information between a first clause statement in the first node and a second clause statement in the second node; calculate the correlation degree between the first semantic vector and the second semantic vector by using the cosine similarity and the point mutual information.
[0018] Based on the above, it can be seen that it can not only effectively identify the explicit and implicit associations between policy clauses, but also significantly improve the accuracy and efficiency of conflict detection through a dual comparison mechanism of structuring and semanticization, providing a basis for policy optimization.
[0019] In at least one embodiment of the present disclosure, extract the keyword sets of a first clause statement and a second clause statement; calculate the point mutual information by using the keyword sets; calculate the cosine similarity between a first semantic vector corresponding to the first clause statement and a second semantic vector corresponding to the second clause statement.
[0020] Based on the above, it can be seen that it can not only effectively identify the explicit and implicit associations between policy clauses, but also significantly improve the accuracy and efficiency of conflict detection through multi-level semantic parsing and quantitative analysis, providing a basis for policy optimization. In addition, before calculating the conflicts of clauses, first evaluate the correlation degree of clauses by using the cosine similarity and the point mutual information, and screen out the clause statements with relatively high correlation degree for conflict analysis, so as to effectively reduce the workload of conflict analysis.
[0021] In at least one embodiment of the present disclosure, obtain a first structural path for representing the context relationship of a first node in a first tree, and a second structural path for representing the context relationship of a second node in a second tree; calculate the structural path similarity between the first structural path and the second structural path.
[0022] Based on the above, first, the structural path similarity is calculated using the clause tree, so that nodes with similar structural paths can be quickly found. Similar structural paths mean that the clauses in the nodes are more likely to have a high degree of correlation. In other words, the evaluation process of structural path similarity is simpler and more efficient. Therefore, by evaluating the structural path similarity first and then calculating the correlation degree, a large amount of correlation degree calculation can be avoided, and the efficiency of correlation degree calculation can be effectively improved.
[0023] In at least one embodiment of the present disclosure, keyword extraction is performed on the first document and the second document respectively using a specified extraction method; wherein, the specified extraction methods include: TF-IDF, BERT attention weight analysis, entity / term extraction; and a keyword set is constructed using the extracted keywords.
[0024] Based on the above, it can be seen that these keyword sets not only provide basic data for subsequent Pointwise Mutual Information (PMI) calculations, but also help the system more accurately quantify the semantic correlation degree between policy clauses, thus significantly improving the accuracy and efficiency of conflict detection.
[0025] In at least one embodiment of the present disclosure, the time difference of the effective time, the effectiveness level, and the correlation degree are determined; and a clause analysis result is calculated using the time difference, the effectiveness level, and the correlation degree.
[0026] Based on the above, it can be seen that the finally generated clause analysis result can not only clearly reflect the severity of the conflict, but also provide a scientific basis for subsequent risk assessment and visual warning, helping decision-makers quickly locate problems and take optimization measures.
[0027] In at least one embodiment of the present disclosure, an adaptive update model is constructed; and a model weight vector after fusing new and old policy knowledge is calculated using the cross-entropy loss corresponding to the obtained new policy document, the KL divergence of the historical policy document, the model parameters, and the penalty coefficient of the KL divergence.
[0028] Based on the above, it can be seen that this mechanism not only improves the response speed of the system to policy changes, but also provides more accurate and reliable support for policy conflict detection, thus significantly enhancing the dynamic adaptability and intelligent level of the system.
[0029] According to a second aspect of the present disclosure, there is provided an electronic device, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, such that the processor executes the method according to the first aspect of any one of the embodiments of the present disclosure.
[0030] According to a third aspect of the present disclosure, there is provided a readable storage medium storing execution instructions, which when executed by a processor are used to implement the method described in the first aspect of any one of the embodiments of the present disclosure.
[0031] According to a fourth aspect of the present disclosure, there is provided a computer program product including a computer program, which when executed by a processor implements the method described in the first aspect of any one of the embodiments of the present disclosure. Description of the Drawings
[0032] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.
[0033] Figure 1 It is a schematic flowchart of a file analysis method provided by the present disclosure.
[0034] Figure 2 It is a schematic structural diagram of the first tree or the second tree illustrated by the present disclosure.
[0035] Figure 3 It is a schematic diagram of a clause analysis and visualization processing flow illustrated by the present disclosure.
[0036] Figure 4 It is a schematic block diagram of the structure of a file analysis device according to an embodiment of the present disclosure.
[0037] Figure 5 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0038] The following further details the present disclosure with reference to the drawings and examples. It can be understood that the specific examples described herein are only for explaining the relevant content and do not limit the present disclosure. Additionally, it should be noted that for ease of description, only parts related to the present disclosure are shown in the drawings.
[0039] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and embodiments.
[0040] In the digital application scenarios of policy documents, more and more electronic documents have achieved digital management. In the existing technologies, the digital transformation of government affairs systems has provided a lot of convenience for staff, reducing the work content such as printing and signing paper documents, and significantly improving work efficiency. However, at the same time, many problems have also arisen. For example, with the improvement of office efficiency, the update frequency and quantity of policy documents have both increased significantly. However, policy interpretation relies on manual judgment, which has problems such as strong subjectivity, low efficiency, and high omission rates. The heterogeneity of government affairs data makes it difficult to quantify policy implementation indicators.
[0041] Although artificial intelligence models have been introduced in some application scenarios to assist in the interpretation of policy documents, often a keyword matching comparison method is used. This method cannot accurately identify the implicit meaning in the clauses, nor can it accurately identify the context relationship between different clauses. Especially for policy documents at different levels of effectiveness and policy documents of different administrative departments, the descriptive terms in the documents are different, and there may be conflict problems. Therefore, there is an urgent need for a solution that can accurately analyze policy documents.
[0042] For the sake of convenience of description and to make the technical solutions of the specific implementation manners of the present disclosure easier to understand, before describing the document analysis method implemented by the present disclosure, the technical terms involved in the specific implementation manners of the present disclosure are explained as follows.
[0043] The level of effectiveness can be understood as the administrative level of issuing the policy document.
[0044] Clause tree: After parsing the policy document, the clause statements in the policy document are used as nodes in the clause tree. In other words, the policy document is represented as a knowledge graph in the form of a clause tree.
[0045] Figure 1 The flowchart of a document analysis method provided by the present disclosure is as follows. As Figure 1 shown, the method includes steps 101 to 104. Among them, this method can be executed by a document analysis device.
[0046] In step 101, policy documents belonging to different levels of effectiveness or different administrative departments are retrieved. Among them, the policy documents include a first document and a second document.
[0047] In practical applications, departments at all levels will abolish some outdated policy documents according to the need for implementation capabilities and supplement them with new policy documents or policy contents that meet current requirements. When formulating policy documents, each administrative level will do so in accordance with the latest policy directions. The policy documents formulated may be for merchants or for lower-level administrative departments. Lower-level administrative departments will then draft and issue policy documents that conform to the local actual situation based on their own interpretations and notify merchants or grass-roots staff to implement them. Generally, there are no conflicts among the policy documents issued by the same administrative department. However, there may be direct or implicit conflicts between policy documents issued by different administrative departments at the same effectiveness level or by different administrative departments at different effectiveness levels.
[0048] The first document and the second document mentioned here can be policy documents belonging to different effectiveness levels respectively. For example, the first document is a policy document with low effectiveness, and the second document is a policy document with high effectiveness. Or, the first document and the second document mentioned here are policy documents belonging to different administrative departments at the same level respectively. For example, the first document is a policy document issued by one department with low effectiveness, and the second document is a policy document issued by another department with low effectiveness.
[0049] Step 102: Parse the policy document to obtain a clause tree. The nodes in the clause tree contain the clause statements in the policy document. The clause tree includes a first tree corresponding to the first document and a second tree corresponding to the second document.
[0050] In the clause tree mentioned here, there are many clause tree nodes. And, each node contains the corresponding clause statement. In other words, in the clause tree, the clause statement is the smallest analysis unit and does not need to be further split into single words or characters. In the clause tree, the hierarchical relationship and context relationship between clause statements are expressed through the hierarchical connection relationship between nodes.
[0051] In the nodes of the clause tree, in addition to storing the parsed clause statements, the effective time and effectiveness level corresponding to the policy document are also stored, providing a basis for subsequent similarity analysis of each clause.
[0052] The process of policy document analysis will be elaborated in subsequent embodiments and will not be repeated here.
[0053] It should be noted that when converting a policy document into a clause tree, it can be converted in batches and the converted clause tree can be stored in a specified location. When needed subsequently, the clause tree can be extracted at any time.
[0054] Step 103: Vectorize the clause statements in the clause tree to obtain the first semantic vector corresponding to the first document and the second semantic vector corresponding to the second document.
[0055] As described above, after generating the clause tree, each clause statement is stored in each node of the clause tree.
[0056] For the convenience of subsequent analysis, it is necessary to vectorize the clause statements in the nodes into semantic vectors. Here, the semantic vector obtained by vectorizing the clause statements in the first document is called the first semantic vector, and the semantic vector obtained by vectorizing the clause statements in the second document is called the second semantic vector.
[0057] When performing vectorization, the clause statements in the specified nodes can be selected for vectorization according to needs. For example, the clause statements in the nodes with the same or similar structural paths are selected for vectorization. The specific vectorization process will be described in detail in the subsequent embodiments and will not be repeated here.
[0058] Step 104: Calculate the policy document analysis result by using the first tree and the first semantic vector, the second tree and the second semantic vector.
[0059] The so-called clause analysis result refers to the result of clause conflict analysis. When analyzing the clauses, it is necessary to comprehensively use the clause tree and the semantic vector for analysis. It should be noted that since some clause trees are relatively large and contain many nodes, it means that there are many clause statements to be analyzed. If analyzed one by one, it will consume a large amount of computing power and the analysis efficiency is very low.
[0060] Therefore, the present disclosure proposes that some nodes with a certain similarity can be selected from the clause tree for conflict analysis according to actual needs.
[0061] That is to say, before performing conflict analysis, first find the nodes with similarity from the first tree and the second tree. The more similar the clause statements in the nodes are, the greater the probability of conflict. Furthermore, conflict analysis is performed on the clause statements in these nodes with a relatively high similarity.
[0062] Based on the above - disclosed solution, when analyzing a document, due to the different levels of issuance, different administrative departments to which the policy documents belong, and different levels of effectiveness. Therefore, when formulating a policy document, it is necessary to fully consider whether the policy document conflicts or contradicts other policies. Since policy documents often have a large amount of content and complex clause levels, for the convenience of analyzing policy documents, the policy document can be converted into a clause tree, and then the clause tree can be converted into a semantic vector. Furthermore, the clause tree and the semantic vector are used to analyze the first document, thereby obtaining an analysis result. This analysis result is used to represent the degree of conflict between the first document and the second document. Through the above - mentioned method, it is possible to comprehensively and detailedly analyze whether the first document has an obvious conflict with other second documents. This analysis process is a comprehensive analysis based on clause statements and the context relationship of clause statements, rather than just analyzing through the similarity of words. Therefore, the analysis result is more comprehensive and accurate.
[0063] In one or more embodiments of the present disclosure, it further includes: according to the analysis result, the result identifier indicating the degree of conflict is marked in different colors in the nodes of the first tree.
[0064] According to the conflict analysis result, the system will mark the result identifier indicating the degree of conflict in different colors in the nodes of the first tree to visually present the conflict situation between policy clauses. Specifically, the system will combine the risk score generated by the dynamic conflict detection engine and color - code each clause node according to a preset warning classification rule. For example, red indicates a high - risk conflict, yellow indicates a medium - risk conflict, and green indicates no conflict or a low - risk conflict. For example, when there is a serious contradiction between a high - effectiveness policy and a low - effectiveness policy in a specific clause and the impact range is wide, the relevant clause node will be marked red, accompanied by a detailed conflict score and analysis report; for clauses with less impact or only potential risks, they may be marked yellow or green. In addition, users can click on the node to view the specific conflict reasons, related clauses, and optimization suggestions, so as to quickly locate the problem and take corresponding measures.
[0065] Based on the above - disclosed solution, it can be seen that this visualization identification method based on the clause tree can not only clearly reflect the internal logical relationship of policies, but also help decision - makers efficiently identify and solve cross - level and cross - departmental policy conflict problems, improving the consistency and transparency of policy implementation.
[0066] In one or more embodiments of the present disclosure, parsing a policy document to obtain a clause tree includes: parsing the policy document to determine the format structure and / or semantic structure included in the policy document; splitting based on the format structure and / or semantic structure to obtain clause statements to be used as nodes in the clause tree; if no format structure or semantic structure is found, segmenting the policy document to obtain clause statements to be used as nodes in the clause tree; and constructing a clause tree using the context relationship of the clause statements.
[0067] The process of parsing a policy document to generate a clause tree is one of the core links of the policy consistency warning system, and its goal is to convert unstructured policy text into structured data with clear logical relationships. First, the system deeply parses the policy document through natural language processing technology to identify the format structure and semantic structure contained therein. The format structure mainly includes numbering rules (such as "Article 1", "1.1", etc.), paragraph identifiers (such as line breaks, indents, etc.), and structural guiding words (such as "stipulate as follows", "apply to", etc.). These explicit features provide clear boundaries for clause splitting. The semantic structure further explores the implicit logic such as the master-slave relationship, causal relationship, or parallel relationship between clauses through dependency syntactic analysis and semantic role annotation. If there is a lack of obvious format or semantic structure in the policy document (such as a pure text or a PDF document converted from a scanned copy), the system will segment the document into independent semantic units based on the language model combined with punctuation and semantic segmentation algorithms. Each semantic unit serves as a node in the clause tree, and the node content includes the specific clause statement and its context information. For example, Figure 2 is a schematic structural diagram of the first tree or the second tree illustrated for the present disclosure. From Figure 2 it can be seen that when constructing the clause tree (which can be the first tree corresponding to the first document, or the second tree corresponding to the second document), the system uses the context relationship of the clause statements to layer-by-layer associate the main clause with the sub-clauses, forming a hierarchical structure of "main clause - sub-clause - sub-clause", and meta-information such as the effective time and the level of effectiveness. For example, "AAA will complete BBB in 2025" may be split into the main clause "AAA needs to perform BBB" and the sub-clauses "apply to AAA" and "the deadline is 2025", and further organized into a tree structure. The finally generated clause tree can not only intuitively reflect the logical framework of the policy text, but also provide an accurate data basis for subsequent semantic comparison and conflict detection.
[0068] Based on the above disclosed solution, each node in the constructed clause tree has a clear and accurate context relationship, as well as the logical relationship implicit in the context. It provides a data basis for subsequent clause analysis, thereby facilitating the improvement of the accuracy of clause conflict analysis.
[0069] In one or more embodiments of the present disclosure, the clause sentences in the clause tree are vectorized, specifically including: obtaining the clause sentences in the nodes of the clause tree; and vectorizing the clause sentences to obtain sentence vectors with specified dimensions.
[0070] In practical applications, vectorizing clause sentences in the clause tree is a key step to achieve deep semantic understanding in the policy consistency early warning system. Specifically, each clause sentence is first extracted from the node of the clause tree. These sentences may be main clauses, subclauses, or further refined semantic units. The extracted clause sentences will be input into a deep learning-based language model (such as Legal-BERT) for vectorization. In this process, the model performs context-aware semantic parsing of clause sentences through a multi-layer Transformer encoder and generates sentence vectors with specified dimensions (such as 768 dimensions). These vectors not only capture the explicit semantic information of the clause sentences, but also mine deep semantic features such as implicit conditions, exception clauses, and complex logical relationships through the attention mechanism. For example, after vectorization, the two sentences "AAA will complete BBB by 2025" and "CCC can continue DDD" can be represented as high-dimensional vectors respectively, thus providing an accurate data basis for further semantic comparison and conflict detection. In addition, in order to improve the efficiency and accuracy of vectorization processing, the system can also fine-tune the model in combination with domain-specific knowledge bases to ensure its applicability and robustness in the field of policy texts.
[0071] Based on the above content, we can know that the sentence vector finally generated will serve as the core input of subsequent modules such as relevance calculation and conflict assessment, supporting the system to realize intelligent policy consistency analysis. It also provides a data basis for subsequent similarity calculation and conflict analysis.
[0072] In one or more embodiments of the present disclosure, a clause tree is constructed using the contextual relationship of clause statements, including: determining the effective time and effectiveness level corresponding to each clause statement according to the release information of the policy document; and constructing a clause tree using clause statements carrying the effective time and effectiveness level as nodes.
[0073] The process of constructing a clause tree using the context of clause statements is a key step in hierarchically organizing the clause statements in a policy document according to their logical structure. First, the system extracts metadata such as the effective time and the level of effectiveness of each clause statement based on the release information of the policy document (such as the title, preface, appendix, etc.). For example, high-efficiency policies usually have a higher level of effectiveness, while low-efficiency policies are relatively lower; the effective time of a clause may be clearly marked in the text (such as "effective as of January 1, 2025"), or it may need to be inferred from the release time. The combination of this metadata and the semantic content of the clause statements provides an important basis for subsequent conflict detection and risk assessment. Next, the system uses the clause statements with metadata such as the effective time and the level of effectiveness as nodes to construct a clause tree based on their context. Specifically, the main clause serves as the root node, and the sub-clauses and supplementary explanations serve as sub-nodes, forming a hierarchical structure of "policy - clause - sub-clause". For example, "AAA shall complete BBB by 2025" as the main clause can be split into sub-clauses such as "applicable to AAA" and "the deadline is 2025", and the corresponding level of effectiveness and effective time are marked. In addition, the construction of the clause tree also combines dependency syntactic analysis technology to ensure that the logical relationships between nodes accurately reflect the actual intention of the policy text.
[0074] Based on the above public content, it can be seen that this structured clause tree not only intuitively presents the logical framework of the policy document, but also provides an accurate data basis for subsequent semantic comparison, conflict location, and dynamic monitoring, thus effectively supporting the intelligent analysis ability of the policy consistency warning system. Different from the prior art that only analyzes based on the words and phrases obtained by splitting, the clause tree provides more logical relationships between clauses and hidden meanings.
[0075] In one or more embodiments of the present disclosure, using the first tree and the first semantic vector, and the second tree and the second semantic vector, a clause analysis result is calculated, specifically including: calculating the cosine similarity and point mutual information between the first semantic vector and the second semantic vector; calculating the correlation degree between the first semantic vector and the second semantic vector using the cosine similarity and point mutual information; when the correlation degree is greater than the correlation degree threshold, the clause analysis result is calculated.
[0076] The process of calculating the policy document analysis result using the first tree and the first semantic vector, and the second tree and the second semantic vector is a key step in achieving accurate conflict detection in the policy consistency warning system. First, the system extracts the semantic vectors corresponding to the nodes in the first tree and the second tree respectively, including the first semantic vector and the second semantic vector. The semantic vector can have a specified dimension, such as a 768-dimensional vector generated based on Legal-BERT. And calculate the cosine similarity and point mutual information between the first semantic vector and the second semantic vector. The cosine similarity is used to measure the degree of direction proximity of two semantic vectors in the high-dimensional space, reflecting the correlation of the overall semantics of the clauses; while the point mutual information captures the potential connection between them in the specific expression by quantifying the co-occurrence probability of the keyword sets of the two clauses.
[0077] For example, for "Policy A requires AAA to complete BBB" and "Policy B provides DDD for AAA", the system will extract the keyword sets (such as "AAA", "BBB", "DDD") and calculate their point mutual information values to evaluate whether there are implicit conflicts between these keywords at the semantic level.
[0078] Next, the system combines the cosine similarity and the point mutual information value, and calculates the comprehensive correlation degree between the first semantic vector and the second semantic vector through a weighted formula, and the weight coefficient can be dynamically adjusted according to actual needs. When the calculated correlation degree exceeds the preset correlation degree threshold, the system determines that there is a potential conflict or overlap relationship between these two policy clauses, and further triggers the conflict evaluation model to generate a detailed clause analysis result.
[0079] Based on the above disclosed solution, this process can not only effectively identify the explicit and implicit correlations between policy clauses, but also significantly improve the accuracy and efficiency of conflict detection through a dual comparison mechanism of structuring and semanticization, providing a basis for policy optimization.
[0080] In one or more embodiments of the present disclosure, calculating the correlation degree between the first semantic vector and the second semantic vector using the cosine similarity and the point mutual information includes: analyzing the structural path similarity between the first node corresponding to the first semantic vector in the first tree and the second node corresponding to the second semantic vector in the second tree; when the structural path similarity is greater than the first threshold, calculating the cosine similarity and the point mutual information between the first clause statement in the first node and the second clause statement in the second node; calculating the correlation degree between the first semantic vector and the second semantic vector using the cosine similarity and the point mutual information.
[0081] Calculating the correlation degree between the first semantic vector and the second semantic vector using the cosine similarity and the point mutual information is a key step in achieving accurate semantic comparison in the policy consistency warning system.
[0082] The system analyzes the structural path similarity between the first node corresponding to the first semantic vector in the first tree and the second node corresponding to the second semantic vector in the second tree. The structural path similarity is used to measure whether the hierarchical relationships and logical contexts of two nodes in the clause tree are similar. For example, whether both nodes belong to sub-clauses under specific topic paths such as "AAA" or "DDD".
[0083] If the structural path similarity exceeds a preset first threshold, the cosine similarity and point mutual information between the first clause statement in the first node and the second clause statement in the second node are further calculated. The cosine similarity reflects the correlation degree of the overall semantics of the clauses by comparing the proximity of the directions of two semantic vectors; while the point mutual information captures the potential connection between them in the specific expressions by quantifying the co-occurrence probability of the keyword sets of the two clauses. For example, "AAA" and "BBB" may frequently co-occur in multiple clauses, and their point mutual information value is relatively high, indicating a strong semantic association between the two.
[0084] Furthermore, by combining the cosine similarity and the point mutual information, the comprehensive correlation degree between the first semantic vector and the second semantic vector is calculated through a weighted formula, and the weight coefficients can be dynamically adjusted according to actual needs.
[0085] Based on the above disclosed solution, it can not only effectively identify the explicit and implicit associations between policy clauses, but also significantly improve the accuracy and efficiency of conflict detection through a dual comparison mechanism of structuring and semanticizing, providing a basis for policy optimization.
[0086] In one or more embodiments of the present disclosure, calculating the cosine similarity and point mutual information between the first clause statement in the first node and the second clause statement in the second node specifically includes: extracting the keyword sets of the first clause statement and the second clause statement; calculating the point mutual information using the keyword sets; calculating the cosine similarity between the first semantic vector corresponding to the first clause statement and the second semantic vector corresponding to the second clause statement.
[0087] The process of calculating the cosine similarity and point mutual information between the first clause statement in the first node and the second clause statement in the second node is a key link in realizing accurate semantic comparison in the policy consistency warning system.
[0088] Extract the keyword sets of the first clause statement and the second clause statement. These keyword sets are generated in various ways: use the TF-IDF algorithm to screen out high-frequency and discriminative words, combine the attention weight analysis of the Legal-BERT model to capture semantic core words, and at the same time use the government affairs domain knowledge base for entity and term extraction. For example, "AAA will complete BBB before 2025" may extract keywords such as "AAA", "BBB", "2025", etc., while "CCC can continue DDD" may extract keywords such as "CCC", "DDD", etc.
[0089] Based on these keyword sets, the system calculates the pointwise mutual information to quantify the co-occurrence probability of keywords in the two clause statements, so as to evaluate their potential relationship in the specific expression. For example, if "AAA" appears frequently in the two clause statements, its pointwise mutual information value is higher, indicating that this keyword has a strong semantic correlation between the two clauses.
[0090] Next, calculate the cosine similarity between the first semantic vector corresponding to the first clause statement and the second semantic vector corresponding to the second clause statement. By comparing the degree of closeness of the directions of the two vectors, it reflects the correlation degree of the overall semantics of the clauses.
[0091] Furthermore, the results of the cosine similarity and the pointwise mutual information are weighted and fused to form a comprehensive correlation score, which is used to judge whether there is a potential conflict or synergy relationship between the two clauses.
[0092] Based on the above public solution, it is not only possible to effectively identify the explicit and implicit relationships between policy clauses, but also to significantly improve the accuracy and efficiency of conflict detection through multi-level semantic parsing and quantitative analysis, providing a basis for policy optimization. In addition, before calculating the conflict of the clauses, first use the cosine similarity and the pointwise mutual information to evaluate the correlation degree of the clauses, screen out the clause statements with relatively high correlation degree and then conduct conflict analysis, so as to effectively reduce the workload of conflict analysis.
[0093] In one or more embodiments of the present disclosure, analyze the structural path similarity between the first node in the first tree and the second node in the second tree, which specifically includes: obtaining the first structural path used to represent the context relationship of the first node in the first tree, and the second structural path used to represent the context relationship of the second node in the second tree; calculating the structural path similarity between the first structural path and the second structural path.
[0094] Analyzing the structural path similarity between the first node in the first tree and the second node in the second tree is an important part of realizing accurate conflict detection in the policy consistency warning system.
[0095] Obtain the first structural path of the first node in the first tree and the second structural path of the second node in the second tree respectively. These structural paths represent the context logical relationship of the nodes through the hierarchical relationship of "policy - clause - sub - clause". For example, "BBB>AAA>Deadline" is a typical structural path.
[0096] Specifically, the extraction of the first structural path and the second structural path depends on the hierarchical organization of the clause tree. The path of each node consists of its complete hierarchical identifier from the root node to the current node, ensuring that it can accurately reflect the position and semantic constraint conditions of the node in the entire policy text.
[0097] By calculating the similarity between the two structural paths, the logical correlation degree of the two nodes in the policy text is evaluated. This similarity calculation usually adopts methods based on tree edit distance or path matching algorithms, comprehensively considering factors such as path length, node label similarity, and hierarchical overlap degree. For example, if both paths involve clauses related to "AAA" and have a similar logical structure at the hierarchical level, their structural path similarity is relatively high; conversely, if one path involves "BBB" and the other path involves "DDD", the similarity is relatively low.
[0098] Based on the above - disclosed solution, first use the clause tree to calculate the structural path similarity, so that nodes with similar structural paths can be quickly found. Similar structural paths mean that the clauses in the nodes are more likely to have a high degree of correlation. In other words, the evaluation process of structural path similarity is simpler and more efficient. Therefore, by first evaluating the structural path similarity and then calculating the correlation degree, a large amount of calculation of the correlation degree can be avoided, effectively improving the calculation efficiency of the correlation degree.
[0099] In one or more embodiments of the present disclosure, extract the keyword sets of the first clause statement and the second clause statement, which specifically includes: using the specified extraction methods to extract keywords from the first file and the second file respectively; where the specified extraction methods include: TF - IDF, BERT attention weight analysis, entity / term extraction; and constructing keyword sets using the extracted keywords.
[0100] Extracting the keyword sets of the first clause statement and the second clause statement is an important step in achieving accurate semantic comparison in the policy consistency warning system. Specifically, multiple specified extraction methods can be used to extract keywords from the clause statements in the first file and the second file respectively.
[0101] For example, high-frequency and discriminative words are selected through the TF-IDF algorithm, which can reflect the core themes and key contents of policy texts. It is also possible to capture words or phrases with high semantic weights in clause statements based on the attention weight analysis technology of the Legal-BERT model. These words are usually the key points affecting semantic understanding, such as "AAA", "BBB", etc. In addition, entity and term extraction methods in the government affairs field can be combined to identify professional terms and named entities in specific fields from policy texts, such as "DDD", "CCC", etc.
[0102] These extraction methods complement each other, ensuring the comprehensiveness and accuracy of the keyword set. After extracting the keywords, the system further integrates these keywords into a standardized keyword set and constructs the final keyword set through processing steps such as deduplication and normalization. For example, "AAA will complete BBB before 2025" may generate the keyword set {"AAA", "BBB", "2025"}, while "CCC can continue DDD" may generate the keyword set {"CCC", "DDD"}.
[0103] Based on the above public solution, these keyword sets not only provide basic data for subsequent pointwise mutual information (PMI) calculations but also help the system more accurately quantify the semantic correlation degree between policy clauses, thus significantly improving the accuracy and efficiency of conflict detection.
[0104] In one or more embodiments of the present disclosure, when the correlation degree is greater than the correlation degree threshold, the clause analysis result is calculated, specifically including: determining the time difference of the effective time, the effectiveness level, and the correlation degree; calculating the clause analysis result using the time difference, the effectiveness level, and the correlation degree.
[0105] When the correlation degree is greater than the preset correlation degree threshold, the system will further calculate the policy document analysis result to quantify the potential conflicts or synergistic relationships between policy clauses.
[0106] Specifically, first, determine the time difference in the effective time between the two clauses, their respective effectiveness levels, and the correlation degree calculated through cosine similarity and pointwise mutual information. These key dimensions together constitute the core input data for conflict analysis. For example, the time difference reflects the time sensitivity of the conflict, and recently issued policy clauses usually have a higher timeliness weight; the effectiveness level reflects the authority and priority of the policy, and higher-level policies may have a binding or covering effect on lower-level policies; the correlation degree measures the similarity or contradiction degree between clauses at the semantic level.
[0107] Next, use the comprehensive scoring formula for these dimensions to calculate the analysis results of policy documents. The formula usually includes weighting coefficients (γ1, γ2, γ3) to dynamically adjust the influence ratio of each dimension on the final result. For example, the conflict timeliness score (T) is calculated based on a time decay function, and the impact of conflicts closer to the current time is greater; the effectiveness level score is weighted and calculated according to policy level differences to ensure that conflicts in high-priority policies can be identified first; the relevance score further combines the context path of the clauses to determine whether there are logical contradictions or goal conflicts.
[0108] Based on the above public solution, it can be seen that the finally generated clause analysis results can not only clearly reflect the severity of conflicts, but also provide a scientific basis for subsequent risk assessment and visual warning, helping decision-makers quickly locate problems and take optimization measures.
[0109] In one or more embodiments of the present disclosure, it further includes: constructing an adaptive update model; using the cross-entropy loss corresponding to the newly obtained policy document, the KL divergence of the historical policy document, model parameters, and the penalty coefficient of the KL divergence , and calculating the model weight vector after fusing new and old policy knowledge.
[0110] Constructing an adaptive update model is the core link for realizing dynamic learning and knowledge fusion in the policy consistency warning system. Its goal is to ensure that the system can quickly adapt to the release of new policies while avoiding forgetting historical policy knowledge. Specifically, when a new policy document is released, the system first calculates the cross-entropy loss corresponding to the new policy data ( L new ), which is used to measure the prediction error of the model on new policy samples and reflects the model's adaptability to new knowledge; at the same time, the system calculates the KL divergence of the historical policy document ( D KL ), quantifies the output distribution difference between the new and old models on historical data, and prevents "catastrophic forgetting" of old policy knowledge due to overfitting to new policies. On this basis, combined with the current model parameters, the penalty coefficient of the KL divergence and the dynamic learning rate η , the model weight vector after fusing new and old policy knowledge ( θ new ) is calculated through a weighted formula. For example, when a new high-effectiveness policy is released, the system will adjust the model parameters based on the conflict detection results between the policy clause and low-effectiveness policies to adapt to this change, and at the same time, use the KL divergence term to constrain the model's understanding ability of existing policies to ensure its stability and robustness in long-term operation.
[0111] Based on the above - disclosed solution, this mechanism not only improves the system's response speed to policy changes but also provides more accurate and reliable support for policy conflict detection, thus significantly enhancing the system's dynamic adaptability and intelligence level.
[0112] For ease of understanding, the following will be described in conjunction with specific embodiments. As Figure 3 This is a schematic diagram of the clause analysis and visualization processing flow illustrated in this disclosure.
[0113] As in step 301, first, data collection is carried out and a clause tree is established. After receiving a policy document, data collection is performed on the policy document. This is because the forms of policy documents issued by different departments and regions are inconsistent. Therefore, data collection needs to be carried out for different data structure types.
[0114] Common data structure types include: (1) Structured data: policies and regulations in the government affairs public database (formats such as XML, JSON). (2) Semi - structured data: implementation rules (formats such as HTML tables, Excel). (3) Unstructured data: policy interpretation documents and meeting minutes (formats such as PDF, Word).
[0115] For example: Assume that Policy A is the first document mentioned above and Policy B is the second document mentioned above. The system will perform deduplication, cleaning, and format conversion on this data to ensure data accuracy and consistency.
[0116] After completing data collection, in - depth semantic analysis of the collected data is required. The so - called in - depth semantic analysis is because many policy clause conflicts are not direct conflicts (not literal conflicts of a certain word or term), but conflicts hidden in semantic details.
[0117] Parsing the clauses in a policy document can be understood as splitting the clause statements in the policy document. Processing is carried out at the sentence level and the paragraph level respectively to ensure parsing accuracy.
[0118] Combining the formal features and semantic logical structures of the policy document at the language level and format level, long - text clauses are reasonably divided into sentences or paragraphs with independent semantic units.
[0119] 1. Format structure (explicit structure) recognition: Using the document structure parsing algorithm in natural language processing, combined with regular rules, title detectors, and syntactic chunkers, the following structural clues are recognized.
[0120] Numbering rules: Clause numbers such as "Article 1", "1.1", "I.", "(I)", "①", etc.
[0121] Paragraph identifiers: Typical format symbols such as line breaks, multiple spaces, indents, etc.
[0122] Structural guiding words: such as "stipulated as follows", "apply to", "not include", "proviso clause", etc.
[0123] Title chunks: such as "scope of application", "penalty clause", "incentive mechanism", etc., used as logical breakpoints.
[0124] The system will first disassemble into hierarchical paragraph units according to the numbering structure (such as "1.2", "1.3.1", etc.). For text where the numbering cannot be recognized, semantic segmentation will be performed based on punctuation marks such as full stops and semicolons in combination with the language model.
[0125] 2. Semantic structure (implicit logic) recognition: On the basis of format parsing, use the fine-tuned Legal-BERT model to perform syntactic dependency analysis and semantic role annotation to judge whether there are the following semantic relationships between clauses.
[0126] Main-clause relationship: The semantic dependence between the main clause and supplementary explanations, conditional restrictions.
[0127] Causality and turning: For example, "if..., then...", "but if... then exception", etc.
[0128] Parallel / contrast: Such as the parallel logic of "enterprises need to meet condition A or condition B".
[0129] These semantic clues are used to further refine the segmentation at the sentence level to ensure that the Legal-BERT model can reasonably disassemble a compound sentence into multiple "clause statements" with independent meanings.
[0130] 3. Paragraph processing method. If the policy clause is in a structural specification (such as the XML / JSON format of a regulatory document), directly use structural tags <clause><sub-clause>, etc. for positioning; if it is pure text or text converted from PDF, perform paragraph segmentation through the structure recognition module (combined with the rule template and the output of the language model).
[0131] For example, a policy paragraph will be disassembled into three semantic units by the system. Each semantic unit will be converted into a vector input for subsequent clause tree construction and semantic comparison.
[0132] After parsing the clauses through the method described above, further, a clause tree can be constructed. According to the structure of the policy text, generate a tree structure with the levels of "policy - clause - sub-clause", and information such as the effective time and the level of validity.
[0133] It should be noted that the reason for converting policy documents into clause trees here is to facilitate subsequent analysis of clauses. Because the "clause tree" is not only a formal representation of the text structure, but also a key data structure for policy semantic comparison, conflict location, and context alignment.
[0134] The clause tree can assist in conflict identification in multiple dimensions. The clause tree reflects the logical dependency relationship between "parent and child clauses", which can ensure that during the comparison process, sentences are not compared in isolation, but the limiting conditions or scope of application of their upper and lower level clauses are considered, thereby avoiding semantic misjudgment. Example: Two sub-clauses may seem contradictory, but if their upper-level main clauses limit different applicable objects, they actually do not constitute a conflict.
[0135] The clause tree also defines the comparison range and path mapping relationship. The clause tree provides a path positioning mechanism for the comparison algorithm: the comparison occurs not only between leaf nodes, but along the path of "policy - clause - sub-clause" for structured alignment, making the comparison more targeted and improving the calculation efficiency.
[0136] The structural hierarchy of the clause tree supports similarity analysis. Providing a structural path in the clause tree can be used as a factor for weighted calculation of semantic similarity. The closer the clauses on the structural path are, the higher the weighted probability of their semantic conflict.
[0137] In addition, in knowledge graph visualization, the nodes of the clause tree are directly mapped to the nodes in the graph, and each node corresponds to a clause text; the conflict detection model writes back the high-risk comparison results to the clause tree nodes, marks colors (red, yellow, green), and outputs them as a structured risk report.
[0138] After parsing the clause statements, it is further necessary to perform semantic vectorization on the clause statements: after each clause is input into Legal-BERT, it is converted into a 768-dimensional vector representation, and the comprehensive semantic vector at the clause level is calculated through the attention mechanism. 。
[0139] Among them is the i-th clause, represents a 768-dimensional vector. For example, Policy A vector: (0.32, -0.45, 0.87, … a total of 768 values); Policy B vector: (0.28, -0.40, 0.83, … a total of 768 values). Traditional methods only check whether "BBB" and "DDD" appear at the same time, and may miss other implicit conflicts. The proposed solution in this disclosure can truly understand the meaning of policy clauses and determine whether these clause statements are contradictory in terms of their objectives.
[0140] As in step 302, conflict detection is performed on the clauses. After completing the vectorization process of the clause statements, the correlation degree between the clause statements in the two clause trees can be further evaluated. When calculating the correlation degree, cosine similarity and point mutual information are used to quantify the semantic correlation degree between different clauses 。 The calculation formula is as follows.
[0141] Among them, 、 are the semantic vectors of clauses i and j, and the cosine similarity reflects the semantic correlation of clauses, 、 are the sets of clause keywords, α、β is the weight coefficient, and the point mutual information quantifies the co-occurrence probability of clauses 。
[0142] It should be noted that clause i and clause j respectively refer to the node pairs in the clause trees of two different policies, including but not limited to the following types: main clause i to main clause j, main clause i to sub-clause j, sub-clause i to sub-clause j, sub-clause i to main clause j.
[0143] The cosine similarity mentioned here can be understood as measuring the overall similarity of semantic vectors; point mutual information: measuring the co-occurrence probability of keyword sets.
[0144] Since some policy documents have a large amount of content, the number of nodes in the generated clause trees is also large. If each node in the clause tree is analyzed, it will consume a large amount of computing power and time. Therefore, non-full calculation can be adopted. There are three ways to choose.
[0145] The first is structural path matching selection: only compare clause pairs with similar structural paths.
[0146] The second is keyword semantic cluster screening: after clustering the clauses through keyword clustering (or topic models such as LDA), only compare the clauses in the same cluster.
[0147] The third is semantic vector rough screening: first use fast dimensionality reduction (such as PCA or quantization index) to roughly screen out the top K clauses that may be relevant in the low-dimensional space, and then perform high-precision comparison.
[0148] After calculating the correlation degree between clauses through the above methods, conflict detection can be further performed. Detect whether there are conflicts between policies at different levels and from different sources, and measure the severity of the conflicts. In other words, the system will obtain the correlation degree of each clause pair after calculating the semantic correlation degree (i.e., the above cosine similarity + point mutual information) . Only when exceeds the first threshold (such as 0.65), the system enters the conflict scoring model (reinforcement learning + reward function) stage. This is equivalent to predicting that only those "possibly in conflict" will continue with conflict analysis, and clauses pairs that are not relevant (i.e., the degree of association is less than the first threshold) are directly skipped, without the risk of conflict.
[0149] In the present disclosure solution, a reinforcement learning system is constructed to comprehensively evaluate the conflict situation between clauses. The reinforcement learning system takes the conflict state as input and constructs a complete policy network structure, including an action selection function, a reward function design, and a policy optimization mechanism.
[0150] Among them, the conflict state is used to represent the severity of the conflict between two clauses. Here, combining the policy effectiveness level and the release time factor, a reinforcement learning system is established. In the reinforcement learning system, the state space S t is used to describe the policy conflict situation at the current moment t, and the state space is defined as follows: Among them, Level represents the policy effectiveness level (high effectiveness = 3, medium effectiveness = 2, low effectiveness = 1), which is used to measure the scope of the conflict. Δt represents the release time difference of the policy, that is, the time interval between the new and old policies. represents the degree of association between clauses, calculates the conflict possibility, and the range is 0 - 1. The closer to 1, the greater the conflict possibility.
[0151] This state space provides input for conflict detection, and the conflict state will be input into the reinforcement learning system to calculate the severity of the conflict in each state.
[0152] To obtain a more comprehensive conflict assessment result, in addition to calculating the conflict state, the timeliness of the conflict and the degree of socio-economic impact are also calculated.
[0153] Step 303, after completing conflict detection, the conflict analysis result can be calculated. In the present disclosure solution, a scheme for dynamically evaluating the conflict risk using a reward function is proposed. Based on the state space , a quantitative index is also needed to measure the severity of the conflict. Define the reward function for conflict detection , covering the following three dimensions.
[0154] Among them, T represents the conflict timeliness score (range 0 - 1, the closer to the current time, the greater the conflict impact), T The value is calculated based on the release time difference of the conflict clauses Δt, obtained through an exponential decay function , can effectively reflect the time sensitivity of conflicts and ensure that conflict policies that come into effect or are revised recently receive a higher early warning weight. Among them, Δt = ∣t now − t policy ∣: represents the interval between the current time and the time when the clause comes into effect (the unit can be days or months). λ is the time decay coefficient. The larger λ is, the faster the decay caused by the time effect. T ∈ (0, 1]: The closer it is to 1, the closer the policy conflict occurs and the more worthy of attention.
[0155] S (that is, the state space S calculated previously t ) represents the severity of the clause conflict (which can be weighted and calculated according to the policy effectiveness level). The calculation process will not be repeated here. For details, please refer to the embodiments described above. L represents the predicted value of the socio-economic impact, and γ1, γ2, γ3 are dynamic adjustment coefficients that control the impact of different factors on the final reward.
[0156] L is based on the LSTM model to model the trend of the possible socio-economic impact after the policy conflict. The model integrates multi-source information such as the semantic content of the policy, historical conflict cases, enterprise portraits, and macroeconomic data, and outputs a normalized score L ∈ [0, 1] to measure the potential destructiveness of the conflict. The core mechanism of L: a multi-factor time series prediction model based on LSTM. Use the long short-term memory network (LSTM) to model the following multi-source data to predict the trend of the socio-economic impact after the conflict event. Finally, a normalized score L ∈ [0, 1] is output.
[0157] The input data dimensions in the LSTM model for calculating L include: the semantic vector of the clause content (obtained from Legal-BERT, encoding the conflict type and field); the feature vector of historical conflict events; macroeconomic indicators; industry sensitivity weights; enterprise portraits, etc. What the LSTM model outputs is the "degree of economic disturbance caused by the predicted conflict" (such as the impact range of GDP, the number of affected enterprises, etc.), which is mapped to the interval [0, 1] through a standardization function: . Among them is the predicted output of LSTM, and Normalize is the normalization function.
[0158] It should be noted that the above LSMT model is trained using historical data as training samples. The sources of historical data include: the historical policy conflict case library (including subsequent economic impact indicators); statistical yearbooks, public economic data; auxiliary information such as credit ratings, enterprise announcements, and industry analysis reports.
[0159] The reward function takes the state space S of the policy conflict tQuantified into a score value, which is used to measure whether the currently detected conflict is serious, moderate or minor. This score value will affect the subsequent adaptive learning process and help the system continuously optimize the decision-making strategy for conflict detection.
[0160] In practical applications, the situation of policy update and release also needs to be considered. When a new policy is released, the present disclosure uses an adaptive learning rate for model update: 。
[0161] Among them, represents the KL divergence of historical policy data, which measures the difference between the new and old policies and prevents the model from forgetting the old policies.
[0162] η represents the dynamic learning rate, which controls the model update amplitude and is adaptively adjusted according to the policy change frequency. The more frequent the policy change, the higher the learning rate, so as to ensure that the system adapts to the new policy. The incremental learning mechanism enables the system to continuously optimize its policy conflict detection ability, ensuring that after the new policy is released, the system can quickly adapt without relying on outdated detection rules. As the penalty coefficient of the KL divergence, its value range is usually 0.01 - 0.1, the larger it is, the less the old policies are forgotten during incremental learning.
[0163] Parameter update result θ new represents the model weight vector after fusing the knowledge of the new and old policies and serves as the current effective parameter of the conflict detection policy network.
[0164] Cross-entropy loss L new Calculated based on the classification prediction error of the model on the new policy samples, which measures the adaptability of the model to the new policy.
[0165] The initial model weight is θ old , which represents the parameters obtained by training under the historical policy. Each time the model receives new policy data, it will perform a "weight update with memory". The finally output θ new is the optimal weight that fuses the knowledge of the new and old policies and can adapt to the current environment.
[0166] L new Calculation method 。Among them: CrossEntropy represents the cross-entropy; y pred : The conflict label predicted by the model based on the current new policy data (such as whether there is a conflict, conflict level); y true : The true conflict result labeled by the new policy data (from manual annotation or rule-guided label).
[0167] KL divergence term D KL (L old ) It represents the output deviation between the new and old models on historical data, which is used to control "catastrophic forgetting", ensure that the model retains the memory ability of existing policy knowledge during update, and thus achieve a stable learning process of dynamic adaptation.
[0168] KL divergence calculation method 。
[0169] P(x) is the output distribution of the historical model weights ( θ old ) on the old policy data; Q(x) is the output distribution of the parameter update result of the current model ( θ new ) on the same old data. The purpose is to measure the output difference between the new and old models on old knowledge. The larger it is, the more the new model "forgets" old knowledge.
[0170] The KL divergence is added to the update formula as a "forgetting penalty term"; if the output of the new model on old data deviates too much, a large penalty will be generated to suppress the bias adjustment of the model; the final achievable effect is to maintain the memory ability of the old policy while absorbing the new policy.
[0171] After completing the conflict assessment and obtaining the conflict analysis result, as described in step 304, the conflict analysis result can be visually displayed. That is, the conflict analysis result is represented by different colors in the clause tree. The conflict analysis result can be quantified first. The formula is as follows: 。
[0172] Among them, represents the weight of the m-th evaluation dimension; it represents the "relative importance" of each risk dimension in the final score; it is a weight coefficient that controls the scoring structure and is used to reflect the contribution degree of different conflict attributes to the overall risk. This can be set according to experience.
[0173] represents the standardized evaluation value of dimension m, which comes from the calculation result of the reward function. Each evaluation dimension has a corresponding conflict score, and these scores need to be standardized (that is, normalized to 0-1) so that different evaluation dimensions can be weighted and calculated. These scores mainly come from the severity of the conflict (S), the policy timeliness score (T), and the predicted value of the socio-economic impact (L). These values will be multiplied by the weight for weighted calculation, and finally affect the risk index; μ is the time decay coefficient, and Δt is the release time difference of the policy.
[0174] It represents the time decay factor, which is used to measure the degree of influence of policy conflicts over time. μ represents the time decay coefficient, which controls the influence speed of time on the risk index. The larger μ is, the more significant the time decay effect is. Δt represents the release time difference of policies. The larger the release time difference is, the more obvious the time decay effect is.
[0175] After calculating Risk (risk index), it can be further marked in the clause tree according to the size of Risk. Red, yellow, and green three-color warnings are set according to the risk index. It is a coloring strategy for knowledge graph nodes judged according to thresholds.
[0176] If the conflict index between Policy A (high-effectiveness policy) and Policy B (low-effectiveness policy) is higher than 0.9, the system will trigger a red alert and recommend modifying Policy B to conform to Policy A.
[0177] It should be noted that when scoring the conflict severity, it does not judge two sentences in isolation, but combines the context path of the clause tree to judge whether the following conditions are met: whether they belong to the same topic path (such as "execution mechanism", "target task"); whether there are contradictions in the upper-level main clause instructions (such as "must", "allow"); whether there is an intersection between high-effectiveness policies and low-effectiveness policies.
[0178] In the solution of the present disclosure, by combining large models, reinforcement learning, knowledge graph (i.e., clause tree) construction, incremental learning, and visual analysis, the limitations of traditional policy consistency detection methods are broken through, the accuracy, dynamic adaptability, and automation of policy comparison are improved, and it can effectively support policy-making and execution agencies to improve decision-making efficiency, reduce compliance risks, enhance the transparency of policy execution at the same time, and provide more intelligent policy warning and evaluation tools.
[0179] Based on any one of the above embodiments, the present disclosure also provides a document analysis device. Figure 4 It is a structural schematic block diagram of the document analysis device according to an embodiment of the present disclosure. As Figure 4 shown, the document analysis method device includes: an acquisition module 41, which is used to acquire policy documents belonging to different effectiveness levels or different administrative departments; wherein, the policy documents include a first document and a second document.
[0180] A parsing module 42, which is used to parse the policy documents to obtain a clause tree. Among them, the nodes in the clause tree contain clause statements in the policy documents. The clause tree includes a first tree corresponding to the first document and a second tree corresponding to the second document.
[0181] The vectorization processing module 43 is used to perform vectorization processing on the clause statements in the clause tree to obtain the first semantic vector corresponding to the first document and the second semantic vector corresponding to the second document.
[0182] The calculation module 44 is used to calculate the policy document analysis result by using the first tree and the first semantic vector, the second tree and the second semantic vector.
[0183] Optionally, it further includes a visualization module 45, which is used to mark the result identifier representing the conflict degree in different colors in the nodes of the first tree according to the analysis result.
[0184] The parsing module 42 is used to parse the policy document to determine the format structure and / or semantic structure included in the policy document; split based on the format structure and / or semantic structure to obtain clause statements used as nodes in the clause tree; if no format structure or semantic structure is found, perform segmentation processing on the policy document to obtain clause statements used as nodes in the clause tree; construct a clause tree by using the context relationship of the clause statements.
[0185] The parsing module 42 is used to obtain the clause statements in the nodes of the clause tree; perform vectorization processing on the clause statements to obtain sentence vectors with a specified dimension.
[0186] The parsing module 42 is used to determine the effective time and validity level corresponding to each clause statement according to the release information of the policy document; construct a clause tree with the clause statements carrying the effective time and validity level as nodes.
[0187] The calculation module 44 is used to calculate the cosine similarity and point mutual information between the first semantic vector and the second semantic vector; calculate the correlation degree between the first semantic vector and the second semantic vector by using the cosine similarity and point mutual information; in the case where the correlation degree is greater than the correlation degree threshold, calculate the clause analysis result.
[0188] The calculation module 44 is used to analyze the structural path similarity between the first node corresponding to the first semantic vector in the first tree and the second node corresponding to the second semantic vector in the second tree; when the structural path similarity is greater than the first threshold, calculate the cosine similarity and point mutual information between the first clause statement in the first node and the second clause statement in the second node; calculate the correlation degree between the first semantic vector and the second semantic vector by using the cosine similarity and point mutual information.
[0189] The calculation module 44 is used to extract the keyword sets of the first clause statement and the second clause statement; calculate the point mutual information by using the keyword sets; calculate the cosine similarity between the first semantic vector corresponding to the first clause statement and the second semantic vector corresponding to the second clause statement.
[0190] A computing module 44, configured to obtain a first structural path for representing a context relationship of a first node in a first tree and a second structural path for representing a context relationship of a second node in a second tree; and calculate a structural path similarity between the first structural path and the second structural path.
[0191] The computing module 44 is configured to perform keyword extraction on a first file and a second file respectively by using a specified extraction method; wherein the specified extraction methods include: TF-IDF, BERT attention weight analysis, entity / term extraction; and construct a keyword set by using the extracted keywords.
[0192] The computing module 44 is configured to determine a time difference, an effectiveness level, and a correlation degree of an effective time; and calculate a clause analysis result by using the time difference, the effectiveness level, and the correlation degree.
[0193] Optionally, it further includes an update module 46, configured to: construct an adaptive update model; and calculate a model weight vector after fusing new and old policy knowledge by using the cross-entropy loss corresponding to the obtained new policy file, the KL divergence of the historical policy file, model parameters, and the penalty coefficient of the KL divergence , and calculate a model weight vector after fusing new and old policy knowledge.
[0194] The implementation processes of the functions and roles of each module in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0195] The execution subject of the file analysis method in the specific implementation manner of the present disclosure may be an electronic device such as a server (including a local server or a cloud server).
[0196] Therefore, based on any one of the above embodiments, the present disclosure further provides an electronic device, which can execute the file analysis method of any one of the above embodiments described in the present disclosure.
[0197] Figure 5 It is a structural schematic block diagram of an electronic device according to an embodiment of the present disclosure.
[0198] The hardware structure of the electronic device 1000 can be implemented by using a bus architecture. The bus architecture may include any number of interconnected buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0199] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only one connecting line is used in this figure, but it does not mean that there is only one bus or one type of bus.
[0200] The present disclosure also provides a readable storage medium storing a computer program, which is used to implement the above method when executed by a processor. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0201] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.
[0202] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed, or a data storage device such as a server or data center integrating one or more available mediums. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; it can also be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0203] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0204] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing method devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing method devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing method devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing method devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0207] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0208] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0209] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made on the basis of the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A method for file analysis, characterized in that, The method includes: Obtaining policy documents belonging to different effectiveness levels or different administrative departments, where the policy documents include a first document and a second document; Parsing the policy documents to obtain a clause tree, where the nodes in the clause tree contain clause statements in the policy documents, and the clause tree includes a first tree corresponding to the first document and a second tree corresponding to the second document; Performing vectorization processing on the clause statements in the clause tree to obtain a first semantic vector corresponding to the first document and a second semantic vector corresponding to the second document; Using the first tree and the first semantic vector, the second tree and the second semantic vector to calculate a policy document analysis result.
2. The method according to claim 1, wherein Parsing the policy documents to obtain a clause tree, including: Parsing the policy documents to determine the format structure and / or semantic structure included in the policy documents; Based on the format structure and / or semantic structure, splitting to obtain clause statements to be used as nodes in the clause tree; If the format structure or the semantic structure is not found, performing segmentation processing on the policy documents to obtain clause statements to be used as nodes in the clause tree; Using the context relationship of the clause statements to construct the clause tree.
3. The method according to claim 2, wherein Performing vectorization processing on the clause statements in the clause tree, including: Obtaining the clause statements in the nodes of the clause tree; Performing vectorization processing on the clause statements to obtain a sentence vector with a specified dimension; Using the context relationship of the clause statements to construct the clause tree, including: According to the release information of the policy documents, determining the effective time and effectiveness level corresponding to each clause statement; Using the clause statements carrying the effective time and the effectiveness level as nodes to construct the clause tree.
4. The method according to claim 1, wherein Using the first tree and the first semantic vector, the second tree and the second semantic vector to calculate a clause analysis result, including: Calculating the cosine similarity and point mutual information between the first semantic vector and the second semantic vector; Using the cosine similarity and the point mutual information to calculate the correlation degree between the first semantic vector and the second semantic vector; In the case where the correlation degree is greater than the correlation degree threshold, calculating the clause analysis result, specifically including: determining the time difference of the effective time, the effectiveness level and the correlation degree; using the time difference, the effectiveness level, the correlation degree to calculate the clause analysis result.
5. The method according to claim 4, wherein Using the cosine similarity and the point mutual information to calculate the correlation degree between the first semantic vector and the second semantic vector, including: Analyzing the structural path similarity between a first node corresponding to the first semantic vector in the first tree and a second node corresponding to the second semantic vector in the second tree; When the structural path similarity is greater than a first threshold, calculating the cosine similarity and point mutual information between a first clause statement in the first node and a second clause statement in the second node; Using the cosine similarity and the point mutual information to calculate the correlation degree between the first semantic vector and the second semantic vector.
6. The method according to claim 5, wherein Calculating the cosine similarity and point mutual information between the first clause statement in the first node and the second clause statement in the second node includes: Extracting the keyword sets of the first clause statement and the second clause statement, specifically including: performing keyword extraction on the first document and the second document respectively using a specified extraction method; wherein, the specified extraction method includes: TF-IDF, BERT attention weight analysis, and entity / term extraction; Constructing a keyword set using the extracted keywords; Calculating the point mutual information using the keyword set; Calculating the cosine similarity between the first semantic vector corresponding to the first clause statement and the second semantic vector corresponding to the second clause statement; Analyzing the structural path similarity between the first node corresponding to the first semantic vector in the first tree and the second node corresponding to the second semantic vector in the second tree, including: Obtaining the first structural path for representing the context relationship of the first node in the first tree and the second structural path for representing the context relationship of the second node in the second tree; Calculating the structural path similarity between the first structural path and the second structural path.
7. The method according to claim 1, wherein Further including: Constructing an adaptive update model; Calculating the model weight vector after fusing new and old policy knowledge using the cross-entropy loss corresponding to the newly obtained policy document, the KL divergence of the historical policy document, the model parameters, and the penalty coefficient of the KL divergence.
8. An electronic device, characterized in that, Including: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, such that the processor executes the method according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that, The readable storage medium stores execution instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for calculating similarity of text data
CN102214232A
Policy file structured decomposition method
CN110609983A
Policy text correlation analysis method and system
CN112580348A
Structured policy knowledge graph construction method and system
CN117520553A
System and method for coupled detection of syntax and semantics for natural language understanding and generation
US20180137110A1
Cited By
Policy text multi-modal acquisition and evolution graph analysis method and policy text multi-modal acquisition and evolution graph analysis system
CN122174844A