Administrative law enforcement case quality supervision method and system
By building a knowledge system for verification points of administrative law enforcement documents and deep learning technology, combining structured information and entity information maps, and using a verification rule engine, all-round quality inspection of administrative law enforcement documents is achieved, the problem of incomplete quality control in the existing technology is solved, and the structure and legal legitimacy of documents are improved.
Patent Information
- Application Number
- CN202110865367.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-29
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-07-29
AI Technical Summary
The existing technology cannot conduct comprehensive and accurate quality inspection of administrative law enforcement documents, especially in the application of law enforcement legal basis, the construction of relationships between illegal behavior requirements, and the utilization of context information, resulting in incomplete and accurate quality control.
Design a knowledge system for verification points of administrative law enforcement documents, combine deep learning and natural language processing technology, build structured information and entity information maps, and use a verification rule engine to conduct comprehensive quality inspection, including structural integrity, content integrity, writing normativeness, etc.
It has achieved comprehensive and precise quality inspection of administrative law enforcement documents, improved the structural integrity, content integrity, writing standardization and legal legality of documents, and ensured that the quality of documents complies with legal requirements.
Smart Images

Figure CN114444477B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a method and system for supervising the quality of administrative law enforcement cases. Background Art
[0002] Administration according to law is a fundamental requirement of a government ruled by law. This practice, in part, manifests itself in the quality of administrative law enforcement documents. Because administrative law enforcement documents define the rights and obligations of the parties involved, they must be rigorous and accurate. Quality control of administrative law enforcement documents is a crucial aspect of their production and management. The quality of administrative law enforcement documents encompasses not only wording but, more importantly, legal principles, legal rules, and practical experience in law enforcement. These issues cannot be addressed through quality control of administrative law enforcement documents themselves; instead, they require the use of knowledge-based intelligent document quality detection technology.
[0003] A method and device for assisting administrative law enforcement (application number / patent number: CN202010269666.2) uses computer semantic analysis and keyword-related matching to obtain legal and regulatory texts and historical case information relevant to the current law enforcement case, then provides penalty recommendations, providing law enforcement personnel with legal decision-making support. Its features include: ① Automated decomposition and matching based on semantic features effectively addresses the difficulties of retrieving valid legal texts and determining the basis for administrative penalties. ② It proposes an effective relevance matching method, which increases query matching speed and accuracy the more basic data available.
[0004] The shortcomings of this patent are: (1) it only addresses the legal basis for law enforcement and the corresponding penalties; (2) it only establishes a simple correlation between case information and legal basis, without building a relationship around the application of legal basis based on the elements of the illegal behavior; and (3) it uses keyword matching technology but does not fully utilize the syntactic and semantic information data in the context. Based on the above three issues, it only serves as an auxiliary reference for the quality control of administrative law enforcement cases, and does not achieve the goal of automatic technical detection and control.
[0005] A method, apparatus, device, and storage medium for evaluating the quality of judicial documents (application number / patent number: CN110851591A): The core method includes (1) obtaining a target judicial document; (2) obtaining a target text from the target judicial document, wherein the target text includes a claim text and a reasoning text, and the claim text includes at least one claim; and (3) using a pre-established reasoning completeness detection model to detect whether each claim in the claim text is responded to by the reasoning text, and obtain a first detection result corresponding to each claim. The method for evaluating the quality of judicial documents provided in this application can automatically and efficiently evaluate the quality of judicial documents.
[0006] The shortcomings of this patent are: (1) it only detects the claim part; (2) when the claim is acknowledged by the other party, there is no need to respond in the reasoning text of the judgment document. This responsive evaluation method has major flaws at the legal professional level; (3) it uses text matching technology, and the specific algorithm uses CNN, which focuses on solving the problem of matching keywords and phrases with the same or similar semantics to the text. Since the claim itself includes multiple elements such as the claimant, the defendant, the type of claim, the content of the claim, and the direction of the claim, these elements are not fully included in the claim text, resulting in information loss, which cannot be solved by relying solely on the CNN algorithm. Based on the above three shortcomings, the quality evaluation results are incomplete and inaccurate.
[0007] A method, device, storage medium and processor for correcting errors in legal documents (application number / patent number: CN110750982A): The core method includes (1) obtaining a legal document to be processed; (2) identifying error points in the legal document to be processed, wherein the categories of the error points include at least one of document format errors, typos, punctuation and symbol errors, document writing errors, legal basis citation errors and litigation personnel name errors; (3) marking each error point in the legal document to be processed, and generating and displaying a modification prompt for each error point.
[0008] The shortcomings of this patent are: (1) the error point system based on the document knowledge system is incomplete; (2) this patent directly uses the error rule library to scan the documents to be inspected for errors. Since the object is unstructured text and the semantic expression of the text is very rich and open, no technical method is used in the preprocessing of the documents to be inspected. Instead, a mode of coupling text recognition algorithm and rule verification algorithm is used, which leads to poor generalization ability of the algorithm.
[0009] A text information quality measurement method under rule constraints (application number / patent number: CN110543628A): The core method includes (1) parsing the text into semi-structured text in XML format according to the rules; (2) classifying the text according to type, extracting the key fields that users are concerned about under each type, and forming a rule set under the type; (3) combining the rule set with the specific text content, and calculating the text information quality based on nine dimensions of information quality measurement indicators.
[0010] This patent focuses on measuring information quality from the information dimension of the document and does not directly point to the error problem of the document. Summary of the Invention
[0011] The technical problem to be solved by the present invention is to provide a method and system for supervising the quality of administrative law enforcement cases, which can comprehensively and accurately verify administrative law enforcement documents.
[0012] In order to solve the above technical problems, the present invention provides a method for supervising the quality of administrative law enforcement cases, the method comprising: designing a knowledge system for checking points of administrative law enforcement documents according to the writing specifications of administrative law enforcement documents, the semi-structural characteristics of administrative law enforcement documents, and the content knowledge system characteristics of administrative law enforcement documents; designing an information model for checking rules of administrative law enforcement documents according to the business characteristics and checking application requirements of each type of checking point in the knowledge system for checking points of administrative law enforcement documents; extracting and constructing rules from the checking rule base according to the information model for checking rules of each type of administrative law enforcement documents based on the business basis source, business basis information carrier and business information storage structure of each type of checking point; and extracting rules from the checking rule base according to the checking rule information model of each type of administrative law enforcement documents. The data structure specification of the law enforcement document verification rule information model stores the extracted and constructed rules in the administrative law enforcement document verification rule library; according to the writing specifications of administrative law enforcement documents, the content knowledge system characteristics of administrative law enforcement documents and the business knowledge corresponding to administrative law enforcement documents, the administrative law enforcement document content information model is designed and defined in XML format; according to the writing specifications and document structure of administrative law enforcement documents, according to the business standards corresponding to administrative law enforcement documents, rule-based natural language processing technology is used to divide the documents from coarse to fine into multi-level text pieces, and a document slicing model is designed; for the structured information in the administrative law enforcement document content information model, according to the writing specifications of administrative law enforcement documents, Based on the writing specifications and business characteristics of each type of entity information graph in the administrative law enforcement documents, an expert rule base is constructed, a sample set with annotations is created, and a hybrid model of deep learning algorithm and rule-based natural language processing technology is used to train the structured information extraction algorithm; for the entity information graph information in the content information model of administrative law enforcement documents, according to the writing specifications and business characteristics of each type of entity information graph in the administrative law enforcement documents, an expert rule base is constructed, a sample set with annotations is created, and a natural language processing technology based on a mixture of rules and syntactic dependencies is used to train the entity information graph extraction algorithm; according to the design characteristics of the administrative law enforcement document verification point knowledge system, the administrative law enforcement document verification rule information model and the administrative law enforcement document content information model, a set of supporting A verification rule engine that supports rule verification scanning; based on the algorithm results of the structured information extraction algorithm and the entity information graph extraction algorithm, it inputs administrative law enforcement documents and outputs administrative law enforcement document content information model instance data in XML format containing structured information and entity graph information; using the administrative law enforcement document verification point knowledge system as the connection standard, it imports the administrative law enforcement document content information model instance data into the verification rule engine, links the verification rules of the administrative law enforcement document verification rule library, performs verification, and outputs the verification results; based on the output verification results and in combination with the content of the administrative law enforcement document to be inspected, it adopts error backtracking positioning technology to mark the original text of document errors, display the error correction results, and display the rule basis.
[0013] In some embodiments, a knowledge system of checkpoints for administrative law enforcement documents is designed based on the writing standards of administrative law enforcement documents, the semi-structured characteristics of administrative law enforcement documents, and the characteristics of the content knowledge system of administrative law enforcement documents, including: designing a four-level classification system for the checkpoints of administrative law enforcement documents according to business standards; designing specific checkpoints for administrative law enforcement documents based on each type of checkpoints; sorting out the corresponding content knowledge system based on the type of administrative law enforcement documents, and designing checkpoints for each type of administrative law enforcement document.
[0014] In some embodiments, based on the business characteristics and verification application requirements of each type of verification point in the administrative law enforcement document verification point knowledge system, an administrative law enforcement document verification rule information model is designed by category, including: analyzing administrative law enforcement documents, and designing information items based on the dimensions of administrative law enforcement cases, administrative counterparts, law enforcement processes, illegal facts, evidence, law enforcement basis, and law enforcement results; analyzing specific verification points, and designing information items that are closely related to the verification points; designing information items that need to be deduced from the original text; and organizing the information items designed in the above steps according to the dimensions of cases, people, events, evidence, processes, basis, and results to form a verification rule information model.
[0015] In some embodiments, based on the business basis source, business basis information carrier and business information storage structure of each type of verification point, the verification rule library rules are extracted and constructed according to the verification rule information model of each type of administrative law enforcement document, including: designing the data structure of the verification rules, including the verification points to which the verification rules belong, verification prompts, the source of the verification rules and normative examples; and constructing the verification rules.
[0016] In some embodiments, constructing verification rules includes: using a visual construction method to drag information nodes and verification expressions from the verification rule information model tree to construct verification rules; or using a rule configuration method to construct verification rules in XML format.
[0017] In some embodiments, the extracted and constructed rules are stored in the administrative law enforcement document verification rule library according to the data structure specifications of the information model of each type of administrative law enforcement document verification rules, including: using XML as the storage format of the verification rules, storing the administrative law enforcement document verification rules in the rule library; and designing the structural content of each verification rule.
[0018] In some embodiments, according to the writing standards and document structure of administrative law enforcement documents and the corresponding business standards of administrative law enforcement documents, rule-based natural language processing technology is used to divide the documents from coarse to fine into multi-level text segments, and a document slicing model is designed, including: summarizing the writing standards and document structure of judicial documents, and dividing each paragraph of the document into multi-level text segments according to logical relationships; designing a document slicing model to store each logical segment of the document, each logical segment contains several fine slices.
[0019] In some embodiments, for the structured information in the content information model of administrative law enforcement documents, according to the writing standards and business characteristics of administrative law enforcement documents, an expert rule library is constructed, an annotated sample set is created, and a hybrid model of deep learning algorithms and rule-based natural language processing technologies is used to train structured information extraction algorithms, including: preprocessing data, for different classifications, some sentence content is meaningless to the classification, and these interfering data can be removed; performing word segmentation on sentences, selectively removing punctuation, line breaks, and stop words; using ALBERT as a word vector; defining a network structure, and building a deep learning model for classification based on LSTM.
[0020] In some embodiments, for the entity information graph information in the content information model of administrative law enforcement documents, according to the writing specifications and business characteristics of each type of entity information graph in the administrative law enforcement documents, an expert rule library is constructed, a labeled sample set is created, and a natural language processing technology based on a mixture of rules and syntactic dependencies is used to train the entity information graph extraction algorithm, including: according to the content characteristics of the basic information slices of the administrative counterparts in the administrative law enforcement documents, according to the business standards of the administrative counterparts in the administrative law enforcement documents, the information contained in the administrative counterparts in the administrative law enforcement documents is decomposed, and an administrative law enforcement document administrative counterpart information graph model is designed; according to the writing specifications and business characteristics of the administrative law enforcement documents, according to Based on the information graph model of administrative counterparts in administrative law enforcement documents, an information graph extraction algorithm for administrative counterparts in administrative law enforcement documents is designed; combining the document slicing model and the semantic features of each type of entity information graph, a document slicing model for this type of entity information graph is designed; according to the writing specifications and business characteristics of this type of entity information graph slices, according to the information model of this type of entity information graph, an information graph extraction algorithm for administrative counterparts in administrative law enforcement documents is linked, and an extraction algorithm for this type of entity information graph is designed; when it is necessary to use multiple associated entity information graph models for secondary association to form a new entity information graph, a new entity information graph comparison and reasoning algorithm is designed according to the business rules of the new entity information graph.
[0021] In some embodiments, based on the design features of the administrative law enforcement document verification point knowledge system, the administrative law enforcement document verification rule information model, and the administrative law enforcement document content information model, a verification rule engine that supports rule verification scanning is designed, including: designing a set of expression parsers. Usually, during verification, multiple values in the verification information model are involved, and these values need to be combined for calculation. By constructing a set of flexible expression parsers, complex business calculation scenarios can be met; designing a set of value parsers, using the underlying capabilities of Xpath, which can extract target values from the document content information model according to the combination of paths and conditions and fill them into calculation factors; designing a set of rule compilers and runners to preprocess the verification rules, compile them into high-performance machine code, and support parallel calculation of multiple sets of rules.
[0022] In some embodiments, according to the algorithm results of the structured information extraction algorithm and the entity information graph extraction algorithm, administrative law enforcement documents are input, and administrative law enforcement document content information model instance data containing structured information and entity graph information in XML format is output, including: based on the analysis of the attributes contained in the information items, information items can be divided into simple information items and complex information items. Information items with a single attribute that can be directly extracted from the original text or simply converted and then applied are called simple information items; after the administrative law enforcement documents are input, after calculation by the structured information extraction algorithm and the entity information graph extraction algorithm, administrative law enforcement document content information model instance data containing structured information and entity graph information is obtained.
[0023] In some embodiments, the administrative law enforcement document content information model instance data is imported into the verification rule engine with the administrative law enforcement document verification point knowledge system as the connection standard, the verification rules of the administrative law enforcement document verification rule library are linked, verification is performed, and the verification results are output, including: when the verification rule engine is started, the verification rules in the rule library are loaded, the rule compiler preprocesses the verification rules and compiles them into high-performance machine code; the administrative law enforcement document content information model instance data is imported into the verification rule engine, the value parser uses the underlying Xpath capability to obtain values from the information model, and fills the corresponding slots in the verification rules as calculation factors; the verification rule engine adopts multimodal computing, and for the verification rules that meet the trigger conditions, they are placed in the runner for parallel computing, and the verification results are output in XML format.
[0024] In addition, the present invention also provides an administrative law enforcement case quality supervision system, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the administrative law enforcement case quality supervision method described above.
[0025] After adopting such a design, the present invention has at least the following advantages:
[0026] The present invention relates to an intelligent quality detection technology for administrative law enforcement documents. Specifically, it adopts technical methods such as administrative law enforcement document structuring technology, administrative law enforcement document entity information map construction technology, verification rule expert engineering method, verification engine technology, etc. to realize all-round quality detection of administrative law enforcement documents from the aspects of structural integrity, content integrity, writing standardization, logicality, content legality, word errors, punctuation errors, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0028] Figure 1 It is a system structure diagram;
[0029] Figure 2 This is a screenshot of the XML storage format of the administrative law enforcement document content information model;
[0030] Figure 3 This is the structural diagram of the ALBERT-BiLSTM-CRM model;
[0031] Figure 4 This is the ALBERT model structure diagram;
[0032] Figure 5 This is the LSTM (Long Short-Term Memory Neural Network) structure diagram;
[0033] Figure 6 It is a model flow chart;
[0034] Figure 7 It is a schematic diagram of the original text of the administrative penalty decision;
[0035] Figure 8 This is a schematic diagram of the output in the application. DETAILED DESCRIPTION
[0036] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0037] The present invention relates to an intelligent quality detection technology for administrative law enforcement documents. The technology specifically adopts administrative law enforcement document structuring technology, administrative law enforcement document entity information map construction technology, verification rule expert engineering method, verification engine technology and other technical methods to realize all-round quality detection of administrative law enforcement documents from the aspects of structural integrity, content integrity, writing standardization, logicality, content legality, word errors, punctuation errors, etc., and belongs to the field of legal natural language processing technology.
[0038] The embodiments of the present application provide a technology and device for intelligent quality detection of administrative law enforcement documents.
[0039] 1. This invention is an intelligent quality detection technology for administrative law enforcement documents
[0040] Focusing on the writing standards of administrative law enforcement documents, the semi-structured characteristics of administrative law enforcement documents, the content knowledge system characteristics of administrative law enforcement documents, and the external business practice rules and domain terminology characteristics based on legal constraints, legal rules and administrative law enforcement documents, an information model contained in administrative law enforcement documents and a rule base information model based on the external rule knowledge system are designed. A technical method for structuring administrative law enforcement documents for extracting structured information from administrative law enforcement documents, a technical method for extracting and constructing administrative entity information maps from administrative law enforcement documents, a technical method for extracting legal knowledge rules, and a verification engine technical method for calling legal knowledge rules to detect administrative law enforcement document information models are proposed. Quality inspection is carried out from the aspects of the structural integrity, content integrity, writing standardization, logicality, content legality, compliance with the discretionary power of administrative law enforcement standards, word errors, punctuation errors, etc. of the documents, ultimately achieving the goal of knowledge rule quality inspection of the text content of administrative law enforcement documents. The invention can be more effectively applied to the quality inspection of various types of administrative law enforcement documents, and play a role in the fields of administrative law enforcement document production, administrative law enforcement document error correction, administrative law enforcement document quality review, business system information proofreading, etc.
[0041] 2. The present application embodiment provides an intelligent error correction device for administrative law enforcement documents
[0042] The present invention comprises the following steps:
[0043] Step (1) Designing a knowledge system of administrative law enforcement document checkpoints based on the writing standards of administrative law enforcement documents, the semi-structured characteristics of administrative law enforcement documents, and the content knowledge system characteristics of administrative law enforcement documents;
[0044] Administrative law enforcement documents, used to declare the rights and obligations of all parties involved, have strict writing standards and are characterized by a distinct semi-structured nature. This implementation design incorporates litigation standards and writing standards into a knowledge system for verifying administrative law enforcement documents. The specific steps are as follows:
[0045] Step (1.1) A four-level classification system is designed for the verification of administrative law enforcement documents according to business standards. The first level includes "formal verification" and "substantive verification". Under the "formal verification" classification, the second level is further divided into "structural integrity verification", "structural rationality verification", "key information must exist verification", "typo verification", "punctuation verification", "legal language and content verification", "inconsistency verification between multiple information items", etc. Under the "substantive verification" classification, it is further divided into "procedural legality verification" and "substantive legality verification". "Procedural legality verification" can be further divided into "administrative reconsideration and administrative litigation legality verification", "applicable procedure legality verification", and "substantive legality verification" can be further divided into "reference of laws and regulations accuracy verification", "penalty result legality verification" and "multi-document mutual verification".
[0046] Step (1.2) Then, based on each type of verification, design specific verification points for administrative law enforcement documents. For example, the specific verification points under "Legal language and content verification" include "verifying whether the name of the evidence type complies with legal provisions."
[0047] Step (1.3) According to the type of administrative law enforcement document, sort out the corresponding content knowledge system and design verification points for each type of administrative law enforcement document. For example, the specific verification points of administrative penalty documents include "verifying whether the cited regulations and laws are accurate" and "whether the applicable penalty type is legal".
[0048] Step (2) designing an administrative law enforcement document verification rule information model by category based on the business characteristics and verification application requirements of each type of verification point in the administrative law enforcement document verification point knowledge system;
[0049] For each checkpoint, it is usually not possible to verify based on a single piece of information. Multiple information items are required to form a complete checkpoint. This embodiment designs a verification rule information model based on the business characteristics and verification application requirements of each type of checkpoint. The specific steps are as follows:
[0050] Step (2.1) Analyze administrative law enforcement documents and design information items based on the dimensions of administrative law enforcement cases, administrative counterparts, law enforcement process, illegal facts, evidence, law enforcement basis, and law enforcement results;
[0051] Step (2.2) Analyze the specific checkpoints and design information items closely related to the checkpoints;
[0052] Step (2.3) Design the information items that need to be inferred from the original text, such as the type of illegal behavior;
[0053] Step (2.4) Organize the information items designed in the above steps according to the dimensions of case, person, event, evidence, process, basis, result, etc. to form a verification rule information model.
[0054] For example, the Administrative Penalty Decision requires: (1) the name or title, address and other basic information of the party. If the party has a principal qualification certificate, the name of the principal qualification certificate, unified social credit code (registration number), residence (address), legal representative (person in charge, operator) and other information shall be written in accordance with the matters recorded in the principal qualification certificate. If the party is an individual business owner and has a business name, the business name shall be used as the party name, and the name of the operator, the name and number of the ID card (other valid certificate) shall be written at the same time. If the principal qualification certificate of the party does not have a unified social credit code, the registration number or other number shall be written. If the party is an individual, the name, address, number and other information shall be written in accordance with the matters recorded in the ID card (other valid certificate).
[0055] According to this requirement, a verification point can be designed as "verification of whether the basic information of the administrative counterpart with the subject qualification certificate is complete". According to the verification point knowledge system, it can be specifically identified as a verification point in the "key information must exist verification" category under the "formal verification" category. The information model involved includes the following information items: whether the administrative counterpart has a subject qualification certificate, the name of the administrative counterpart's subject qualification certificate, the administrative counterpart's unified social credit code (registration number), the administrative counterpart's residence (address), and the legal representative of the administrative counterpart (person in charge, operator).
[0056] To identify whether the administrative counterpart has a subject qualification certificate, in the administrative penalty decision, extract entities such as "Administrative counterpart's subject qualification certificate name", "Administrative counterpart's unified social credit code (registration number)", "Administrative counterpart's residence (address)", "Administrative counterpart's legal representative (person in charge, operator)", etc., and verify whether each piece of information exists.
[0057] Step (3) extracting and constructing a verification rule base according to the business basis source, business basis information carrier, and business information storage structure of each type of verification point and the verification rule information model of each type of administrative law enforcement document;
[0058] Specific steps:
[0059] Step (3.1) Design the data structure of the verification rules, including the verification points to which the verification rules belong, verification prompts, the source of the verification rules and normative examples.
[0060] Step (3.2) Construct the validation rules. Method 1: Use the visual construction method to drag and drop information nodes and validation expressions from the validation rule information model tree to construct the validation rules. Method 2: Use the rule configuration method to construct the validation rules in XML format.
[0061] Step (4) storing the extracted and constructed rules in the administrative law enforcement document verification rule library according to the data structure specification of the information model of each type of administrative law enforcement document verification rule;
[0062] Specific steps:
[0063] Step (4.1) uses XML as the storage format of the verification rules and stores the administrative law enforcement document verification rules in the rule base.
[0064] Step (4.2) designs the structure of each verification rule, including: configuring the triggering conditions of the verification rule, the calculation factors involved in the verification, the verification expression, etc.
[0065] Step (5) designing an administrative law enforcement document content information model based on the writing standards of administrative law enforcement documents, the characteristics of the content knowledge system of administrative law enforcement documents, and the business knowledge corresponding to administrative law enforcement documents, and defining it in XML format;
[0066] Step (5.1) This embodiment designs an XML node specification for representing the document content information model. Each information item corresponds to an XML node. The node name is uniformly named "<information group>", the node attribute "name" is the Chinese name of the information item, and the node attribute "value" is the content of the information item.
[0067] Step (6) Based on the writing specifications and document structure of administrative law enforcement documents and the corresponding business standards of administrative law enforcement documents, rule-based natural language processing technology is used to divide the documents into multi-level text slices from coarse to fine, and a document slicing model is designed;
[0068] By reading a large number of administrative law enforcement documents and summarizing their writing patterns, this patent proposes a method for structuring administrative law enforcement documents from coarse to fine scale. Based on the court's requirements for document writing standards, the documents are divided into multiple text segments and a document slicing model is designed to store each logical paragraph. The specific steps are as follows:
[0069] Step (6.1) summarizes the writing standards and structure of judicial documents, and divides each paragraph of the document into multi-level text segments according to the logical relationship. For example, the first-level text segment of the judgment document is analyzed, including the "text header", "administrative counterpart section", "case source and investigation process section", "illegal facts and related evidence", "administrative penalty basis and decision" and "text tail".
[0070] Step (6.2) Design a document slicing model to store each logical segment of a document. Each logical segment contains several subslices. Based on the subslices contained in each paragraph, design a document slicing model. As shown in the figure, each subslice is stored as a string type and named according to the content contained. The entire slicing model uses a tree structure for storage.
[0071] Step (7) for the structured information in the content information model of administrative law enforcement documents, according to the writing standards and business characteristics of administrative law enforcement documents, an expert rule library is constructed, a sample set with annotations is created, and a hybrid model of deep learning algorithm and rule-based natural language processing technology is used to train the structured information extraction algorithm;
[0072] For structured information extraction, the specific methods of this implementation are:
[0073] 1) Define several named entity recognition models to identify entities such as names, dates, subject matter, amounts, evidence, and types of penalties in administrative law enforcement documents;
[0074] The entity recognition model mainly consists of three parts: the ALBERT pre-trained language model, the BILSTM layer, and the CRF layer. The ALBERT encoding output is used as the input of the BILSTM layer, and a CRF layer is added after the hidden layer of the BILSTM for decoding, and finally the label type of each character is obtained. The specific structure is as follows Figure 3 .
[0075] Traditional language models, such as neural network language models, are unidirectional and therefore unable to incorporate contextual information. Furthermore, because the trained word embeddings are fixed, they cannot represent word polysemy. The ALBERT model structure effectively addresses both of these issues. The ALBERT model structure is shown in the figure. It uses a bidirectional Transformer as the encoder, replacing the LSTM with the more effective Transformer. The bidirectional language model allows BERT to capture contextual information, thereby enriching word embeddings with richer semantic information.
[0076] 2) Define several classification models to identify the categories of information items in administrative law enforcement documents;
[0077] Since the original information items that need to be classified in documents are mostly short texts, commonly used short text classification methods include CNN, RNN, LSTM, Attention, etc. This implementation uses LSTM for model training to achieve a relatively good effect.
[0078] The structure of LSTM (Long Short-Term Memory Neural Network) is shown in the figure. It adds the concept of gates (input gate, forget gate, output gate) to RNN. Simply put, it is like a valve that can control the degree of memory and forgetting of previous and current information, thereby giving the RNN network a long-term memory function.
[0079] step:
[0080] 1. Preprocess the data. For different classifications, some sentence contents are meaningless to the classification, and these interfering data can be removed;
[0081] 2. Segment the sentences and selectively remove punctuation, line breaks, stop words, etc.
[0082] 3. Use ALBERT as word vector;
[0083] 4. Define the network structure and build a deep learning model for classification based on LSTM, such as Figure 6 shown.
[0084] Step (8) for the entity information graph information in the content information model of the administrative law enforcement document, according to the writing specifications and business characteristics of each type of entity information graph in the administrative law enforcement document, build an expert rule library, create a sample set with annotations, and use the natural language processing technology based on the hybrid of rules and syntactic dependencies to train the entity information graph extraction algorithm;
[0085] Step (8-1) decomposes the information contained in the administrative counterparty in the administrative law enforcement document according to the content characteristics of the basic information slice of the administrative counterparty in the administrative law enforcement document and the business standards of the administrative counterparty in the administrative law enforcement document, and designs an information map model of the administrative counterparty in the administrative law enforcement document;
[0086] Specific steps:
[0087] Step (8-1-1) Design the administrative counterpart information map model.
[0088] Step (8-1-2) designs the Schema representation of the information graph, encapsulates RDFS through the Schema, and provides an object-oriented description system that supports class inheritance and attribute polymorphism.
[0089] Step (8-2) designs an administrative law enforcement document administrative counterpart information graph extraction algorithm based on the administrative law enforcement document administrative counterpart information graph model according to the writing standards and business characteristics of the administrative law enforcement document;
[0090] Specific steps:
[0091] Step (8-2-1) first identifies the cause of the administrative action through the document. Taking the cause of the administrative action as a prerequisite can more accurately determine the type of subject of the administrative legal relationship.
[0092] Step (8-2-2) defines a named entity recognition model to identify administrative counterpart information and administrative legal relationship subject types.
[0093] Step (8-2-3) defines a relationship recognition model to identify the relationship between entities and construct a subject information graph through entities and relationships;
[0094] Step (8-3) combines the document slicing model with the semantic features of each type of entity information graph to design a document slicing model for this type of entity information graph;
[0095] Step (8-3-1) of this implementation usually identifies a large amount of entity information in an administrative law enforcement document. If the entire document is used as a connection between the relationships between entities to construct a graph, not only will the amount of graph calculation increase, but it will also cause interference between information. This implementation uses the entity information graph to limit a certain range. On the basis of document slicing, the text segments of each type of entity information are further sliced. For example, for the entity of administrative penalty results, the "administrative penalty decision result segment" under the "administrative penalty basis and decision segment" is sliced.
[0096] Step (8-3-2) designs a document slicing model of this type of entity information graph to store each logical segment. Each logical segment contains several fine slices, and the entire slicing model is stored in a tree structure.
[0097] Step (8-4) is to design an algorithm for extracting the information graph of the administrative counterpart of administrative law enforcement documents based on the writing specifications and business characteristics of the slice of the entity information graph of this type, according to the information model of the entity information graph of this type, and link it to the algorithm for extracting the information graph of the administrative counterpart of administrative law enforcement documents;
[0098] Specific steps:
[0099] Step (8-4-1) defines a relationship recognition model to identify the relationship between entities;
[0100] Through syntactic dependency and Chinese semantic role analysis, the parallel relationship between entities is identified to construct an entity information graph.
[0101] In this example, the sample data is first labeled. The target sequence labeling model can be used to label the Chinese semantic roles of the corpus in the test set, that is, to predict the Chinese semantic roles of the words and obtain predicted labels. For example, after inputting "The party compensates the consumer for nursing fees of 1205.5 yuan" into the target sequence labeling model, the output is "[Party Agent][Compensation V][Consumer Dative][Nursing Fee Patient][1205.5 yuan]." Where "compensation" is the predicate verb, "Party", "Consumer", and "Nursing Fee" are labeled respectively, corresponding to the agent, the subject, and the patient.
[0102] Step (8-4-2) Link the administrative law enforcement documents to the administrative counterpart information map, apply the reference resolution technology, and match the entity information in the document with the administrative counterpart information to form a more complete map.
[0103] Step (8-5) is to design a new entity information graph comparison and reasoning algorithm according to the business rules of the new entity information graph when it is necessary to use multiple related entity information graph models for secondary association to form a new entity information graph;
[0104] In administrative law enforcement documents with complex case descriptions, such as those involving multiple individuals, multiple illegal administrative acts, and multiple facts, reasoning is necessary to complete or exclude certain elements. For example, by matching the rights and obligations of the obligees and obligors in the administrative penalty decision with the corresponding rights and obligations of the obligees and obligors in the illegal act section, and then reasoning, this allows for "verification of administrative penalty decisions beyond the scope of the illegal act."
[0105] Step (9) designing a set of verification rule engines that support rule verification scanning based on the design characteristics of the administrative law enforcement document verification point knowledge system, the administrative law enforcement document verification rule information model, and the administrative law enforcement document content information model;
[0106] This embodiment is designed as an independent verification rule engine, which separates the verification rule logic of data and developer technology, improves the maintainability of verification rules that implement complex logic, and supports the order of verification rules and rule conflict detection.
[0107] Specific steps:
[0108] Step (9-1) Design an expression parser. Verification typically involves multiple values in the verification information model, and requires combined calculations on these values. By constructing a flexible expression parser, we can meet complex business calculation scenarios.
[0109] Step (9-2) Design a value parser that uses the underlying capabilities of XPath to extract target values from the document content information model and fill them into the calculation factors based on the combination of paths and conditions.
[0110] Step (9-3) Design a set of rule compilers and runners to preprocess the verification rules, compile them into high-performance machine code, and support parallel computing of multiple sets of rules.
[0111] Step (10) inputs administrative law enforcement documents according to the algorithm results of step (7) and step (8), and outputs administrative law enforcement document content information model instance data containing structured information and entity graph information in XML format;
[0112] Step (10-1) analyzes the attributes contained in the information items and divides the information items into simple information items and complex information items. Information items with a single attribute that can be directly extracted from the original text or simply converted and then applied are called simple information items. For example, the reasons for administrative actions, types of illegal acts, names of administrative law enforcement agencies, etc. are relatively easy to extract from the original text. However, due to historical factors, the original text of the document may be written in abbreviations or legal standards at the time. For such information items, this embodiment provides a conversion method by adding a dictionary table and configuring matching rules for identification, and converting them into standard values stipulated by existing laws.
[0113] After the administrative law enforcement document is input in step (10-2), after the calculation in steps (7) and (8), the administrative law enforcement document content information model instance data containing structured information and entity graph information is obtained. This embodiment designs a set of conversion instructions to output the instance data into XML format;
[0114] Step (11) uses the administrative law enforcement document verification point knowledge system as the connection standard, imports the administrative law enforcement document content information model instance data into the verification rule engine, links the verification rules of the administrative law enforcement document verification rule library, performs verification, and outputs the verification results;
[0115] Specific steps:
[0116] Step (11-1) When the verification rule engine is started, the verification rules in the rule base are loaded, and the rule compiler pre-processes the verification rules and compiles them into high-performance machine codes.
[0117] Step (11-2) imports the instance data of the administrative law enforcement document content information model into the verification rule engine. The value parser uses the underlying Xpath capability to obtain values from the information model and fills the corresponding slots in the verification rule as calculation factors.
[0118] Step (11-3) The verification rule engine uses multimodal computing. For verification rules that meet the triggering conditions, they are put into the runner for parallel computing and the verification results are output in XML format.
[0119] Step (12) uses error backtracking positioning technology based on the output verification results and the content of the administrative law enforcement document to be inspected to mark the original text of the document errors, display the correction results, and display the rule basis.
[0120] This step belongs to the human-computer interaction interface. It uses error backtracking positioning technology to directly display the correction results, the basis of the correction rules, and examples of correct expressions in the original text, which can bring a better user experience.
[0121] Specifically:
[0122] One method is plain text retrospective positioning. The administrative law enforcement document being verified is plain text, and an HTML5 web page can be used to display the original text and highlight errors.
[0123] One method is to retrospectively locate Word documents or WPS documents. When identifying such documents, they are converted into plain text for analysis. Since Word or WPS documents may contain implicit typesetting symbols or tab characters and other non-document content data, and the verification analysis needs to be performed in plain text, during the conversion to plain text, this embodiment will record the position difference between the plain text and the original Word document as an offset. In order not to change the format of the original document, when displaying, it will be relocated to the position in the Word document according to the offset of the verification body.
[0124] For marking erroneous original text, one way is to use a single verification target, such as a typo, and then directly trace back to the typo; one is to use multiple targets, such as the result of consistency verification, and then trace back to multiple targets; and another is missing verification, and then trace back to the paragraph where it is located, or the adjacent paragraph.
[0125] For the display of verification results and rule basis, the Web-based Html5 technology is used for display.
[0126] Example: In an administrative penalty decision, the original description of the penalty result is as follows: Figure 7 shown.
[0127] Article 62 of the Regulations for the Implementation of the Tobacco Monopoly Law states: Anyone who sells illegally produced tobacco products in violation of Article 27 and the first paragraph of Article 40 of these Regulations shall be ordered by the tobacco monopoly administrative department to stop sales, confiscate illegal gains, and be fined not less than 20% and not more than 50% of the total amount of illegal sales. The illegally sold tobacco products shall be publicly destroyed. The "Guiding Opinions on the Discretionary Power of Administrative Penalties for Tobacco Monopoly in Jiangsu Province (Trial)" stipulates that if the total amount of illegal sales exceeds 5,000 yuan, the person shall be ordered to stop sales, confiscate illegal gains, and be fined not less than 40% and not more than 50% of the total amount of illegal sales. The illegally sold tobacco products shall be publicly destroyed.
[0128] According to the above technical solution, the following steps are performed:
[0129] (1) Design a knowledge system for administrative law enforcement document verification points:
[0130] (2) Design verification rule information model
[0131] (3) Verification rule construction
[0132] (4) Storage verification rule base
[0133] (5) Design content information model
[0134] (6) Design document slicing model
[0135] (7) Training structured information extraction algorithm
[0136] (8) Training entity information graph extraction algorithm
[0137] (9) Design verification rule engine
[0138] (10) Output model instance data
[0139] (11) Verify and output verification results
[0140] (12) Human-computer interaction
[0141] Finally, the output in the application is as follows: Figure 8 shown.
[0142] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Those skilled in the art can make some simple modifications, equivalent changes or modifications based on the technical content disclosed above, which all fall within the scope of protection of the present invention.
Claims
1. A method for supervising the quality of administrative law enforcement cases, characterized in that: include: Based on the writing standards of administrative law enforcement documents, the semi-structured characteristics of administrative law enforcement documents, and the content knowledge system characteristics of administrative law enforcement documents, a knowledge system of administrative law enforcement document checkpoints is designed, including: A four-level classification system has been designed for the verification of administrative law enforcement documents according to business standards. The first level of the four-level classification system includes formal verification and substantive verification. The formal verification includes structural integrity verification, structural rationality verification, verification of the existence of key information, verification of typos, verification of punctuation, verification of legal language and content, and verification of inconsistencies between multiple information items. The substantive verification includes procedural legality verification and entity legality verification. The procedural legality verification includes the legality verification of administrative reconsideration and administrative litigation, and the legality verification of applicable procedures. The entity legality verification includes the accuracy verification of cited regulations and laws, the legality verification of penalty results, and the mutual verification of multiple documents. Design specific checkpoints for administrative law enforcement documents based on each type of checkpoint; According to the types of administrative law enforcement documents, sort out the corresponding content knowledge system and design checkpoints for each type of administrative law enforcement document; According to the business characteristics and verification application requirements of each type of verification point in the administrative law enforcement document verification point knowledge system, the administrative law enforcement document verification rule information model is designed by category; According to the business basis source, business basis information carrier and business information storage structure of each type of verification point, the verification rule base is extracted and constructed according to the verification rule information model of each type of administrative law enforcement document; According to the data structure specification of the information model of each type of administrative law enforcement document verification rule, the extracted and constructed rules are stored in the administrative law enforcement document verification rule library; According to the writing standards of administrative law enforcement documents, the characteristics of the content knowledge system of administrative law enforcement documents and the business knowledge corresponding to administrative law enforcement documents, the content information model of administrative law enforcement documents is designed and defined in XML format; Based on the writing specifications and structure of administrative law enforcement documents and the corresponding business standards, rule-based natural language processing technology is used to segment documents into multi-level text slices from coarse to fine, and a document slicing model is designed; For the structured information in the content information model of administrative law enforcement documents, based on the writing standards and business characteristics of administrative law enforcement documents, we build an expert rule library, create annotated sample sets, and use a hybrid model of deep learning algorithms and rule-based natural language processing technology to train structured information extraction algorithms; For the entity information graph information in the content information model of administrative law enforcement documents, according to the writing standards and business characteristics of each type of entity information graph in administrative law enforcement documents, an expert rule library is built, a labeled sample set is created, and a natural language processing technology based on a mixture of rules and syntactic dependencies is used to train the entity information graph extraction algorithm; According to the design characteristics of the administrative law enforcement document verification point knowledge system, the administrative law enforcement document verification rule information model and the administrative law enforcement document content information model, a verification rule engine that supports rule verification scanning is designed; According to the algorithm results of the structured information extraction algorithm and the entity information graph extraction algorithm, administrative law enforcement documents are input and administrative law enforcement document content information model instance data containing structured information and entity graph information in XML format is output; Taking the administrative law enforcement document verification point knowledge system as the connection standard, the administrative law enforcement document content information model instance data is imported into the verification rule engine, the verification rules of the administrative law enforcement document verification rule library are linked together, verification is performed, and the verification results are output; Based on the output verification results and combined with the content of the administrative law enforcement documents to be inspected, error backtracking positioning technology is used to mark the original text of document errors, display the correction results, and display the rule basis.
2. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: According to the business characteristics and verification application requirements of each type of verification point in the administrative law enforcement document verification point knowledge system, the administrative law enforcement document verification rule information model is designed by category, including: Analyze administrative law enforcement documents and design information items based on the dimensions of administrative law enforcement cases, administrative counterparts, law enforcement process, illegal facts, evidence, law enforcement basis, and law enforcement results; Analyze specific checkpoints and design information items closely related to the checkpoints; Design information items that need to be derived from the original text; The information items designed in the above steps are organized according to the dimensions of case, person, event, evidence, process, basis, and result to form a verification rule information model.
3. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: Based on the business basis source, business basis information carrier and business information storage structure of each type of verification point, the verification rule base is extracted and constructed according to the verification rule information model of each type of administrative law enforcement document, including: Design the data structure of the verification rules, including the verification points to which the verification rules belong, verification prompts, the source of the verification rules, and normative examples; Build validation rules.
4. The method for supervising the quality of administrative law enforcement cases according to claim 3, characterized in that: Build validation rules, including: Use a visual construction method to drag information nodes and verification expressions from the verification rule information model tree to build verification rules; or Use rule configuration method and XML format to build verification rules.
5. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: According to the data structure specification of the information model of each type of administrative law enforcement document verification rules, the extracted and constructed rules are stored in the administrative law enforcement document verification rule library, including: XML is used as the storage format for verification rules, and administrative law enforcement document verification rules are stored in the rule base; Design the structure and content of each verification rule.
6. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: Based on the writing specifications and structure of administrative law enforcement documents and the corresponding business standards, rule-based natural language processing technology is used to segment documents into multi-level text slices from coarse to fine. A document slicing model is designed, including: Summarize the writing standards and structure of judicial documents, and divide each paragraph of the document into multi-level text segments according to the logical relationship; A document slicing model is designed to store the logical segments of a document, and each logical segment contains several fine slices.
7. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: For the structured information in the content information model of administrative law enforcement documents, based on the writing standards and business characteristics of administrative law enforcement documents, we build an expert rule library, create annotated sample sets, and use a hybrid model of deep learning algorithms and rule-based natural language processing technology to train structured information extraction algorithms, including: Preprocess the data. For different classifications, some sentence contents are meaningless to the classification, so remove these interfering data. Segment the sentences and selectively remove punctuation, line breaks, and stop words; Use ALBERT as word vector; Define the network structure and build a deep learning model for classification based on LSTM.
8. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: For the entity information graph information in the content information model of administrative law enforcement documents, based on the writing standards and business characteristics of each type of entity information graph in administrative law enforcement documents, an expert rule library is constructed, a labeled sample set is created, and a natural language processing technology based on a mixture of rules and syntactic dependencies is used to train the entity information graph extraction algorithm, including: According to the content characteristics of the basic information slices of the administrative counterparts in administrative law enforcement documents and the business standards of the administrative counterparts in administrative law enforcement documents, the information contained in the administrative counterparts in administrative law enforcement documents is decomposed to design an information map model of the administrative counterparts in administrative law enforcement documents; According to the writing standards and business characteristics of administrative law enforcement documents, and based on the administrative counterpart information graph model of administrative law enforcement documents, an algorithm for extracting administrative counterpart information graphs of administrative law enforcement documents is designed; Combining the document slicing model with the semantic features of each type of entity information graph, a document slicing model for this type of entity information graph is designed. According to the writing specifications and business characteristics of this type of entity information graph slices, according to the information model of this type of entity information graph, link the administrative law enforcement document administrative counterpart information graph extraction algorithm, and design this type of entity information graph extraction algorithm; When it is necessary to use multiple associated entity information graph models for secondary association to form a new entity information graph, a new entity information graph comparison and reasoning algorithm is designed according to the business rules of the new entity information graph.
9. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: According to the algorithm results of the structured information extraction algorithm and the entity information graph extraction algorithm, administrative law enforcement documents are input and the administrative law enforcement document content information model instance data containing structured information and entity graph information in XML format is output, including: Based on the attributes contained in information items, information items are divided into simple information items and complex information items. Information items with a single attribute that can be directly extracted from the original text or simply converted and applied are called simple information items; After inputting the administrative law enforcement documents, the structured information extraction algorithm and the entity information graph extraction algorithm are used to calculate and obtain the administrative law enforcement document content information model instance data containing structured information and entity graph information.
10. The method for supervising the quality of administrative law enforcement cases according to claim 1, characterized in that: Using the administrative law enforcement document verification point knowledge system as the connection standard, the administrative law enforcement document content information model instance data is imported into the verification rule engine, the verification rules of the administrative law enforcement document verification rule library are linked together, verification is performed, and the verification results are output, including: When the verification rule engine starts, it loads the verification rules in the rule library, and the rule compiler preprocesses the verification rules and compiles them into high-performance machine code; Import the instance data of the administrative law enforcement document content information model into the verification rule engine. The value parser uses the underlying XPath capability to obtain values from the information model and fill the corresponding slots in the verification rule as calculation factors. The verification rule engine uses multimodal computing. For verification rules that meet the triggering conditions, they are placed in the runner for parallel computing and the verification results are output in XML format.
11. An administrative law enforcement case quality supervision system, characterized by: include: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the administrative law enforcement case quality supervision method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Text information quality measurement method under rule constraints
CN110543628A
Legal document error correction method and device, storage medium and processor
CN110750982A
Quality evaluation method, device and equipment for judgment document, and storage medium
CN110851591A
Administrative law enforcement auxiliary method and device
CN111581327A
Civil complaint and judgment map construction method and system
CN113010684A