Formal conversion method of supervision rules, terminal equipment and storage medium
Through regular expressions and abstract syntax tree construction, combined with formal verification of Z3 solver, syntax errors and conflicts in the conversion of unstructured regulatory rules text into structured data representations are solved, and the automation and consistency of regulatory rules are improved.
Patent Information
- Application Number
- CN202510440451.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
AI Technical Summary
It is difficult for the prior art to efficiently convert unstructured regulatory rule text into structured data representations, and find potential syntax errors or rule conflicts during the conversion process, resulting in inefficiency in supervision.
Regular expressions are used to identify lexical units, build an abstract syntax tree (AST), generate syntax tree nodes through recursive descent parsing functions, identify the rule scope, and use the Z3 solver for formal verification to ensure logical consistency.
The automatic regulatory rules are converted into structured data representations, and the detection of potential conflicts and errors is improved through formal verification, which improves the logical consistency and management efficiency of regulatory rules.
Smart Images

Figure CN120278121A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and formal verification, and particularly to a method for formal transformation of regulatory rules, a terminal device, and a storage medium, which transform regulatory rules from semi-structured and unstructured texts into structured data representations and perform formal verification. Background Art
[0002] With the rapid development of information technology, especially driven by Internet, big data, and artificial intelligence technologies, the service industry is undergoing profound changes. Against this background, regulatory agencies are facing unprecedented challenges and need to effectively manage and supervise an increasingly complex service environment to ensure service compliance and security. To address these challenges, formal modeling and transformation technologies for regulatory languages have emerged, aiming to convert unstructured regulatory regulations into structured forms that can be understood and executed by computers.
[0003] Regulatory language is a formal language for describing regulatory rules and policies, which is crucial for ensuring that service providers comply with laws and regulations, maintaining market order, and protecting consumers' rights and interests. However, traditional regulatory regulations are usually written in natural language and have problems such as ambiguous expression, diverse interpretations, and difficulty in automated processing. These problems limit regulatory efficiency and effectiveness and increase regulatory costs. To improve the automation and intelligence level of regulation, it is necessary to formalize regulatory regulations into machine-readable and executable models. This process not only includes extracting elements such as regulatory intentions, conditions, and actions from natural language texts, but more importantly, systematically detecting and verifying the rules through formal means to ensure their logical consistency and correctness. The research and development of formal regulatory languages aim to create a method for representing regulatory rules that can be understood and executed by computer systems, thereby better supporting the automated detection, verification, and management of regulatory rules.
[0004] Currently, existing work can achieve the automatic transformation of natural language text regulations into CDSRL regulatory language, mainly based on a two-stage regulatory language conversion method using large models. However, due to limitations such as large model hallucinations, the CDSRL regulatory language transformed by the current large model may have potential syntax errors or constraint conflicts and therefore needs to be further parsed and verified.
[0005] Regulation text parsing is the basis of regulation text analysis, and its goal is to convert unstructured or semi-structured regulation texts into structured data representations for subsequent computer processing. Traditional manual parsing methods are costly and inefficient. Since regulatory rules are automatically generated by LLM, there may be syntax errors or constraint omissions, and different degrees of conflicts may occur in the different scopes of multiple rules. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a formal transformation method, a terminal device and a storage medium for regulatory rules, which can automatically convert unstructured regulatory rule texts into structured data representations, and perform logical detection and verification on the rules through formal methods to discover potential conflicts, inconsistencies or errors in view of the deficiencies of the prior art.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is: A formal transformation method for regulatory rules, comprising the following steps:
[0008] Use regular expressions to identify each lexical unit in the regulation text, and split each lexical unit into basic lexical units Token. The basic lexical units include keywords, identifiers, operators, etc., and output a Token stream;
[0009] Adopt a recursive descent method to convert the Token stream into an Abstract Syntax Tree (AST) according to predefined syntax rules. For each type of rule, write a parsing function that is responsible for parsing the input text of the corresponding type of rule and gradually generating syntax tree nodes to form an AST for subsequent verification;
[0010] Identify the scopes in the regulatory rules, and represent the regulatory rules as a set of rule scopes. Adopt a metadata-based identification method to extract the explicitly marked scope strings from the rule documents;
[0011] By parsing the scope strings of each specific rule, convert the scope strings of each rule into a set form, calculate the intersection and union of the scope sets between the rules, and thus output multiple independent scopes, each scope containing several rules;
[0012] Extract the verifiable constraints in the regulatory rules according to the Abstract Syntax Tree (AST) corresponding to the relevant rules, parse the conditional part of the regulatory rules, and convert it into a combination of boolean variables and logical operators;
[0013] In different independent scopes, use the Z3 solver as a formal verification tool within each independent scope to check the satisfiability of the logical expressions.
[0014] The present invention is specifically optimized for the characteristics of regulatory languages and performs excellently in dealing with scope conflicts and formal verification. One of the core features of regulatory languages is the existence of scopes. Rule conflicts between different scopes do not affect the operation of the overall system, but the rules within the same scope must maintain logical consistency. To this end, the present invention ensures that there are no logical conflicts or inconsistencies in the rules within the same scope by introducing key technologies such as scope identification, intersection and union calculations of scope sets, and adaptive constraint formal verification. Therefore, the present invention has significant technical advantages in the formal transformation and verification of regulatory rules.
[0015] The specific implementation process of splitting each lexical unit of the present invention into keywords and logical operators includes:
[0016] Define a set of regular expression rules for matching keywords and logical operators in the regulation text;
[0017] Scan the regulation text line by line, apply the regular expression rules to each line, extract the keywords and logical operators, and store the keywords and logical operators as lexical units in a list or an array.
[0018] The specific implementation process of generating the corresponding nodes in the AST of the present invention includes:
[0019] Define a set of grammar rules, and each grammar rule corresponds to a parsing function;
[0020] Each parsing function receives a list of lexical units as input, performs matching and parsing according to the grammar rules, and generates nodes of the AST.
[0021] The specific implementation process of verifying the AST includes: traversing each node of the AST and checking whether it conforms to the predefined grammar rules, including checking whether the child nodes of each node exist, whether the attributes are complete, and whether the node order is correct.
[0022] The specific implementation process of using the Z3 solver as a formal verification tool to check the satisfiability of logical expressions within each independent scope includes:
[0023] Within each independent scope, if the Z3 solver finds an assignment such that all logical expressions are true, it is determined that there are no conflicts between the regulatory rules within this scope; otherwise, it is determined that there are conflicts or inconsistencies in the regulatory rules under this scope.
[0024] As an inventive concept, the present invention also provides a terminal device, including a memory, a processor, and a computer program stored on the memory; the processor executes the computer program to implement the steps of the above method.
[0025] As an inventive concept, the present invention also provides a computer-readable storage medium having a computer program / instructions stored thereon; when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0026] As an inventive concept, the present invention also provides a computer program product, including a computer program / instructions; when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention can automatically convert unstructured regulatory rule texts into structured data representations and perform formal verification. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a flowchart of the method of the embodiment of the present invention;
[0029] Figure 2 is an example diagram of the parsing and verification process of the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] Embodiment 1
[0032] As Figure 1 shown in the flowchart of the method of the embodiment of the present invention, the formal modeling and transformation method of the regulatory rules provided by the embodiment of the present invention includes the following steps:
[0033] S1. Use regular expressions to identify each lexical unit in the regulation text and split it into keywords and logical operators;
[0034] S2. Construct a syntax analysis module and use the recursive descent method to construct an AST according to predefined syntax rules. For each rule, write a parsing function that is responsible for parsing the input text of the corresponding structure and generating the corresponding node in the AST;
[0035] S3. Verify the AST to check whether its syntax structure is incorrect, including whether each node contains necessary child nodes and attributes, whether the node order conforms to the predefined syntax rules, and cross-line attributes, etc.;
[0036] S4. Identify the scopes in the regulatory rules and represent them as sets. Adopt a metadata-based identification method to extract the explicitly marked scope information from the rule documents.
[0037] S5. Calculate the intersection and union of the sets of rule scopes. Define a function to parse the scope strings and convert them into set forms.
[0038] S6. Extract the verifiable constraints in the regulatory rules. First, parse the conditional part of the rules and convert it into a combination of boolean variables and logical operators. Next, use a predefined function to parse the scope strings and convert them into logical expressions for formal verification.
[0039] S7. Within different scopes, use the Z3 solver as a formal verification tool to check the satisfiability of the logical expressions. Determine whether there is a variable assignment that makes all expressions true, thereby detecting conflicts, inconsistencies, or errors among the rules.
[0040] The use of regular expressions to identify each lexical unit in the regulation text described in step S1 and split it into keywords and logical operators is specifically implemented as follows: First, define a set of regular expression rules for matching keywords (such as "Rule", "Entity", etc.) and logical operators (such as "and", "or", etc.) in the regulation text. Then, scan the regulation text line by line, apply the regular expressions to each line, extract the keywords and logical operators, and store them as lexical units in a list or array. The purpose of this step is to prepare the input data for subsequent syntactic analysis and ensure that each lexical unit can be accurately identified and classified.
[0041] The construction of the syntactic analysis module described in step S2 adopts a recursive descent method to construct an AST according to predefined syntactic rules. For each rule, write a parsing function that is responsible for parsing the input text of the corresponding structure and generating the corresponding nodes in the AST. The specific implementation is as follows: Define a set of syntactic rules, each rule corresponding to a parsing function. These parsing functions call each other recursively to handle nested and complex syntactic structures. Each parsing function receives a list of lexical units as input, matches and parses according to the syntactic rules, generates the nodes of the AST, and constructs the parent-child relationships between the nodes. For example, for a rule structure, its parsing function will identify the "Rule" keyword and look for subsequent attributes such as "id" and "type", and add these attributes as child nodes under the rule node.
[0042] The verification of the AST described in step S3, checking whether its syntax structure is incorrect, including whether each node contains necessary child nodes and attributes, whether the node order conforms to predefined syntax rules, and cross-line attributes, etc., is specifically implemented as follows: Design a verification module that traverses each node of the AST and checks whether it conforms to the predefined syntax rules. This includes checking whether the child nodes of the node exist, whether the attributes are complete, and whether the node order is correct. For example, if a "Rule" node lacks the "id" attribute or the order of the "Entity" nodes is incorrect, the verification module will mark these errors and provide error messages.
[0043] The following are some key syntax rules:
[0044] 1) RuleDefinition: "RuleDefinition" is the root node, containing child nodes such as "MetaDataSet",
[0045] "ActionSet", and "Rule".
[0046] 2) MetaDataSet: The "MetaDataSet" node contains two child nodes, "Entities" and "Externals".
[0047] 3) Entities: The "Entities" node contains multiple "Entity" child nodes.
[0048] 4) Entity: The "Entity" node contains attributes such as "id", "name", and "Property", where
[0049] "Property" is a list containing multiple "Property" elements.
[0050] 5) Property: The "Property" node contains attributes such as "id", "name", "dataType", "unit", etc., where "dataType" and "unit" are optional attributes.
[0051] 6) Rule: The "Rule" node contains attributes such as "id", "type", "requirement", "relation",
[0052] "PreCondition", "RegulatoryMeasure", "Coverage", and "Constraint". Among them, the "relation" attribute is an optional attribute; "PreCondition", "RegulatoryMeasure",
[0053] The "Coverage" and "Constraint" attributes may contain multiple lines of content and require special handling.
[0054] 7) When "Property" appears, "Entity", "Entities", and "MetaDataSet" must appear; when "Entity" appears, "Entities" and "MetaDataSet" must appear; when "Entities" or "Externals" appears, "MetaDataSet" must appear; when "PreCondition", "RegulatoryMeasure", "Description", "Relation", "Coverage", or "Constraint" appears, "Rule" must appear; when "MetaDataSet", "ActionSet", and "Rule" appear, "RuleDefinition" must appear.
[0055] For attributes such as "PreCondition", "RegulatoryMeasure", "Coverage", and "Constraint" that may contain multiple lines of content, the parser processes them using a state machine. When a keyword in "PreCondition", "Coverage", "Constraint", or "RegulatoryMeasure" is encountered, the keyword is assigned to multiline_key, and the content after the keyword is assigned to multiline_value, indicating the start of recording an attribute across multiple lines. In the subsequent loop, if multiline_key is not empty and the current line is not "end", the content of the current line is appended to multiline_value to accumulate the content across multiple lines. When the "end" keyword is encountered and multiline_key is not empty, it indicates the end of the current attribute across multiple lines. According to the value of multiline_key, the accumulated multiline_value is assigned to the corresponding node's attribute. Finally, multiline_key and multiline_value are reset to empty to prepare for processing the next attribute across multiple lines. The specific matching of attributes across multiple lines is shown in Table 1.
[0056] Table 1 Matching Table for Attributes across Multiple Lines
[0057]
[0058] Identify the scope in the regulatory rules described in step S4 and represent it as a set. The scope refers to the specific conditions or range to which the rule applies. In service supervision, the scope may include geographical regions (such as countries, provinces, cities, etc.), time ranges (such as dates, time periods, etc.), service types (such as financial services, medical services, etc.), and other factors that may affect the applicability of the rules. To accurately identify the scope, the system needs to define a complete scope model and identifier system. The scope identification methods mainly include metadata-based identification and keyword matching-based identification. The metadata-based identification method relies on the scope information clearly marked in the rule document, such as domain information or geographical information, which is clearly marked in the rule file header; the keyword matching-based identification method determines the scope by matching specific keywords or phrases in the rule text. For example, "no more than 5 road traffic safety violations occurred while driving a motor vehicle within one year before the application date" belongs to "entity existence or activation", and "those who have engaged in cruise vehicle services and have not been included in the serious illegal information database for taxis;" belongs to "entity attribute check". In the rule document, the scope information usually appears in a specific metadata form, such as the "Coverage" field. By parsing these metadata, the scope information is extracted and stored as a set. For example, if the "Coverage" field contains "OnlineDoctor", it is parsed into a set containing "OnlineDoctor".
[0059] Calculate the intersection and union of the scope sets of the rules described in step S5. Define a function to parse the scope string and convert it into a set form. The elements in the set can be geographical location codes, timestamps, service type IDs, etc. Intersection calculation: For any two rules R1 and R2, calculate the intersection R1.scope ∩ R2.scope of their scope sets. If the intersection is non-empty, it indicates that these two rules apply simultaneously under the conditions represented by the intersection. Union calculation: Similarly, calculate the union R1.scope ∪ R2.scope of the rule scope sets to understand the combined applicable ranges of these rules. The specific implementation is as follows: The parser traverses all the rules, extracts their scopes, and calculates the intersection and union. To achieve this, first parse and store the scopes of all the rules in a dictionary, where the key is the ID of the rule and the value is the scope set of that rule. Then, use nested loops to traverse all possible pairs of rules, calculate their intersection and union, and store the results in the corresponding data structures. For example, if the scopes of two rules are {A, B} and {B, C} respectively, then their intersection is {B}, and the union is {A, B, C}.
[0060] Extract the verifiable constraints in the supervision rules described in step S6. First, parse the conditional part of the rules and convert it into a combination of Boolean variables and logical operators. Next, use a predefined function to parse the scope string and convert it into a logical expression for formal verification. The specific implementation is as follows: For the conditional part of each rule, use the predefined parsing function to convert it from text form to a Boolean logic expression. These expressions are composed of Boolean variables and logical operators (such as "and", "or", "not") and can be processed by formal verification tools. For example, the condition "OnlineDoctor.Qualification == 'Licensed Physician'" can be converted into the logical expression "OnlineDoctor.Qualification = Licensed Physician".
[0061] In different scopes described in step S7, use the Z3 solver as a formal verification tool to check the satisfiability of the logical expressions. Determine whether there is a variable assignment that makes all expressions true, so as to detect conflicts, inconsistencies or errors between rules. The specific implementation is as follows: Input the logical expressions obtained in step S6 into the Z3 solver, and the solver will try to find an assignment for the variables that makes all expressions true. If the solver finds such an assignment, it means there are no conflicts between the rules; if not, it means there are conflicts or inconsistencies. For example, if two rules have conflicting requirements for the same service in the same scope, such as two rules, one requires "OnlineDoctor.Experience >= 5" and the other requires "OnlineDoctor.Experience < 5", the Z3 solver will be able to detect this conflict and report it. An important advantage of this method is that it can systematically check all possible pairs of rules, not just the obvious conflicts. This enables the validator to discover potential problems in advance and correct them before the rules are formally applied, thus improving the accuracy and reliability of rule management. Then the Z3 solver will not be able to find a variable assignment that satisfies both conditions, thus detecting the conflict between the rules.
[0062] The following further elaborates on the present invention in combination with a practical case:
[0063] First, input the set of supervision languages composed of any supervision rule files. First, identify each lexical unit by applying regular expressions to identify keywords and logical operators, and then apply the recursive descent method to the identified lexical units to construct an AST. Each grammar rule corresponds to a parsing function to generate AST nodes.
[0064] Next, using predefined syntax rules, verify the constructed AST to ensure the correctness of the syntax structure, including the necessary child nodes, attributes, and order of the nodes.
[0065] Finally, extract the scope information from the rule document and represent it as a set. Calculate the intersection and union of the scope sets to provide the necessary set operation results for the transformation and verification of logical expressions. Parse the conditional part of the rule, convert the conditions into a combination of boolean variables and logical operators, and convert the scope strings into logical expressions. Use the Z3 solver to check the satisfiability of the logical expressions within different scopes to ensure the logical consistency and correctness among the rules.
[0066] Embodiment 2
[0067] Embodiment 2 of the present invention provides a terminal device corresponding to Embodiment 1 above. The terminal device can be a processing device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of the above embodiment.
[0068] The terminal device of this embodiment includes a memory, a processor, and a computer program stored on the memory; the processor executes the computer program on the memory to implement the steps of the method of Embodiment 1 above.
[0069] In some implementations, the memory can be a high-speed random access memory (RAM: Random Access Memory), and may also include non-volatile memory, such as at least one disk memory.
[0070] In other implementations, the processor can be various types of general-purpose processors such as a central processing unit (CPU), a digital signal processor (DSP), etc., which are not limited herein.
[0071] Embodiment 3
[0072] Embodiment 3 of the present invention provides a computer-readable storage medium corresponding to Embodiment 1 above, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the method of Embodiment 1 above are implemented.
[0073] A computer-readable storage medium can be a tangible device that holds and stores instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.
[0074] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript, etc.
[0075] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0077] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0078] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A formal transformation method for regulatory rules, characterized in that, Including the following steps: Use regular expressions to identify each lexical unit in the regulation text, and split each lexical unit into basic lexical units Token to obtain a Token stream; The basic lexical units include keywords, identifiers, and operators; Adopt a recursive descent method to convert the Token stream into an Abstract Syntax Tree (AST) according to predefined syntax rules; Identify the scopes in the regulatory rules and represent the regulatory rules as a set of rule scopes. Adopt a metadata-based identification method to extract the explicitly marked scope strings from the rule documents; By parsing the scope strings of each specific rule, convert the scope strings of each rule into a set form, calculate the intersection and union of the scope sets between the rules, and obtain multiple independent scopes, with each scope containing several rules; According to the Abstract Syntax Tree (AST) corresponding to the relevant rules, extract the verifiable constraints in the regulatory rules, parse the conditional part of the regulatory rules, and convert it into a combination of boolean variables and logical operators; In different independent scopes, use the Z3 solver as a formal verification tool within each independent scope to check the satisfiability of the logical expressions.
2. The formal transformation method of the supervision rules according to claim 1, characterized in that The specific implementation process of splitting each lexical unit into basic lexical units includes: Define a set of regular expression rules for matching keywords and logical operators in the regulation text; Scan the regulation text line by line, apply the regular expression rules to each line, extract the keywords and logical operators, and store the keywords and logical operators as lexical units in a list or array.
3. The formal transformation method of the supervision rules according to claim 1, characterized in that The specific implementation process of generating the Abstract Syntax Tree (AST) nodes includes: For each type of rule, write a parsing function that is used to parse the input text of the corresponding type of rule and gradually generate the Abstract Syntax Tree (AST) nodes.
4. The formal transformation method of the supervision rule according to claim 1, characterized in that After converting the Token stream into an Abstract Syntax Tree (AST) according to the predefined syntax rules, verify the AST. The specific implementation process includes: Traverse each node of the AST and check whether it conforms to the predefined syntax rules, including checking whether the child nodes of each node exist, whether the attributes are complete, and whether the node order is correct.
5. The formal transformation method of the supervision rule according to claim 1, characterized in that The specific implementation process of using the Z3 solver as a formal verification tool to check the satisfiability of the logical expressions within each independent scope includes: Within each independent scope, if the Z3 solver finds an assignment that makes all logical expressions true, it is determined that there is no conflict between the regulatory rules within this scope; otherwise, it is determined that there is a conflict or inconsistency in the regulatory rules under this scope.
6. A terminal device, comprising a memory, a processor, and a computer program stored on the memory; characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program / instructions stored thereon; characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer program product, comprising a computer program / instructions; characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.