Rule extraction method, device and equipment of unstructured text and storage medium
Patent Information
- Application Number
- CN202611005694.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请实施例的目的是提供一种非结构化文本的规则提取方法、装置、设备及存储介质,能够解决现有的规则审查方案难以兼顾准确性和效率的问题
在本申请实施例中,首先,通过对非结构化文本进行解释性文本与指令性文本的物理分流,并利用解释性文本构建只读的专业术语字典,能够减少大语言模型对专业术语的自由解释与语义漂移,有利于消除指令性文本的拆解过程的歧义;其次,在步骤拆解和规则节点生成阶段,引入预定义的实体列表和动作列表,对审查对象、审查动作和审查目标进行强类型约束与校验,避免大语言模型自行创造动作或引用未定义实体,降低了规则提取过程中的错误率;再次,根据执行步骤之间的逻辑关系为每个规则节点配置分支跳转逻辑,并将多个规则节点串联为审查规则有向无环图,能够确定审查流程的无环性、确定终止性和执行路径可追溯性,将复杂的多条件的合规审查转换为线性开销,构建了明确的审查流程;最后,将三元组机器指令编译为抽象语法树,将抽象语法树与大语言模型解耦,既保障了审查流程的执行安全性,又能够实现批量化的自动化合规审查流程。
Smart Images

Figure CN122817461A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of document analysis technology, specifically relating to a method, apparatus, device, and storage medium for rule extraction from unstructured text. Background Technology
[0002] In rule review, especially legal review scenarios, the mainstream solutions include: making compliance determinations based on the basic legal knowledge built into the pre-training stage of the Large Language Model (LLM); or relying on manually written review rules and injecting prompt words into the model to achieve constraint verification.
[0003] However, if reviews are conducted solely based on the built-in knowledge of the large model, the black-box nature of the model makes it impossible for staff to verify the accuracy of its reasoning logic, easily leading to omissions and misjudgments. Manually compiling and organizing review rules is costly and inefficient, and errors cannot be fully understood. Therefore, existing rule-based review schemes struggle to balance accuracy and efficiency. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and storage medium for extracting rules from unstructured text, which can solve the problem that existing rule review schemes are difficult to balance accuracy and efficiency.
[0005] The technical solution adopted by this application to solve its technical problem is: In a first aspect, embodiments of this application provide a method for extracting rules from unstructured text, the method comprising: The unstructured text to be processed is classified into explanatory text and instructional text; the explanatory text is used to construct a dictionary of technical terms. Based on the aforementioned terminology dictionary, preset step prompts, and action list, the instruction text is broken down to obtain multiple execution steps; Based on the preset entity list, the action list, and the multiple execution steps, a triplet machine instruction is generated and multiple rule nodes are obtained. Based on the logical relationship between the multiple execution steps, branch jump logic is configured for each rule node, and the multiple rule nodes are connected in series to form a directed acyclic graph of review rules; Based on the triplet machine instructions corresponding to each rule node, the triplet machine instructions are compiled into an abstract syntax tree.
[0006] Secondly, a data review device for multi-source documents, the device comprising: The text classification module is used to classify the unstructured text to be processed into explanatory text and instructional text; the explanatory text is used to construct a dictionary of technical terms. The text decomposition module is used to decompose the instruction text into multiple execution steps based on the technical terminology dictionary, preset step prompts, and action list; The node generation module is used to generate triplet machine instructions and obtain multiple rule nodes based on the preset entity list, the action list and the multiple execution steps; The process generation module is used to configure branch jump logic for each rule node according to the logical relationship between the multiple execution steps, and to connect the multiple rule nodes into a directed acyclic graph of review rules. The instruction compilation module is used to compile the triplet machine instructions into an abstract syntax tree based on the triplet machine instructions corresponding to each rule node.
[0007] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the steps of the rule extraction method for unstructured text as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the rule extraction method for unstructured text as described in the first aspect.
[0009] The beneficial effects of this application are: In this embodiment, firstly, by physically separating unstructured text into explanatory and instructional texts, and using the explanatory text to construct a read-only dictionary of technical terms, the free interpretation and semantic drift of technical terms by the large language model can be reduced, which helps to eliminate ambiguity in the decomposition process of instructional text. Secondly, in the step decomposition and rule node generation stages, predefined entity lists and action lists are introduced to impose strong type constraints and verifications on review objects, review actions, and review targets, preventing the large language model from creating actions on its own or referencing undefined entities, thus reducing the error rate in the rule extraction process. Thirdly, branch jump logic is configured for each rule node according to the logical relationship between execution steps, and multiple rule nodes are connected to form a directed acyclic graph of review rules, which can determine the acyclicity of the review process, the determination of termination, and the traceability of the execution path, transforming complex multi-condition compliance review into linear overhead and constructing a clear review process. Finally, triple machine instructions are compiled into an abstract syntax tree, and the abstract syntax tree is decoupled from the large language model, which not only ensures the execution security of the review process but also enables a batch automated compliance review process. Attached Figure Description
[0010] Figure 1 This is a flowchart of a rule extraction method for unstructured text provided in an embodiment of this application.
[0011] Figure 2 This is a flowchart illustrating the specific steps of a rule extraction method for unstructured text provided in this application embodiment.
[0012] Figure 3 This is a schematic diagram of a step identification provided in an embodiment of this application.
[0013] Figure 4 This is a schematic diagram of the compilation result of a rule node provided in an embodiment of this application.
[0014] Figure 5 This is a schematic diagram of a branch jump logic for a rule node provided in an embodiment of this application.
[0015] Figure 6 This is a schematic diagram of a rule extraction and compilation process provided in an embodiment of this application.
[0016] Figure 7 This is a block diagram of a rule extraction device for unstructured text provided in an embodiment of this application.
[0017] Figure 8 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] The present application will be further described below with reference to the accompanying drawings and embodiments.
[0019] The following will clearly and completely describe the concept, specific structure, and resulting technical effects of this application in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of this application. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application. Furthermore, all connections / linkages involved in the patent do not simply refer to direct contact between components, but rather to the ability to form a better connection structure by adding or reducing connecting accessories according to specific implementation conditions. The various technical features in this application can be combined interactively without contradicting each other.
[0020] Currently, large language models have been widely applied in legal scenarios such as contract review and procurement compliance review. Existing legal review rule extraction solutions are mainly divided into two categories: one is to complete compliance judgment based on the basic legal knowledge built into the pre-training stage of the large model; the other is to rely on manually written review rules and inject prompt words into the model to achieve constraint verification.
[0021] The existing mainstream methods for obtaining legal review rules include manually sorting out and compiling various legal compliance rules; or using massive legal professional datasets to train large models specifically, relying on the model's own memory to carry compliance judgment logic.
[0022] However, relying solely on the legal knowledge built into large models for review is prone to the "artificial intelligence (AI) illusion" problem due to the black-box nature of these models. This can lead to rule omissions, misjudgments, and other issues. Furthermore, the model's reasoning logic cannot be traced back to its source, making it difficult to adjust and optimize judgment criteria. If rules are manually written and delivered to the model via prompts, it not only consumes significant manpower and results in inefficient rule compilation, but the natural language form of the rules is inherently ambiguous, failing to completely eliminate the comprehension biases of the large model and ultimately affecting the accuracy and stability of legal review results.
[0023] It is evident that the existing legal review rule extraction scheme has the problem of creating the illusion of artificial intelligence (AI) and low accuracy and stability of legal review results.
[0024] To address the aforementioned problems, this application provides a method, apparatus, device, and storage medium for extracting rules from unstructured text. The method for extracting rules from unstructured text provided in this application will be described in detail below with reference to the accompanying drawings.
[0025] Figure 1 This is a flowchart of a rule extraction method for unstructured text provided in an embodiment of this application. See also... Figure 1 The method includes the following steps.
[0026] Step 101: Classify the unstructured text to be processed to obtain explanatory text and instructional text.
[0027] In this embodiment of the application, explanatory text is used to construct a dictionary of technical terms.
[0028] In this application embodiment, unstructured text includes rule documents in the legal and financial fields.
[0029] In some embodiments, step 101 may include: performing character recognition and layout analysis on the unstructured text to obtain the original text content; slicing the original text content according to chapter structure, heading level, or semantic boundary to obtain multiple text fragments; and dividing the multiple text fragments into explanatory text and instructional text according to preset triage prompts.
[0030] In some other embodiments, step 101 may further include: performing rule-based initial screening on unstructured text based on chapter titles and preset keyword templates to obtain candidate explanatory text fragments and candidate instructional text fragments; calling a large model to perform structured discrimination on candidate text fragments based on preset classification prompts to determine the category label of each text fragment; performing consistency verification on the category labels, and if there are conflicts or discrimination results with confidence levels lower than preset thresholds, triggering re-discrimination or manual review to finally obtain explanatory text and instructional text.
[0031] In some embodiments, after step 101, the above-mentioned rule extraction method for unstructured text may further include: constructing a terminology dictionary based on explanatory text; wherein the terminology dictionary includes at least one term entry, each term entry records the term name, alias and standardized definition, and the terminology dictionary is a read-only closed mapping table used for standardized parsing and consistency verification of terminology.
[0032] It should be noted that explanatory texts are texts with explanatory, defining, and declarative characteristics, used to define terms, explain concepts, and specify the scope of application. They do not have direct enforceability and can be understood as term definitions. Typical sentence structures include: "X as used in this specification refers to...", "X refers to...", "Unless otherwise defined...", etc. Instructive texts are texts with binding, directive, and conditional characteristics, used to stipulate what should and should not be done. They can be directly converted into enforcement rules and can be understood as compliance rules. Typical sentence structures include: "shall...", "must...", "if...then...", "shall not...", etc.
[0033] Step 102: Based on the professional terminology dictionary, preset step prompts and action list, decompose the instruction text to obtain multiple execution steps.
[0034] In the embodiments of this application, the terminology dictionary can be a mapping relationship between terms and standardized definitions, used to eliminate ambiguity in the understanding of terminology by large models and to prevent term definitions from being misjudged as execution rules.
[0035] In this embodiment, step prompts are used to constrain the large model to decompose the instructional text into pseudocode-like execution steps that conform to predefined syntax rules. The predefined syntax rules include step type constraints, action whitelist constraints, and structured output constraints.
[0036] In some embodiments, step type constraint information is used to limit text statements to conditional steps or verification steps; wherein, conditional steps are used to determine the applicable scenario of the rule, and verification steps are used to perform compliance judgments.
[0037] In this embodiment of the application, the action list is an executable verification action, including multiple predefined and fixed standard actions; the review action must completely match a certain standard operator in the list, and large models are prohibited from creating or using actions outside the list, and users cannot modify the action list.
[0038] In some embodiments, the standard actions include numerical comparison actions, semantic judgment actions, and set judgment actions. The numerical comparison actions include greater than, less than, or equal to; the semantic judgment actions include belonging to or meeting a condition; and the set judgment actions include including or being included.
[0039] In some embodiments, step 102 may include: standardizing and parsing the technical terms in the instructional text according to a technical terminology dictionary; converting the statements in the instructional text into multiple rule steps according to step prompts; the multiple rule steps include condition steps or verification steps; verifying whether the actions in the multiple rule steps belong to the action list to obtain multiple execution steps.
[0040] Step 103: Based on the preset entity list, action list, and multiple execution steps, generate triplet machine instructions and obtain multiple rule nodes.
[0041] In this embodiment of the application, the entity list includes contract entities, technical rule entities, and process record material entities.
[0042] In this embodiment, the triplet machine instruction includes a review object, a review action, and a review target, wherein the review target includes a constant review target or a variable review target.
[0043] In some embodiments, the execution steps include a conditional step and a verification step; wherein the conditional step is a control flow step and the verification step is a computation flow step.
[0044] In this embodiment, each triplet machine instruction corresponds to a rule node, and the data structure of the rule node conforms to a format recognizable by the sandbox engine.
[0045] In some embodiments, step 103 may include: compiling the step contents of multiple execution steps with strong type constraints according to the entity list and the action list to generate triple machine instructions corresponding to each execution step; and generating multiple rule nodes according to the triple machine instructions.
[0046] Step 104: Based on the logical relationship between multiple execution steps, configure branch jump logic for each rule node, and connect multiple rule nodes into a directed acyclic graph of review rules.
[0047] In this embodiment, the triplet machine instruction includes the review object, the review action, and the review target.
[0048] In some embodiments, the review object includes the target entity, the target attribute of the target entity, and the data type of the target attribute; the target entity is a predefined entity type in the entity list, the target attribute is an element in the set of legal attributes in the entity list, and the data type is an element in the data type whitelist in the entity list; the review action is a predefined standard action in the action list.
[0049] In some embodiments, each entity has different attribute information, such as the contract amount and the name of the party A, which can be obtained from the contract. This attribute information is generated by the large model based on rule requirements and entity descriptions.
[0050] In some embodiments, the review target includes constant review targets or variable review targets; variable review targets include target entities, target attributes, and data types. Constant review targets have fixed data types and values, such as strings and company names, or numbers and 220; variable review targets are standard objects that need to be extracted from the document to be reviewed.
[0051] In this embodiment of the application, the Directed Acyclic Graph (DAG) of the review rules is the flow control logic corresponding to the specific rules in legal review.
[0052] In some embodiments, step 104 may include: configuring branch jump logic for each rule node, the branch jump logic including the node flow direction determined according to the review result of the current rule node, the node flow direction including: the next judgment node, the inapplicable final state node, the compliant final state node, or the non-compliant final state node; determining the node jump relationship according to the node flow direction of each rule node; and connecting all rule nodes in series according to the node jump relationship to obtain a directed acyclic graph of review rules.
[0053] Step 105: Compile the triplet machine instructions into an abstract syntax tree based on the triplet machine instructions corresponding to each rule node.
[0054] In this embodiment, the Abstract Syntax Tree (AST) is the set of executable instructions corresponding to specific rules in legal review. Specifically, the AST is an executable AST of the sandbox engine.
[0055] In some embodiments, step 105 may include: reading the review object and review target, and the left and right leaf nodes of the root node; wherein the leaf nodes are encapsulated as data access objects for extracting specific values from the context during execution; performing type consistency checks on the root node and its child nodes according to the attribute data types in the entity list by the compiler; after completing the checks, constructing a tree structure with a defined parent-child relationship, wherein the internal nodes represent specific computational logic and the leaf nodes represent data sources; and serializing the constructed abstract syntax tree into a JavaScript Object Notation (JSON) format that the sandbox engine can recognize.
[0056] Understandably, since the JSON format only contains node types, operators, data paths, and constant values, and does not contain any natural language descriptions, it helps ensure that the sandbox engine can execute the logic safely and efficiently in an isolated environment.
[0057] In summary, in this embodiment, firstly, by physically separating unstructured text into explanatory and instructional texts, and using the explanatory text to construct a read-only dictionary of technical terms, the free interpretation and semantic drift of technical terms by the large language model can be reduced, which helps to eliminate ambiguity in the decomposition process of instructional text. Secondly, in the step decomposition and rule node generation stages, predefined entity lists and action lists are introduced to impose strong type constraints and verifications on review objects, review actions, and review targets, preventing the large language model from creating actions or referencing undefined entities, thus reducing the error rate in the rule extraction process. Thirdly, branch jump logic is configured for each rule node according to the logical relationship between execution steps, and multiple rule nodes are connected to form a directed acyclic graph of review rules, which can determine the acyclicity, termination, and traceability of the execution path of the review process, transforming complex multi-condition compliance review into linear overhead and constructing a clear review process. Finally, triple machine instructions are compiled into an abstract syntax tree, and the abstract syntax tree is decoupled from the large language model, which not only ensures the execution security of the review process but also enables a batch automated compliance review process.
[0058] Figure 2 This is a flowchart illustrating the specific steps of a rule extraction method for unstructured text provided in this application embodiment. The method includes the following steps.
[0059] Step 201: Perform character recognition and layout analysis on the unstructured text to obtain the original text content.
[0060] In this application embodiment, unstructured text includes rule documents in the legal and financial fields.
[0061] In some embodiments, rule documents typically exist in Portable Document Format (PDF), scanned, or image formats. Text recognition includes Optical Character Recognition (OCR), which identifies character information in images, while layout analysis identifies the physical and logical structure of a document.
[0062] In some embodiments, step 201 may include: performing OCR on the unstructured text to obtain text data; performing layout analysis on the unstructured text to obtain header and footer, paragraph boundaries, table areas, illustration areas, and heading level information.
[0063] For example, after OCR and layout analysis, the document "XX Group Contract Management Standard" was identified to contain chapter titles such as "Chapter 1 General Provisions", "Chapter 2 Terminology Definitions", "Chapter 3 Contract Formation", and "Chapter 4 Contract Review", and the start and end positions of each chapter were located.
[0064] Step 202: Slice the original text content according to chapter structure, heading level, or semantic boundaries to obtain multiple text fragments.
[0065] In some embodiments, step 202 includes at least one of steps 2021 to 2023.
[0066] Step 2021: Divide the content within a chapter into a text segment according to the explicit chapter markers as slice boundaries.
[0067] For example, explicit chapter markers can be chapter numbers such as "Chapter X" or "Section X".
[0068] Step 2022: Using the title hierarchy information obtained from the layout analysis, aggregate consecutive content under the same title hierarchy into a single text fragment.
[0069] For example, heading level information can be first-level heading, second-level heading, third-level heading, etc.
[0070] Step 2023: For texts without obvious chapter markers, aggregate consecutive sentences expressing the same complete semantic meaning into a text segment based on paragraph spacing, line break density, or semantic integrity.
[0071] For example, regarding the "XX Group Contract Management Standards", the entire "Chapter Two Terminology Definitions" is divided into one text segment, and the entire "Chapter Three Contract Conclusion" is divided into another text segment.
[0072] In this way, the slicing process can effectively avoid breaking a complete term definition or a complete compliance rule into multiple fragments, thereby ensuring the accuracy of subsequent classification.
[0073] Step 203: Based on the preset triage prompts, divide multiple text fragments into explanatory text and instructional text.
[0074] In this embodiment, the traffic detour prompt can be a semantic firewall rule set, used to classify text fragments into explanatory text or instructional text based on modal verbs, conditional conjunctions, and definitional verbs in the text fragments. In some embodiments, the triage prompts are predefined structured compilation instructions used to constrain the output behavior of large language models and output classification results in a structured format.
[0075] In some embodiments, the triage prompts include explanatory text judgment rules, imperative text judgment rules, conflict resolution rules, and structured output constraints. The explanatory text judgment rules are used to determine whether a text fragment is a terminology definition based on the absence of defining verbs, scope verbs, and obligatory modal verbs. The imperative text judgment rules are used to determine whether a text fragment is a compliance rule based on the presence of obligatory modal verbs, conditional connectors, and executable actions. The conflict resolution rules are used to arbitrate priority when the same text fragment simultaneously matches two types of judgment rules. The structured output constraints are used to force the classification results to be output in machine-readable JSON format.
[0076] For example, the diversion prompts are explanatory text judgment rules, including: "major contracts" as referred to in these Measures refer to all types of transaction contracts with a contract value of more than 50 million yuan.
[0077] For example, the diversion prompts are directive text judgment rules, including: if the contract amount exceeds 50 million yuan, it should be reported to the group headquarters for approval.
[0078] For example, the flow-off prompt is a conflict resolution rule. If a text fragment satisfies two types of features at the same time, it is processed according to the following priority: if the fragment is mainly a definition with accompanying illustrative explanations, it is determined to be explanatory text; if the fragment is mainly an obligation with accompanying definition references, it is determined to be instructional text; if it cannot be clearly distinguished, it is marked as "AMBIGUOUS" and an error code is returned.
[0079] In some embodiments, step 203 may include: classifying text fragments with explanatory, delimiting, and declarative characteristics into explanatory text, and classifying text fragments with binding, directive, and conditional characteristics into directive text, based on preset triage prompts.
[0080] In some embodiments, step 203 may further include: dividing multiple text fragments into explanatory text and instructional text according to explanatory text determination rules, instructional text determination rules, conflict resolution rules and structured output constraints.
[0081] In some embodiments, after step 203, the above method may further include: inputting explanatory text into a terminology dictionary construction module to generate a terminology dictionary; and inputting instructional text into an execution step decomposition module as the sole input source for subsequent rule compilation.
[0082] In this way, physical isolation fundamentally prevents large language models from misjudging the objective definitions of terms as review rules that need to be enforced, greatly reducing the possibility of extraction errors.
[0083] Step 204: Standardize and parse the technical terms in the instruction text according to the technical terminology dictionary.
[0084] In the embodiments of this application, the terminology dictionary includes at least one term entry. Each term entry records the term name, alias, and standardized definition. The terminology dictionary is a read-only closed mapping table used for standardized parsing and consistency verification of terminology.
[0085] In some embodiments, step 204 may include: traversing the words in the instructional text according to the terminology dictionary; if a word is detected to match an “alias” or “term variant” in the terminology dictionary, replacing it with a “standard term name” defined in the dictionary; and binding the standardized term to the set of attributes defined in the dictionary to provide metadata support for subsequent triple generation.
[0086] For example, if expressions such as "related party," "related enterprise," or "holding company" appear in the directive text, they will be uniformly replaced with the standard term "related party" after standardized parsing; "related party" will be bound to the "contractor type" attribute under "contract entity," and its data type will be confirmed as "enumerated type."
[0087] Step 205: Based on the step prompts, convert the statements in the instruction text into multiple rule steps.
[0088] In this embodiment, step prompts are used to constrain the large model to decompose the instructional text into pseudocode-like execution steps that conform to predefined syntax rules. The predefined syntax rules include step type constraints, action whitelist constraints, and structured output constraints.
[0089] In some embodiments, the step prompt includes step type constraint information, action whitelist information, and output format constraint information; the step type constraint information is used to limit the text statement to a conditional step or a verification step; the action whitelist information limits the actions that can be used in the step to only the predefined standard actions in the action list; the output format constraint information forces the step to output in a structured pseudocode format; wherein, the conditional step is used to determine the applicable scenario of the rule, and the verification step is used to perform compliance judgment.
[0090] In the embodiments of this application, multiple rule steps include condition steps or verification steps.
[0091] In some embodiments, step 205 may include: identifying statements in the instructional text that contain logical judgment relationships but do not directly produce compliance conclusions as conditional steps, and identifying sentences containing explicit review objects, review actions, or review targets as verification steps, based on step prompt words.
[0092] For example, "If this contract is a material contract, the provisions of this chapter shall apply" is identified as a conditional step, and "The contract amount shall not exceed RMB 5 million" is identified as a verification step.
[0093] For example, the statement "If the contract amount exceeds 5 million yuan and the contracting party is a foreign-invested enterprise, it shall be submitted to the headquarters for approval" is identified as: Condition 1: Determine if the contract amount exceeds 5 million yuan; Condition Step 2: Determine whether the contracting party is a foreign-invested enterprise; Verification step 3: If both condition step 1 and condition step 2 are true, then it is determined that approval from headquarters is required.
[0094] For example, see Figure 3 The rule descriptions on the left represent statements in the instructional text, while the multiple rule steps on the right are conditional and validation steps transformed by LLM. Conditional steps determine whether the current rule applies to the current review scenario, and validation steps execute specific review actions if the rule applies. This shifts rule extraction from a "reading comprehension" mode to a "pseudocode generation" mode, breaking down complex rules into ordered logical execution steps, providing a clean and structured logical blueprint for the subsequent generation of rigorous abstract tree nodes.
[0095] Step 206: Verify whether the actions in multiple rule steps belong to the action list to obtain multiple execution steps.
[0096] In some embodiments, based on the action list, actions in multiple conditional steps or verification steps are mapped to predefined standard actions in the action list according to the timing sequence, resulting in multiple rule steps.
[0097] For example, "must not exceed" can be mapped to the negative form of "greater than", or "belong to" can be mapped to "belong to".
[0098] In some embodiments, the standard actions in the action list include numerical comparison actions, semantic judgment actions, and set judgment actions. Numerical comparison actions include greater than, less than, or equal to; semantic judgment actions include belonging to or meeting a condition; and set judgment actions include including or being included.
[0099] In some embodiments, the attributed action is used to determine whether the value of the object under review belongs to the enumerated value set defined by the review target; the conditional action is used to determine whether the text content of the object under review satisfies the rule expression or regular expression pattern defined by the review target; wherein, the enumerated value set and the rule expression are predefined by the entity list or the review target, rather than being freely generated by a large model.
[0100] In some embodiments, step 206 may include: traversing each of the multiple rule steps and verifying whether the action taken in that step fully matches a predefined standard action in the action list.
[0101] In some other embodiments, step 206 may further include: if the action belongs to the action list, determining multiple rule steps as multiple execution steps; if the action does not belong to the action list, triggering the regeneration of multiple rule steps until all actions belong to the action list, resulting in multiple execution steps.
[0102] In this way, the ambiguity of natural language is eliminated through the standardized parsing of the professional terminology dictionary; and the dual constraints of step prompts and action lists prevent the possibility of large models generating unexecutable actions such as "the first step is to check it", ensuring that the multiple rule steps output can be directly scheduled and executed by the underlying sandbox engine.
[0103] Step 207: Based on the entity list and action list, compile the step content of multiple execution steps to generate triplet machine instructions corresponding to each execution step.
[0104] In this embodiment, the entity list includes contract entities, technical rule entities, and process record material entities. Each entity corresponds to a preset set of legal attributes and a whitelist of data types.
[0105] In some embodiments, a contract entity is a structured data object corresponding to a legally binding agreement text identified from the document to be reviewed. Each entity possesses different attribute information, which is referred to as a "legal attribute set". For example, the legal attribute set corresponding to a contract entity includes contract amount, contracting parties, signing date, effective clause, and breach of contract clause. The data type whitelist corresponding to a contract entity includes numeric, string, date, or enumerated types.
[0106] In some embodiments, a technical rule entity is a structured data object corresponding to a document describing the technical specifications, functional requirements, or performance standards that a system, product, or project should meet. For example, the set of legal attributes corresponding to a technical rule entity includes technical indicator items, indicator thresholds, applicable environments, and testing methods, and the whitelist of data types corresponding to a technical rule entity includes numeric, string, date, or enumeration types.
[0107] In some embodiments, a process record material entity is a structured data object corresponding to a document that records the implementation, management, or acceptance process of a project. For example, the set of valid attributes corresponding to a process record material entity includes activity name, activity time, responsible party, activity results, and attachment list, and the whitelist of data types corresponding to a process record material entity includes numeric, string, date, or enumeration types.
[0108] In the embodiments of this application, the triplet machine instructions include a candidate, an action, and an objective.
[0109] In some embodiments, the triplet machine instruction can also be simply referred to as Review Object-Review Action-Review Objective (CAO). The triplet machine instruction is used to constrain "what to review," "how to review," and "what review standards to use" in each step. Due to the existence of this triplet, it is possible to ensure that each rule step has clear executableness.
[0110] In some embodiments, step 207 may include: for each execution step, extracting the entities and attributes to be examined from the execution step, and verifying whether the entities and attributes belong to the entity list; extracting action operators from the execution step, and verifying whether the action operators completely match the standard actions in the action list; extracting the benchmark or target value from the execution step, and compiling it to generate triplet machine instructions.
[0111] In one possible implementation, verifying whether entities and attributes belong to the entity list includes: verifying whether the review objects involved in this step belong to the legal entities and attributes defined in the entity list; and verifying whether each step contains the necessary fields.
[0112] For example, the process of compiling the step content of multiple execution steps based on the entity list and action list to generate the triplet machine instructions corresponding to each execution step is shown below.
[0113] 1. For the step “[Contract.Amount]>5 million”, extract “Contract” as an entity and “Amount” as an attribute; query the entity list to confirm that “Contract” is a predefined “Contract Entity” and “Amount” is one of its legal attributes, with the data type being “Decimal”; if the extracted entity or attribute is not in the entity list, it is determined to be a compilation error, triggering the step repair mechanism; 2. Extract the ">" symbol, query the action list, and map the ">" symbol to the standard action "greater than (GREATER_THAN)". If the extracted action is not in the action list, it is considered a compilation error. 3. If the target value is a fixed value, then directly convert the target value to a data type consistent with the reviewed object; if the target value is another entity attribute, then repeat the compilation process of the reviewed object to ensure that the reviewed object belongs to the entity list. 4. After the compilation process, the execution step "[Contract.Amount]>5 million" is converted into a triplet machine instruction.
[0114] For example, see Figure 4 , Figure 4 This includes transforming the review objects described in natural language into a machine-recognizable entity-attribute structure, and defining what kind of judgments or calculations each rule node can perform.
[0115] Figure 4 In the upper part, the entity represents the subject to which the rule applies, such as the "main contract"; the property represents the specific characteristics of the entity, such as "Party A's name," "total contract amount," and "breach of contract clauses"; and, Figure 4 This demonstrates how an entity contains multiple attributes, forming a tree structure. The Entity List defines all entity types existing in the system; for example, ENTITY_MAIN_CONTRACT represents the master contract, ENTITY_PROCESS_RECORD represents a business process record, and ENTITY_DELIVERABLE_SPEC represents a deliverable specification. The Data Type defines the data format for attribute values: STRING represents string text, NUMBER represents a uniform number, and DATE represents a date.
[0116] Figure 4The lower left section encapsulates review actions into three standard interfaces: Semantic (semantic judgment action), Compute (numerical comparison action), and List-List (set judgment action). Semantic handles semantic judgments of text or concepts, including SEMANTIC_BELONGS_TO (belongs to); when all verification steps pass, the system directs the process to the final node marked "SEMANTIC_COMPLY," indicating that the content under review conforms to the relevant specifications. Compute handles comparisons of numbers or dates, including EQUAL (standard numerical comparison operator) and GREATER_THAN (greater than, less than, or equal to). List-List handles inclusion or subset relationships between sets, including LIST_EXACTLY_MATCH (included) and LIST_IS_SUBSET_OF (included).
[0117] Figure 4 In the lower right section, the review objectives include CONSTANT (constant review objectives) and ENTITY (variable review objectives). For example, constant review objectives can be "500,000" or "2025-01-01", while variable review objectives can be attribute values of "contract counterparties".
[0118] Step 208: Generate multiple rule nodes based on triplet machine instructions.
[0119] Each triplet machine instruction corresponds to a rule node, and the data structure of the rule node conforms to a format that the sandbox engine can recognize.
[0120] For example, the data structure of a rule node can be in JSON Schema format.
[0121] In some embodiments, after step 208, the method further includes: serializing the generated multiple rule nodes according to a unified schema and storing them in a rule repository.
[0122] In this way, transforming unstructured natural language rules into a set of rules that are machine-understandable and executable, consisting of multiple rule nodes, helps to eliminate semantic ambiguity; and, through the dual constraints of entity lists and action lists, it reduces and curbs the divergent illusion of large models, thereby improving the accuracy of rule extraction.
[0123] Step 209: Configure branch jump logic for each rule node.
[0124] In this embodiment of the application, the branch jump logic includes the node flow direction determined according to the review result of the current rule node. The node flow direction includes: the next judgment node, the inapplicable final state node, the compliant final state node, or the non-compliant final state node.
[0125] In some embodiments, the next judgment node is used to point to the next rule node in the process to achieve continuous condition judgment or verification; the inapplicable final state node is used to indicate that the current rule is not applicable to the scenario to be reviewed, the process is terminated directly and no compliance evaluation is performed; the compliant final state node is used to indicate that after review, the current scenario fully complies with the specification requirements, and the process is terminated in a compliant state; the non-compliant final state node is used to indicate that after review, there is a violation in the current scenario, and the process is terminated in a non-compliant state.
[0126] In some embodiments, step 209 may include: configuring branch jump logic for each rule node based on the Boolean value of the review result. The Boolean value may be either true or false.
[0127] In some embodiments, if the review result of the current rule node is pass or fail, and if it is logically necessary to further determine other compliance elements, the node flow is to the next judgment node.
[0128] In some embodiments, if the review result of the current rule node indicates that the rule is not applicable to the current review scenario, the node flow is to an inapplicable final state node.
[0129] In some embodiments, if the review result of the current rule node indicates that all preset compliance conditions have been met, the node flow is to the compliance final state node.
[0130] In some embodiments, if the review result of the current rule node indicates that there are unmet compliance conditions or that a clear prohibition is met, the node flow is to a non-compliant final state node.
[0131] For example, if the current node verifies "whether the contract amount is greater than 5 million" and it is true, then jump to the next judgment node to verify "whether the contracting party is a foreign-funded enterprise"; if the current contract type is determined to be "personal gift contract", then it is not applicable to the "procurement compliance review rules" and directly jumps to the inapplicable final state node; if all verification steps return true values and there are no violations, then jump to the compliant final state node; if the key verification steps return false values or hit a clear prohibition, then jump to the non-compliant final state node.
[0132] Step 210: Determine the node jump relationship based on the node flow direction of each rule node.
[0133] In some embodiments, step 210 includes: traversing the branch jump logic of all rule nodes, extracting and parsing the node flow direction identifier of each rule node, and determining the jump relationship between nodes.
[0134] For example, define the on_true and on_false fields respectively to indicate the node flow identifier when the review result is "pass" or "fail".
[0135] In some embodiments, the string identifiers in the branch logic are parsed into the actual address of the rule node object or a unique identifier in memory through a global node index table; based on the parsing result, a one-way connection is established between nodes; three types of final state nodes are registered as the legal endpoints of the process and do not participate in subsequent node concatenation.
[0136] Step 211: Connect all rule nodes in series according to the node jump relationship to obtain a directed acyclic graph of review rules.
[0137] In this embodiment of the application, the directed acyclic graph has a unique entry node and multiple legal exit nodes.
[0138] In some embodiments, each edge in the directed acyclic graph has a definite direction, representing a unidirectional and irreversible flow of the review process; if a loop is detected during the execution of a series of steps, a rule reconstruction mechanism is triggered.
[0139] In some embodiments, step 211 may include: creating an empty vertex set and edge set, and storing all rule nodes in the vertex set; traversing each rule node in the vertex set and reading the identity (ID) of the subsequent node in the branch jump logic; if the ID of the subsequent node is not empty, adding a directed edge from the node to the subsequent node in the edge set; if the subsequent node is a predefined final state node, adding the final state node as a special vertex to the vertex set; after all edges are established, calling the topological sorting algorithm to check the graph structure; if the algorithm can successfully output a linear sequence of all nodes, it proves that the graph structure is an acyclic graph (DAG), and outputs the graph structure as a directed acyclic graph of the review rules; if a cycle is detected, an error is reported and the process construction is terminated to prevent the review logic from falling into an infinite loop.
[0140] For example, such as Figure 5 As shown, the topology on the left illustrates the logical relationships between the rule nodes. Figure 5Nodes 1 to 5 represent rule nodes. Each node encapsulates a triplet machine instruction, representing a specific review action. Two dashed arrows extend from node 1, marked as pass and fail, respectively, indicating branching logic. That is, when a rule node completes its execution, the process will flow down different paths based on its review result (true / false). Figure 5 There are three rectangles at the bottom left, representing inapplicable final state nodes, compliant final state nodes, and non-compliant final state nodes. The nodes on the right contain a triple consisting of the review object, review action, and review target. Review actions include Semantic (semantic analysis) and Compute (numerical calculation), and the sandbox engine also supports various types of review actions. Figure 5 The pass and fail nodes at the bottom right correspond to the pass and fail exits in the left diagram, respectively. This structured node design enables the topology diagram on the left to be transformed into machine instructions that the sandbox engine can recognize, thereby achieving automated execution of complex review logic.
[0141] In this way, by configuring branch jump logic for all rule nodes, multiple rule nodes that were originally discrete and isolated can be organically linked together to form a directed acyclic graph of review rules. Moreover, this directed acyclic graph clearly depicts the complete execution path from the rule entry point to the final review conclusion, ensuring the controllability, determinism and schedulability of the review process.
[0142] In some embodiments, after step 211, the above-described method for extracting rules from unstructured text further includes: Step 212: Perform a validity check on the triplet machine instruction or rule node.
[0143] In some embodiments, step 212 may include: confirming whether the entity and attribute referenced by the review object exist in the entity list and whether the data types match; if an illegal entity not defined in the entity list is referenced, or the attribute type does not match, the verification is deemed to have failed.
[0144] In some other embodiments, step 212 may further include: confirming whether the review action completely matches the action list; if an action not found in the action list is used, the verification is deemed to have failed.
[0145] In some other embodiments, step 212 may also include: confirming that all node flow identifiers point to valid target node IDs or final state node identifiers; and determining that the verification fails if there are dangling pointers or circular self-references.
[0146] Step 213: If the legality check fails, call the large model to repair the abnormal content.
[0147] In some embodiments, step 213 may include: if the legality check fails, using context-enhanced prompts to invoke a large model to repair the abnormal content.
[0148] In this embodiment of the application, the number of times the context-enhanced prompt words are used to call the large model to repair abnormal content does not exceed once.
[0149] In some embodiments, if the legality check fails, an error diagnosis report is generated, which records the error type, the field in which the error occurred, and the cause of the error. Based on the error diagnosis report, a repair prompt is constructed. The repair prompt locks the current entity list, action list, and terminology dictionary as inviolable hard constraints and issues precise repair instructions to the large model, so that the large model can only make local corrections to the illegal fields at the smallest granularity, rather than regenerating the entire rule.
[0150] For example, if the object under review in the triplet references the "Contract Importance" attribute which is not defined in the entity list, an error diagnosis report is generated, recording the error type as INVALID_ATTRIBUTE, and indicating that the attribute does not belong to the "Contract Entity".
[0151] In some embodiments, after step 213, the above method may further include: re-verifying the legality of the repaired node; if the verification passes, determining that the repair is successful and adding the node to the repository; if the legality verification fails, determining that the repair has failed.
[0152] Step 214: In the event of repair failure, regenerate the triplet machine instructions for the execution steps.
[0153] In some embodiments, step 214 may include: in the event of repair failure, backtracking to the original instructional text fragment corresponding to the rule node; clearing the previous erroneous decomposition results, while retaining global constraint information such as the terminology dictionary, entity list, and action list; and calling the large language model to re-decompose the text fragment into steps and compile the triples.
[0154] In some embodiments, the number of times the node is regenerated is no more than once; if the validity verification of the regenerated node fails, the generation of the rule is deemed to have failed, it is isolated and logged, and this does not affect the generation and deployment of other rules in the same batch.
[0155] In this way, by performing strict legality checks before operation, most low-level errors are intercepted outside the sandbox engine, ensuring the stability of the production environment; and by using the error correction capabilities of the large model itself for targeted repair, the cost of manual intervention is greatly reduced.
[0156] For example, see Figure 6 The rule document refers to unstructured text. After entering the LLM (Local Management Model), this unstructured text is broken down into three parallel processing streams: term definitions, action lists, and entity lists. Terminology definitions represent the definitions of proper nouns, enumerated values, and specific concepts involved in the extraction rules; these are essentially explanatory texts. Action lists represent the allowed operation types within the extraction rules, and entity lists represent the main objects involved. Rule descriptions represent the specific business logic extracted from the text; they can also be understood as statements in the instruction text. LLM combines rule descriptions with terminology definitions, action lists, and entity lists, breaking them down into specific execution steps and encapsulating multiple rule steps into independent rule nodes. Each node typically corresponds to a specific triple, representing a smallest granularity judgment unit. Finally, all rule nodes converge into the rule extraction result and are output. The extraction result is the source data subsequently compiled into an AST or directly input into the sandbox engine for execution.
[0157] In summary, in this embodiment, firstly, based on the extraction rules, the manually extracted rules are changed to automatically extracted rules from large models, reducing manual costs and improving extraction efficiency; secondly, by converting the rules into abstract syntax trees, the problem that general large models or manually extracted rules are too abstract and cannot be actually executed is solved; finally, by converting the rules into directed acyclic graphs, the problem that general rules are conceptually vague and prone to illusion is solved.
[0158] Figure 7 This is a block diagram of a rule extraction device for unstructured text provided in an embodiment of this application, such as... Figure 7 As shown, the rule extraction device 400 for unstructured text includes the following modules.
[0159] The text classification module 401 is used to classify the unstructured text to be processed into explanatory text and instructional text; the explanatory text is used to construct a dictionary of technical terms. The text decomposition module 402 is used to decompose the instruction text into multiple execution steps based on a professional terminology dictionary, preset step prompts and action lists; The node generation module 403 is used to generate triplet machine instructions and obtain multiple rule nodes based on a preset entity list, action list and multiple execution steps; The process generation module 404 is used to configure branch jump logic for each rule node according to the logical relationship between multiple execution steps, and to connect multiple rule nodes into a directed acyclic graph of review rules. The instruction compilation module 405 is used to compile the triplet machine instructions into an abstract syntax tree based on the triplet machine instructions corresponding to each rule node.
[0160] Optionally, the text classification module 401 is specifically used for: The text is analyzed by character recognition and layout analysis to obtain the original text content; unstructured text includes rule documents in the legal and financial fields. The original text content is sliced according to chapter structure, heading level, or semantic boundaries to obtain multiple text fragments; Based on preset triage prompts, multiple text fragments are divided into explanatory text and instructional text.
[0161] Optionally, the text decomposition module 402 is specifically used for: Based on a dictionary of technical terms, standardized parsing of technical terms in directive texts is performed. Based on the step prompts, convert the statements in the instructional text into multiple rule steps; Based on the action list, the conditional steps or validation steps are broken down into multiple rule steps; Verify whether the actions in multiple rule steps belong to the action list to obtain multiple execution steps.
[0162] Optionally, the node generation module 403 is specifically used for: Based on the entity list and action list, the step content of multiple execution steps is compiled to generate the triplet machine instructions corresponding to each execution step; Multiple rule nodes are generated based on triplet machine instructions; Each triplet machine instruction corresponds to a rule node, and the data structure of the rule node conforms to a format that the sandbox engine can recognize.
[0163] Optionally, the process generation module 404 is specifically used for: For each rule node, configure branch jump logic; the branch jump logic includes the node flow determined based on the review result of the current rule node, and the node flow includes: the next judgment node, the inapplicable final state node, the compliant final state node, or the non-compliant final state node; Determine the node jump relationship based on the node flow direction of each rule node; By connecting all rule nodes according to their jump relationships, a directed acyclic graph of review rules is obtained.
[0164] Optionally, the rule extraction device 400 for unstructured text also includes a model fallback module for: Perform legality verification on triplet machine instructions or rule nodes; If the validity check fails, the larger model is invoked to repair the abnormal content; If the repair fails, regenerate the triplet machine instructions for the execution steps.
[0165] Optionally, the action list includes numerical comparison actions, semantic judgment actions, and set judgment actions; the entity list includes contract entities, technical rule entities, and process record material entities; the triple machine instruction includes the review object, review action, and review target, and the review target includes constant review target or variable review target.
[0166] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0167] Figure 8 This is a structural block diagram of an electronic device according to an exemplary embodiment. For example... Figure 8 As shown, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the steps of the rule extraction method for unstructured text in the aforementioned embodiment.
[0168] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including instructions that can be executed by a processor of an electronic device to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0169] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described unstructured text rule extraction method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0170] This application also provides a computer program product, including a computer program that, when executed by a processor, implements a method for extracting rules from unstructured text.
[0171] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar parts between the various embodiments can be referred to each other.
[0172] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks of the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0174] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a predictive manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0176] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0177] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0178] The above is a detailed description of the preferred embodiments of this application. However, the invention of this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method for extracting rules from unstructured text, characterized in that, The method includes: The unstructured text to be processed is classified into explanatory text and instructional text; the explanatory text is used to construct a dictionary of technical terms. Based on the aforementioned terminology dictionary, preset step prompts, and action list, the instruction text is broken down to obtain multiple execution steps; Based on the preset entity list, the action list, and the multiple execution steps, a triplet machine instruction is generated and multiple rule nodes are obtained. Based on the logical relationship between the multiple execution steps, branch jump logic is configured for each rule node, and the multiple rule nodes are connected in series to form a directed acyclic graph of review rules; Based on the triplet machine instructions corresponding to each rule node, the triplet machine instructions are compiled into an abstract syntax tree.
2. The method according to claim 1, characterized in that, The process of classifying the unstructured text to be processed into explanatory text and instructional text includes: The unstructured text is subjected to character recognition and layout analysis to obtain the original text content; the unstructured text includes rule documents in the legal and financial fields. The original text content is sliced according to chapter structure, heading level, or semantic boundaries to obtain multiple text fragments; Based on preset triage prompts, the multiple text fragments are divided into explanatory text and instructional text.
3. The method according to claim 1, characterized in that, The process involves breaking down the instruction text into multiple execution steps based on the terminology dictionary, preset step prompts, and action lists, including: Based on the aforementioned terminology dictionary, the terminology in the instructional text is standardized and parsed. Based on the step prompts, the statements in the instruction text are converted into multiple rule steps; the rule steps include conditional steps or verification steps. Verify whether the actions in the multiple rule steps belong to the action list to obtain the multiple execution steps.
4. The method according to claim 1, characterized in that, The process of generating triplet machine instructions and obtaining multiple rule nodes based on a preset entity list, an action list, and multiple execution steps includes: Based on the entity list and the action list, the step content of the multiple execution steps is compiled to generate triplet machine instructions corresponding to each execution step; Multiple rule nodes are generated based on the triplet machine instructions; Each of the triplet machine instructions corresponds to a rule node, and the data structure of the rule node conforms to a format recognizable by the sandbox engine.
5. The method according to claim 1, characterized in that, The step of configuring branch jump logic for each rule node based on the logical relationship between the multiple execution steps, and connecting the multiple rule nodes into a directed acyclic graph of review rules, includes: For each rule node, branch jump logic is configured; the branch jump logic includes the node flow determined according to the review result of the current rule node, and the node flow includes: the next judgment node, the inapplicable final state node, the compliant final state node, or the non-compliant final state node; Determine the node jump relationship based on the node flow direction of each rule node; By connecting all rule nodes in a chain according to the node jump relationship, the directed acyclic graph of the review rules is obtained.
6. The method according to claim 1, characterized in that, After configuring branch jump logic for each rule node based on the logical relationship between the multiple execution steps, and concatenating the multiple rule nodes into a directed acyclic graph of review rules, the method further includes: Perform a validity check on the triplet machine instruction or the rule node; If the legality check fails, the larger model is invoked to repair the abnormal content; If the repair fails, the triplet machine instructions for the execution steps are regenerated.
7. The method according to any one of claims 1 to 6, characterized in that, The action list includes numerical comparison actions, semantic judgment actions, and set judgment actions; the entity list includes contract entities, technical rule entities, and process record material entities; the triple machine instructions include review objects, review actions, and review targets, and the review targets include constant review targets or variable review targets.
8. A rule extraction device for unstructured text, characterized in that, The device includes: The text classification module is used to classify the unstructured text to be processed into explanatory text and instructional text; the explanatory text is used to construct a dictionary of technical terms. The text decomposition module is used to decompose the instruction text into multiple execution steps based on the technical terminology dictionary, preset step prompts, and action list; The node generation module is used to generate triplet machine instructions and obtain multiple rule nodes based on the preset entity list, the action list and the multiple execution steps; The process generation module is used to configure branch jump logic for each rule node according to the logical relationship between the multiple execution steps, and to connect the multiple rule nodes into a directed acyclic graph of review rules. The instruction compilation module is used to compile the triplet machine instructions into an abstract syntax tree based on the triplet machine instructions corresponding to each rule node.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the rule extraction method for unstructured text as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the rule extraction method for unstructured text as described in any one of claims 1-7.