Agent-based policy automatic labeling method and device, and computer equipment
Patent Information
- Application Number
- CN202610726932.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-25
AI Technical Summary
首先是基于规则引擎与正则表达式的提取方法,这种方法虽然简单直接,但其泛化能力和维护成本限制了其应用范围和效果;其次是基于传统深度学习的信息抽取方法,如BiLSTM-CRF和BERT序列标注,这类方法虽然提升了语义理解能力,但在处理强逻辑、长距离依赖的政策文本时仍显不足;最后是基于大规模预训练语言模型的提示词提取方法,尽管利用了零样本或少样本学习的优势,但在实际应用中面临着严重的“幻觉”问题和格式不稳定的问题,难以满足高准确率的要求
[0033]本发明与现有技术相比的有益效果是:本发明通过宏观字段串行遍历调度器对待处理字段进行选择,加载其所有约束条件和标准枚举值集合以获取字段信息,并基于此信息从政策原文中精确提取字段值;随后,它审查这些字段值是否符合预设规则和要求,详细确定错误详情并分析之,结合政策原文上下文生成反思指导提示以便调整字段值提取策略,再次尝试提取与审查直至达到收敛条件。整个过程构建了一个自我校验与纠错机制,确保了对政策文本中复杂逻辑关系的理解与解析、字段值提取的完整性和准确性。这种方法不仅保证了高效性,还能够适应不断变化的政策环境,提供稳定可靠的服务,从而满足不同应用场景下的需求。
Smart Images

Figure CN122820108A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computers, and more specifically to a method, apparatus, and computer equipment for automated policy labeling based on intelligent agents. Background Technology
[0002] In the current context of digitalizing government services and intelligentizing enterprise services, the automated, structured parsing and annotation technologies for policy texts have become particularly important. To obtain government subsidies, tax breaks, or qualification certifications, enterprises need to accurately match policies that meet their specific conditions from a vast and complex slew of policy documents. However, policy texts are often in unstructured or semi-structured form, containing numerous long and complex sentences, implicit logical conditions, and technical jargon. This makes automatically converting policy texts into structured field tags, such as "industry to which the policy pertains," "maximum subsidy amount," and "enterprise honors and qualifications," a key technical challenge.
[0003] For structured extraction of policy texts, existing technologies mainly rely on three types of methods. First, there are extraction methods based on rule engines and regular expressions. While simple and direct, these methods limit their application scope and effectiveness due to their limited generalization ability and maintenance costs. Second, there are information extraction methods based on traditional deep learning, such as BiLSTM-CRF and BERT sequence labeling. Although these methods improve semantic understanding, they are still insufficient when dealing with policy texts with strong logic and long-distance dependencies. Finally, there are cue word extraction methods based on large-scale pre-trained language models. Although they utilize the advantages of zero-shot or few-shot learning, they face serious "illusion" problems and format instability issues in practical applications, making it difficult to meet high accuracy requirements.
[0004] Given the limitations of the aforementioned technologies, there is currently no technical solution that can perfectly solve the task of high-precision structured annotation of policy texts.
[0005] Therefore, it is necessary to design a new method that can understand and analyze the complex logical relationships in policy texts, has a self-verification and error correction mechanism, ensures the completeness and accuracy of information extraction, and provides stable and reliable services while ensuring efficiency and adapting to the ever-changing policy environment. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, apparatus and computer equipment for automated policy annotation based on intelligent agents.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: an agent-based automated policy annotation method, comprising:
[0008] The scheduler selects the field to be processed by serially traversing the macro fields, and loads all constraints and standard enumeration value sets of the field to be processed to obtain the field information;
[0009] Receive the original policy text and the name of the field to be extracted, and extract the field value from the original policy text based on the field information;
[0010] Review whether the field values conform to preset rules and requirements to determine the error details;
[0011] Analyze the error details and, in conjunction with the context of the original policy text, generate reflection guidance prompts;
[0012] Based on the reflection guidance, adjust the field value extraction strategy and extract the field values from the original policy text again, and perform the review to determine whether the field values meet the preset rules and requirements in order to determine the error details;
[0013] When the field extraction meets the convergence criteria, the field values are saved to a local file or database.
[0014] The further technical solution is as follows: after the field extraction reaches the convergence condition and the field value is saved to a local file or database, it also includes:
[0015] The scheduler selects the next field to be processed by serially traversing the macro fields, loads all constraints and standard enumeration value sets of the field to be processed to obtain field information, and executes the receiving policy text and the name of the field to be extracted. Based on the field information, the field value is extracted from the policy text until all fields to be processed are processed.
[0016] The further technical solution is as follows: the fields to be processed include enumeration fields, absolute value fields, interval value fields, and fuzzy semantic fields.
[0017] The further technical solution is as follows: the review of whether the field value conforms to preset rules and requirements to determine error details includes:
[0018] Based on the preset rules and requirements for each field, the field values are scored, and structured error localization is determined to obtain error details.
[0019] The further technical solution is as follows: The scoring of the field values based on preset rules and requirements for each field includes:
[0020] Indicative functions are used to score the field values based on preset rules and requirements for each field.
[0021] The present invention also provides an automated policy annotation device based on intelligent agents, comprising:
[0022] The loading unit is used to select the field to be processed by the macro-field serial traversal scheduler, and load all the constraints and standard enumeration value set of the field to be processed to obtain the field information;
[0023] The receiving and extraction unit is used to receive the original policy text and the name of the field to be extracted, and extract the field value from the original policy text based on the field information;
[0024] The review unit is used to review whether the field value conforms to preset rules and requirements in order to determine the error details;
[0025] The analysis unit is used to analyze the error details and, in conjunction with the context of the original policy text, generate reflection guidance prompts;
[0026] The adjustment unit is used to adjust the field value extraction strategy according to the reflection guidance prompts and extract the field values from the original policy text again, and perform the review to determine whether the field values meet the preset rules and requirements in order to determine the error details;
[0027] The storage unit is used to save the field value to a local file or database when the field extraction reaches the convergence condition.
[0028] Its further technical solutions include:
[0029] The next selection unit is used to select the next field to be processed by the macro field serial traversal scheduler, load all constraints and standard enumeration value sets of the field to be processed to obtain field information, and execute the received policy text and the name of the field to be extracted, and extract the field value from the policy text based on the field information until all fields to be processed have been processed.
[0030] The further technical solution is as follows: the review unit is used to score the field value based on the preset rules and requirements of each field, and determine the structured error location to obtain error details.
[0031] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0032] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0033] The advantages of this invention compared to existing technologies are as follows: This invention selects the fields to be processed through a macro-level field serial traversal scheduler, loads all constraints and standard enumeration value sets to obtain field information, and accurately extracts field values from the original policy text based on this information. Subsequently, it examines whether these field values meet preset rules and requirements, identifies and analyzes error details, generates reflective guidance prompts based on the context of the original policy text to adjust the field value extraction strategy, and attempts extraction and examination again until convergence conditions are met. The entire process constructs a self-verification and error correction mechanism, ensuring the understanding and parsing of complex logical relationships in policy texts and the completeness and accuracy of field value extraction. This method not only guarantees high efficiency but also adapts to the constantly changing policy environment, providing stable and reliable services to meet the needs of different application scenarios.
[0034] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0035] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart illustrating the agent-based automated policy annotation method provided in this embodiment of the invention;
[0037] Figure 2 A schematic block diagram of an agent-based automated policy annotation device provided in an embodiment of the present invention;
[0038] Figure 3 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0041] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0042] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0043] Please see Figure 1 , Figure 1 This is a flowchart illustrating the agent-based automated policy annotation method provided in this embodiment of the invention. The agent-based automated policy annotation method is applied to a server, using an intelligent scheduler to serially traverse macro-level fields, loading the constraints and standard enumeration value sets for each field to be processed, and extracting the corresponding field values from the original policy text. The extracted field values are reviewed and scored using preset rules. After determining the error details, reflection guidance prompts are generated based on the policy context, thereby adjusting the extraction strategy until convergence, ensuring the accuracy and completeness of information extraction. This process not only supports the processing of various field types (such as enumeration, numerical, range numerical, and fuzzy semantic fields) but also incorporates a built-in self-verification and error correction mechanism, enabling adaptive adjustments to different policy environments and providing efficient and stable information extraction services to meet ever-changing needs. After all fields have been processed, the results are saved to a local file or database, ensuring data persistence and facilitating subsequent processing. This method achieves the understanding and parsing of complex logical relationships in policy texts, maintaining efficiency while enhancing the system's flexibility and reliability.
[0044] Figure 1 This is a flowchart illustrating the agent-based automated policy annotation method provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S160.
[0045] S110. Select the field to be processed by the macro-field serial traversal scheduler, and load all constraints and standard enumeration value sets of the field to be processed to obtain the field information.
[0046] In this embodiment, field information refers to the set of all constraints and standard enumeration values of the field to be processed, which is used to guide the accurate extraction of the corresponding field values from the original policy text.
[0047] First, the scheduler selects the field to be processed by serially traversing the macro fields, and loads all constraints and standard enumeration values for this field (e.g., valid values for enterprise size can be "large", "medium", etc.). This step ensures that each field has clear rules to guide its processing, avoiding logical interference between different fields.
[0048] Specifically, the fields to be processed include enumerated fields, absolute numerical fields, interval numerical fields, and fuzzy semantic fields.
[0049] For policy matching scenarios, a highly rigorous and refined system of 19 structured fields was constructed. The values of these fields must theoretically be "standardized," which is the sole criterion for Evaluator's "0 / 1 binary evaluation." Based on the mathematical and logical characteristics of the fields, they are divided into four categories, as shown in Table 1 below.
[0050] Table 1. Fields to be processed
[0051] Enumeration field Company size, independent legal person status, listing status, honors and qualifications Absolute numeric fields Maximum subsidy amount Range numeric fields Previous year's operating revenue lower limit, previous year's operating revenue upper limit, previous year's operating revenue growth rate lower limit, previous year's operating revenue growth rate upper limit, previous year's R&D expenses lower limit, previous year's R&D expenses upper limit, previous year's R&D expense growth rate lower limit, previous year's R&D expense growth rate upper limit, previous year's fixed asset investment lower limit, previous year's fixed asset investment upper limit, previous year's fixed asset investment growth rate lower limit, previous year's fixed asset investment growth rate upper limit Fuzzy semantic class fields Policy application requirements and industry
[0052] Enumeration fields include:
[0053] Enterprise size: Only one or more labels from the four values of [Large], [Medium], [Small], and [Micro] are allowed to be extracted. In the policy application condition analysis, if there is no restriction, it will be returned as an empty value. Any non-standard enterprise size terms such as "above-scale", "small and micro enterprise", or "above-quota" are strictly prohibited.
[0054] Independent legal entity status: Select "Yes" or "No".
[0055] Listing status: Select "Yes" or "No".
[0056] Honors and qualifications: such as "High-tech Enterprise" and "Specialized, Refined and Innovative Little Giant" and other standard enterprise honors and qualifications labels. Policy texts may require one or more honors and qualifications labels.
[0057] Absolute numeric fields include:
[0058] Maximum subsidy amount: Extracted as a specific number (the unit of amount is uniformly converted to "ten thousand yuan", such as the original text "1 million yuan" output 100, the original text "100 million yuan" output 10000).
[0059] Range-based numeric fields include:
[0060] Such fields are extremely common in corporate policies, typically expressed as "operating revenue between 10 million and 50 million". To avoid the model extracting fuzzy strings, the method in this embodiment refines them into independent numerical node fields:
[0061] Operating revenue range: [Lower limit of operating revenue in the previous year], [Upper limit of operating revenue in the previous year] (unit is uniformly 10,000 yuan. For example, if the original text is "not less than 10 million yuan", then fill in 1,000 for the lower limit and leave the upper limit blank).
[0062] Operating revenue growth rate range: [Lower limit of operating revenue growth rate in the previous year], [Upper limit of operating revenue growth rate in the previous year] (unit is uniformly %. For example, "20%" in the original text is recorded as 20).
[0063] R&D expense range: [Lower limit of R&D expenses in the previous year], [Upper limit of R&D expenses in the previous year].
[0064] R&D expense growth rate range: [Lower limit of previous year's R&D expense growth rate], [Upper limit of previous year's R&D expense growth rate].
[0065] Fixed asset investment range: [Lower limit of fixed asset investment in the previous year], [Upper limit of fixed asset investment in the previous year].
[0066] The growth rate range for fixed asset investment is: [Lower limit of the growth rate of fixed asset investment in the previous year] and [Upper limit of the growth rate of fixed asset investment in the previous year].
[0067] Fuzzy semantic fields include:
[0068] Policy application requirements: Extract a standardized short text or logical expression text.
[0069] Industry classification: strictly extracted according to the national economic industry classification standards (e.g., "Manufacturing - Computer, Communication and Other Electronic Equipment Manufacturing").
[0070] S120: Receive the original policy text and the name of the field to be extracted, and extract the field value from the original policy text based on the field information.
[0071] In this embodiment, the system receives the original policy text and the name of the field to be extracted, and extracts the corresponding field value from the original policy text based on the loaded field information. Taking "enterprise size" as an example, if the original policy text mentions "this project focuses on supporting large-scale industrial enterprises within the jurisdiction", it attempts to extract the enterprise size label that conforms to predefined rules.
[0072] S130. Review whether the field value meets the preset rules and requirements to determine the error details.
[0073] In this embodiment, error details refer to a structured description of the specific error type and cause determined based on preset rules after evaluating the extracted field values.
[0074] Specifically, based on the preset rules and requirements for each field, the field values are scored, and structured error localization is determined to obtain error details.
[0075] Indicative functions are used to score the field values based on preset rules and requirements for each field.
[0076] Use the Evaluator module to review the extracted field values and check if they comply with preset rules and requirements. For non-compliant cases, the system will provide specific error details. For example, "above-scale enterprise" is not a valid standard enterprise size label and will therefore be marked as an error.
[0077] Error details include not only a simple score (e.g., 1 for correct, 0 for incorrect), but also detailed structured error localization, such as indicating whether the error type is due to reasons such as unconverted units or unnormalized synonyms.
[0078] In this embodiment, the reflection of traditional large models is often just "talking to oneself" at the natural language level, which cannot be precisely controlled in engineering. The absolute core of the method in this embodiment lies in: fully mathematicalizing and formulating the interaction among Actor, Evaluator, and Reflection, especially by introducing a strict 0 / 1 characteristic function, using mathematical formulas to quantify the evaluation results and determine the termination condition of the iteration.
[0079] Assume the current macro scheduler is processing the first... fields ( ), in the During the next iteration ( ),in This indicates the maximum number of iterations.
[0080] The actor generates extracted values based on the current context. Define the first... The field, the first The extracted value of the wheel : ;
[0081] Among them, PolicyText is the original policy text. Specify the current field name and the required data to be extracted from the field. Reflection guidance prompts generated from the previous round of Reflection (when) hour, You can directly set it to an empty character.
[0082] The Evaluator agent evaluates the Actor's output based on predefined rules and requirements for each field. Scoring is performed. The method in this embodiment explicitly states that the evaluation result is either correct (score of 1 point) or incorrect (score of 0 points), using an indicator function. Achieved with a fault tolerance rate of 0.
[0083] For enumeration type fields (such as...) (This corresponds to the enterprise size field and its content)
[0084] Set the standard enumeration whitelist set as such as enterprise size ={"Large", "Medium", "Small", "Micro", "Large;Medium;Small", "Large;Medium", "Medium;Small", "Medium;Small;Micro", ""}. Note: If the policy text does not mention specific information about enterprise size, please extract it directly as a null value (""). The specific mathematical expression is as follows: ;
[0085] For absolute numeric and range numeric fields (such as...) This corresponds to the maximum subsidy amount, or the lower limit of the previous year's operating revenue and related details.
[0086] For a numeric type to be considered, two conditions must be met simultaneously: first, the data type must be numeric (without string units); second, the numeric logic must conform to the semantics of the original text. A type checking characteristic function should be defined. and semantic correctness check indicator function The specific formula is as follows: Comprehensive evaluation score formula: ;
[0087] The Evaluator error details output function, when the score is 0, not only assigns a score but also outputs structured error location data for Reflection to use. Error Details The definition is as follows: For example: {"ErrorType": "A company above the designated size is not a standardized company size value, and it cannot be used as a rule for identifying company size.", "ActualValue": "A company above the designated size"}.
[0088] The Reflection policy generation function allows the Reflection agent to receive error details but not directly modify the values. Instead, it generates reflection guidance instructions for the Actor. The definition is as follows: ;
[0089] This function is implemented internally within the large language model as follows: combining the local context of the original policy text, it analyzes the root cause of the ErrorType (e.g., is it due to unconverted units? Unnormalized synonyms? Or incorrect extraction and identification of field values?), and outputs a corrective natural language guidance.
[0090] To avoid infinite loops, the maximum number of iterations in the single-field content extraction process needs to be set. (e.g., 5 rounds). Since a 0 / 1 evaluation is used, once the score is 1, it indicates that the field has reached absolute standardization and no further iteration is needed. The termination function of the convergence determination module is defined as follows: ;
[0091] The final corresponding number The logic for determining the output result of each field is as follows (including a fallback mechanism for error prevention): ;
[0092] The above formula means that if a score of 0 is still obtained after reaching the maximum number of rounds (e.g., 5 rounds), the system will absolutely not enter the erroneous value into the database. Instead, it will output "" and trigger a manual verification warning. This ensures the absolute cleanliness of the underlying data in the system.
[0093] S140. Analyze the error details and, in conjunction with the context of the original policy text, generate reflection guidance prompts.
[0094] In this embodiment, the reflection guidance prompt refers to the specific operational suggestions provided to the executor agent based on the error details analysis results, for correcting errors and improving the extraction strategy.
[0095] Once an error is detected, the Reflection module analyzes the error details in conjunction with the original policy text context to generate specific guidance prompts for reflection. These prompts aim to guide Actors on how to correct the error, for example pointing out that "enterprises above a certain size cannot be used as a basis for identifying enterprise size."
[0096] S150. Based on the reflection guidance, adjust the field value extraction strategy and extract the field values from the original policy text again, and perform the review to determine whether the field values meet the preset rules and requirements in order to determine the error details.
[0097] Based on the generated reflection guidance, the Actor adjusts its extraction strategy and attempts again to extract field values from the policy text, then repeats the review process of S130. This iterative cycle continues until the field values reach the preset accuracy standard or the maximum number of iterations is reached.
[0098] S160. When the field extraction reaches the convergence condition, the field value is saved to a local file or database.
[0099] When a field value reaches the convergence condition (i.e., meets all preset rules), it is saved to a local file or database to ensure data persistence and ease of subsequent processing.
[0100] This mechanism not only enables efficient parsing of policy texts, but also ensures the integrity and accuracy of information extraction through a rigorous self-verification and error correction process, adapting to the ever-changing policy environment and providing stable and reliable services.
[0101] In addition, it also includes:
[0102] The scheduler selects the next field to be processed by serially traversing the macro fields, loads all constraints and standard enumeration value sets of the field to be processed to obtain field information, and executes the receiving policy text and the name of the field to be extracted. Based on the field information, the field value is extracted from the policy text until all fields to be processed are processed.
[0103] In this embodiment, to address the shortcomings of existing technologies, the traditional paradigm of "one-time full extraction + overall self-reflection" is abandoned. Instead, an isolated architecture combining macroscopic field serial traversal scheduling with microscopic single-field independent closed loops is designed. Within the microscopic closed loop, three physically independent intelligent agent modules are strictly decoupled: Actor, Evaluator, and Reflection. In particular, a strict "0 / 1 binary absolute evaluation standard" is introduced into the Evaluator module, combined with mathematical quantification formulas to drive reflection iteration, ensuring that the output of each field fully complies with policy annotation standardization requirements.
[0104] In terms of deep architecture analysis, the macro-level field serial traversal scheduler employs a strict serial mechanism for each field in the schema, ensuring that the attention mechanism of the large model is fully focused on the extraction of a single target field, thus avoiding attention diversion. For example, after extracting the "Company Size" field and saving it to a local file, the next field, such as "Maximum Subsidy Amount," is extracted. This mechanism effectively isolates logical interference between different fields, allowing only one field to be processed at a time. In specific implementation, for an array containing 19 fields, the program internally uses a simple For loop to sequentially traverse and obtain each field, then performs an Actor-Evaluator-Reflection iteration process only on that field until the convergence condition is met.
[0105] The Actor, acting purely as an information extractor, receives the original policy text, the name of the field to be extracted, and modification guidance from the previous round of Reflection. It is not responsible for rule validation and focuses on outputting field values that meet the prompt requirements. The Evaluator, on the other hand, is a rigorous reviewer, specifically responsible for verifying whether the Actor's output values meet the specific rule requirements in the prompts. It only provides an absolute binary result: 1 point for correct and 0 points for incorrect. If a score of 0 is given, the error type must be explained in detail, but no modification suggestions are provided. Finally, the Reflection, acting as a senior business supervisor, is responsible for analyzing the reasons for the Evaluator's output score of "0," and, combined with the local context of the original policy text, provides the Actor with specific reflection guidance to help it improve and accurately extract the correct field values. This series of steps constitutes a single-field independent extraction and reflection iterative loop mechanism, achieving efficient automated policy annotation.
[0106] In this embodiment, based on the above mathematical formula, the complete lifecycle of each field (taking "enterprise size" as an example) is as follows:
[0107] Assuming the macro scheduler selects a field =“Enterprise Size”, load its corresponding condition constraints (enumeration type is {“Large”, “Medium”, “Small”, “Micro”, “Large;Medium;Small”, “Large;Medium”, “Medium;Small”, “Medium;Small;Micro”, “”}).
[0108] First round of Actor execution ( The Actor receives the original policy text: "This project focuses on supporting large-scale industrial enterprises within the jurisdiction," along with basic extraction instructions. Based on its own language comprehension capabilities, the Actor may incorrectly output "large-scale enterprises" as "large-scale enterprises."
[0109] Specifically: Using the system prompts from the Actor (executor) intelligent agent, and with the help of a large model, the value of the enterprise size in the phrase "This project focuses on supporting large-scale industrial enterprises within the jurisdiction" is "large-scale enterprises".
[0110] Step S3: First round of Evaluator evaluation ( )
[0111] The Evaluator uses the original policy text, extraction rules, and Actor extraction results. It was found that "above-scale enterprises" is not a legal standard enterprise name, and above-scale enterprises cannot be used as a basis for identifying enterprise size.
[0112] Specifically, using the system prompts from the Evaluator agent, and leveraging a large model, the output is directly:
[0113] {
[0114] "score": 0,
[0115] "error_type": "“Above-scale enterprise” is not a valid standard enterprise name, and “above-scale enterprise” cannot be used as a basis for identifying enterprise size."
[0116] },
[0117] Calculate according to the formula: ,therefore Here, This corresponds to a score of 0 in the system prompt word extraction results using the Evaluator agent.
[0118] The details of the output evaluation results are as follows:
[0119] {"ErrorType": "“Above-scale enterprise” is not a valid standard enterprise name, and “Above-scale enterprise” cannot be used as a basis for identifying enterprise size.", "ActualValue": "Above-scale enterprise"}.
[0120] Among them, ErrorType corresponds to the error_type field in the system prompt words extracted using the Evaluator intelligent agent, and ActualValue corresponds to the result extracted using the system prompt words extracted using the Actor intelligent agent: "above-scale enterprise".
[0121] Determining convergence: .
[0122] First round of reflection ( Reflection receives the original policy text and related rule information, and Actor extracts... as well as Generate reflection guidance prompts =“Enterprises above the designated size cannot be used as the basis for identifying the enterprise size field, therefore enterprises above the designated size cannot be directly extracted as the label for enterprise size, and empty characters can be directly output.”
[0123] Specifically, the system prompts are used with the Reflection agent, and then a large model is used to generate the reflection guidance prompts: "Enterprises above the designated size cannot be used as the basis for identifying the enterprise size field, so enterprises above the designated size cannot be directly extracted as the label of enterprise size, and empty characters can be directly output."
[0124] Second round of Actor execution ( The Actor now receives new context (including the original policy text, rule information, and reflection guidance), rereads the original text, and performs semantic transformation based on the guidance of Reflection. Output .
[0125] Specifically, using the system prompts from the Actor (executor) agent, the Reflection (reflector) agent outputs the prompt "Large-scale enterprises cannot be used as the basis for enterprise size field identification, therefore large-scale enterprises cannot be directly extracted as the label of enterprise size, and an empty character can be directly output", and then using the large model, the value of the enterprise size of "This project focuses on supporting large-scale industrial enterprises within the jurisdiction." is extracted as "".
[0126] Second round of Evaluator evaluation ( Evaluator assessment Calculate the score .
[0127] Specifically, using the system prompts from the Evaluator agent, and leveraging a large model, the output is directly:
[0128] {
[0129] "score": 1,
[0130] "error_type":""
[0131] },
[0132] Then determine convergence: The convergence requirement has been met; you can directly use "" as the field value for enterprise size.
[0133] Once the micro-loop for the "Enterprise Size" field is complete, the macro-scheduler writes "Enterprise Size":"" to the final result dictionary. This then triggers an independent loop for the next field, continuing until all 19 fields are completed.
[0134] Furthermore, through extremely sophisticated cue word design, it is ensured that the large model maintains a strict sense of boundary when playing the three roles of Actor, Evaluator, and Reflection, avoiding role confusion. Each agent has its own specific system cue words to guide its behavior.
[0135] For the Actor (executor) intelligent agent, the system prompt explicitly states that this role is a pure information extraction expert whose task is to accurately extract specified field values from policy texts. The Actor's behavior is strictly constrained, including being responsible only for information extraction and format conversion without performing any logical judgments or rule checks, adhering to the instructions in the [Reflection Guidelines], and ensuring that the output is concise and clear, limited to the final extracted values, and does not contain explanatory statements or other symbols.
[0136] The Evaluator agent is defined as a rigorous "data evaluation reviewer," possessing natural language understanding and rule-based check capabilities. Its primary responsibility is scoring and outputting error types; it is not allowed to offer modification suggestions or explain the reasons for errors. During the evaluation process, the Evaluator strictly adheres to the JSON format for output and employs a black-and-white scoring standard: 1 point for complete compliance and 0 points for any non-compliance. Furthermore, it verifies the evaluated value against a pre-defined set of validation rules to ensure compliance.
[0137] The Reflection agent acts as a seasoned "policy data analysis expert." Its task is not to directly extract data, but rather to analyze the reasons for errors made by the extractors and provide clear guidance on these mistakes. This guidance must be highly actionable, for example, by providing specific mapping dictionaries as transformation references. Importantly, the guidance provided by Reflection should be direct and rigorous, focusing on policy recommendations rather than the generation of final results, aiming to help the Actor improve the accuracy of its extraction.
[0138] In summary, through carefully designed system prompts, the method in this embodiment ensures clear boundaries between the three intelligent agent roles—Actor, Evaluator, and Reflection—each performing its specific function to jointly achieve an efficient and accurate automated policy annotation process. This mechanism not only improves the accuracy of data processing but also enhances the stability and reliability of the entire system.
[0139] For example, here is a sample prompt template style for the system's internal operation (sample prompt words):
[0140] The system prompt for the Actor (Executor) intelligent agent is as follows:
[0141] """
[0142] #
Role Definition
[0143] You are a pure "information extraction expert". Your task is to accurately understand the semantics and requirements of the provided policy text and extract the values of specified fields.
[0144] #
Behavioral Constraints
[0145] 1.You are only responsible for extraction and format conversion, and never make any logical judgment or rule verification by yourself.
[0146] 2.You must strictly follow the instructions in
Reflection Guidance
Reflection Guidance
[0147] 3.Your output must be extremely brief, only output the final value, do not output any explanatory statements, do not add quotation marks, do not carry punctuation marks.
[0148] #
Input Information
[0149] <Original Policy Text>
[0150] {Dynamically insert full text of the policy or relevant paragraphs}
[0151] < / Original Policy Text>
[0152] <Target Field>
[0153] Enterprise scale
[0154] < / Target Field>
[0155] <Reflection Guidance>
[0156] {Dynamically insert P_{i,t-1}, which is "" in the first round}
[0157] < / Reflection Guidance>
[0158] #
Output Format
[0159] Directly output the extracted value, for example: large
[0160] """
[0161] The system prompt of the Evaluator agent is as follows:
[0162] """
[0163] #
Role Definition
[0164] You are a strict "data evaluation reviewer". You have natural language semantic understanding ability and can automatically check based on the rule table.
[0165] #
Behavioral Constraints
[0166] 1.You are prohibited from putting forward any modification suggestions! You are prohibited from explaining why it is wrong! Your only tasks are scoring and outputting error types.
[0167] 2. The output must be in strict JSON format, and no other characters are allowed.
[0168] 3. Your scoring standard is black or white: 1 point is given for full compliance with the rules, 0 point is given for any slight non-compliance. There is no intermediate score.
[0169] #
Check Rule Base
[0170] Field name: Enterprise Scale
[0171] Allowed value set: ["Large","Medium","Small","Micro","Large;Medium;Small","Large;Medium","Medium;Small","Medium;Small;Micro","Small;Micro",""]
[0172] Data type requirement: String
[0173] #
Input Information
[0174] <Original Policy Text>
[0175] {Dynamically insert the full text of the policy or relevant paragraphs}
[0176] < / Original Policy Text>
[0177] <Value to be Evaluated>
[0178] {Dynamically insert v_{i,t}, for example: "above-scale enterprise"}
[0179] < / Value to be Evaluated>
[0180] #
Evaluation Logic Execution Flow
[0181] Step 1: Check whether the value to be evaluated is exactly equal to an element in the whitelist (note that it is an exact match, case-sensitive, and no extra spaces are allowed).
[0182] Step 2: If it is in the whitelist, the score is 1, and the error type is "".
[0183] Step 3: If it is not in the whitelist, or contains characters other than those in the whitelist, the score is 0, and a specific and accurate error type description (error_type) shall be given.
[0184] #
Output Format
[0185] Strictly output the following JSON:
[0186] {
[0187] "score":0,
[0188] "error_type":""
[0189] }
[0190] """
[0191] The system prompt of the Reflection agent is as follows:
[0192] """
[0193] # [Role Definition]
[0194] You are a senior "policy data analysis expert". You do not directly extract data. Your task is to analyze why the "extractor" made a mistake, and write clear "error guidance" for him.
[0195] # [Behavior Constraints]
[0196] 1. Your guidance must be extremely executable, and clearly tell him how to convert (e.g., provide a mapping dictionary).
[0197] 2. Do not generate the final extraction result by yourself, only provide guidance strategies.
[0198] 3. The tone must be severe and direct, no nonsense.
[0199] # [Input Information]
[0200] <Local policy context>
[0201] {Dynamically insert the original paragraph containing the field information}
[0202] < / Local policy context>
[0203] <Extractor's output result>
[0204] {Dynamically insert v_{i,t}, for example: "enterprises above designated size"}
[0205] < / Extractor's output result>
[0206] <Quality inspector's error report>
[0207] {Dynamically insert E_{i,t}, for example: score 0, error type ErrorType is xxx}
[0208] < / Quality inspector's error report>
[0209] # [Analysis Task]
[0210] The extractor made the mistake of "regarding enterprises above designated size as the identification basis for enterprise scale".
[0211] What we require is: "enterprises above designated size" cannot be used as a basis for identifying enterprise scale, and enterprise scale cannot be extracted in this case.
[0212] #
Output Format
[0213] <Reflection Guidance>
[0214] "Enterprises above designated size" is not a legal standard enterprise name, and at the same time, it cannot be used as a basis for identifying enterprise scale.
[0215] < / Reflection Guidance>
[0216] """
[0217] Compared with the prior art, the method of this embodiment has significant advantages and beneficial effects, which are mainly reflected in the following aspects:
[0218] First, the method of this embodiment completely solves the common "attention dilution" problem during multi-field concurrent extraction, and maximizes the extraction depth of a single field. By adopting the architecture design of "macro serial traversal scheduling + single-field independent closed-loop", it ensures that all attention resources of the large model only focus on one specific field at any time, for example, "enterprise scale". This strategy not only eliminates the interference caused by multi-objective processing, but also ensures that each field can obtain the deepest semantic analysis and high-precision results.
[0219] Second, the "execution-evaluation-reflection" separation of powers architecture proposed by the method of this embodiment, combined with "0 / 1 binary evaluation", greatly improves the quality of automatic data annotation by large models. Different from the traditional method that allows the same large model to conduct self-verification, the method of this embodiment physically isolates these three capabilities into three independent Agents: Actor focuses on information extraction, Evaluator strictly finds errors according to rules and uses an all-or-nothing evaluation standard, and Reflection is responsible for error cause analysis and guidance compilation. This method enables the system to truly implement adversarial error correction and improve the accuracy of data annotation.
[0220] Furthermore, the introduced 0 / 1 absolute evaluation mechanism ensures the complete controllability of industrial-level system data. This mechanism accurately identifies successful cases and exits the agent cycle immediately, and triggers manual early warning through the maximum round limit when encountering intractable errors, preventing dirty data from flowing into the downstream database. This not only improves the accuracy of the system, but also effectively controls the inference cost.
[0221] Furthermore, the method in this embodiment implements "soft coding" of complex business logic through a reflection mechanism, enhancing the maintainability of the system. Faced with complex normalization requirements in policy texts, traditional hard-rule methods require the maintenance of numerous regular expressions, while the method in this embodiment allows for dynamic updates of mapping logic via natural language instructions, without modifying the underlying code. This approach significantly shortens the system's iterative maintenance cycle from weeks to minutes, greatly improving the system's flexibility and responsiveness.
[0222] Finally, the core innovations of this embodiment are concentrated in three aspects: first, a decoupled intelligent agent scheduling architecture based on a combination of single-field micro-closed-loop and macro-serial traversal; second, a "0 / 1 binary absolute evaluation" reflection mechanism based on the separation of powers of Actor-Evaluator-Reflection; and third, a standardization and reflective iterative safe convergence method for strongly constrained policy texts. These innovations provide strong underlying technical support for the system, build a solid technical barrier, and achieve efficient and accurate automated labeling of policy data.
[0223] Specifically, a decoupled agent scheduling architecture is proposed, which integrates a micro-level closed loop for single fields with a macro-level serial traversal. This innovative two-level scheduling architecture performs serial traversal and attention isolation on 19 heterogeneous fields, including interval values, enumerations, and text, at the macro level. At the micro level, it establishes an independent "extraction-evaluation-reflection" lifecycle closed loop for each individual field. Through this design, the architecture fundamentally eliminates the attention decay and logical cross-interference problems commonly encountered by large language models when handling complex tasks with multiple constraints, ensuring the purity of the single-field extraction logic.
[0224] Furthermore, the method in this embodiment introduces a "0 / 1 binary absolute evaluation" reflection mechanism based on the separation of powers among Actor-Evaluator-Reflection. This mechanism overcomes the limitations of existing agent systems where a single Prompt performs multiple tasks and model self-evaluation has ambiguity, constructing a three-agent collaborative model with absolutely isolated responsibilities within a single-field closed loop. Specifically for the Evaluator module, the method in this embodiment creatively adopts a binary evaluation formula based on indicator functions, forcibly discretizing the evaluation result into 1 (completely correct) or 0 (incorrect), thereby completely depriving the large model of its "discretionary power" in the rule verification process. Combined with the root cause analysis and policy generation functions of the Reflection module, a rigorous digital closed loop is formed, from "0-point alarm" to "policy guidance," then to "directional correction," and finally achieving "1-point pass."
[0225] Finally, the method in this embodiment provides a standardized and reflective iterative safe convergence method for policy texts with strong constraints. This method abandons the traditional fuzzy evaluation approach that relies on the subjective judgment of "whether it is correct" by a large model. Instead, it combines a 0 / 1 evaluation formula to construct a triple safe convergence function, including "early victory termination (immediate interruption when the score is 1)," "forced termination after the maximum number of rounds (to prevent getting stuck in an infinite loop)," and "failure fallback warning (outputting all scores of 0 and intercepting them for storage)." This calculation method is the first to achieve the engineering implementation of the Reflection mechanism for automated annotation of policy data under the stringent data quality requirements of government-level data, solving the problem of automatically mapping non-standardized policy terminology to a standard database schema and ensuring the accuracy of the output results. These innovations together constitute the core technical barrier of the method in this embodiment, providing strong underlying technical support.
[0226] In this embodiment, to implement the agent-based automated policy annotation method, the first step is field selection and information loading. During this process, the system identifies and determines the specific fields to be processed through a macro-level field serial traversal scheduler. This requires a deep understanding and analysis of the policy text structure to ensure accurate selection of all relevant parts. For each selected field, the system loads all its constraints and standard enumeration value sets, which are predefined based on domain knowledge, historical data, or specific rules. This step lays the foundation for subsequent field value extraction.
[0227] The next step is receiving and parsing the original policy text. In this step, the system receives the original text containing the policy content and identifies the fields that need to be extracted. This may involve the application of natural language processing (NLP) technology to more accurately understand the actual meaning of the policy text. Once the field names to be extracted are identified, the system will precisely extract the corresponding field values from the original policy text based on previously loaded field information. This process requires a high degree of accuracy and attention to detail to ensure that the extracted information truly reflects the content of the policy. Specifically, after receiving the original text containing the policy content, the system first needs to perform preliminary processing and analysis. The key to this stage is using natural language processing (NLP) technology to understand the content structure of the policy text and identify key information. The specific steps are as follows:
[0228] Text preprocessing: This is the first step in the entire parsing process, mainly including removing irrelevant characters and punctuation marks, and converting to a uniform format to facilitate subsequent processing. In addition, word segmentation is performed to divide continuous text into manageable smaller units.
[0229] Field name matching: Based on the previously loaded field information, the system uses predefined keyword or pattern matching rules to locate relevant fields in the original policy text. This step may involve complex string matching algorithms or the application of machine learning models to improve matching accuracy.
[0230] Contextual understanding and semantic analysis: To extract field values more accurately, the system relies not only on surface-level text matching but also on a deep understanding of the context of the policy text. For example, semantic role labeling (SRL) technology can identify the relationships between components in a sentence, thereby better grasping the specific meaning of the field values.
[0231] Field value extraction: Once the location and contextual meaning of a specific field are determined, the system can begin extracting the corresponding field value. This process may utilize technologies such as regular expressions and Named Entity Recognition (NER) to ensure that the extracted information is both comprehensive and accurate.
[0232] Validation and Adjustment: The extracted field values must undergo a series of validation steps to ensure they conform to the previously loaded constraints and standard enumeration value set. If any inconsistencies are found, the system will automatically adjust the extraction strategy according to preset rules, or provide prompts for manual intervention and correction.
[0233] Results logging: Finally, all validated field values will be logged as part of the final output. This data can be further used to generate reports, support decision-making, or for other purposes.
[0234] After the initial extraction of field values, the next step is to review them. The goal of this stage is to comprehensively check the extracted field values to confirm that they conform to the preset rules and requirements. This includes, but is not limited to, verification of syntactic correctness, format consistency, and logical coherence. If any non-compliance is found, the system will record detailed error information, such as the specific error type and location. This step is crucial for ensuring the quality of the final result because it allows for the timely detection and correction of potential problems.
[0235] Once error details are logged, the system generates reflective guidance prompts to help improve field value extraction strategies. These prompts are generated based on known error details and the context of the original policy text, aiming to provide specific suggestions, such as adjusting keyword matching algorithms or optimizing the use of regular expressions. This feedback mechanism helps the system continuously learn and improve, thereby enhancing its ability to handle similar situations in the future.
[0236] The next step involves adjusting the strategy based on the generated reflection guidance and retrying to extract field values. Here, the system dynamically adjusts the previous extraction strategy and attempts to extract field values from the original policy text again. This review process is then repeated until the field values meet all preset rules and requirements, or reach predetermined convergence conditions, such as no new errors being generated after several consecutive iterations. This iterative process ensures that the final field values have high quality and accuracy.
[0237] The aforementioned agent-based automated policy annotation method selects fields to be processed through a macro-level field serial traversal scheduler, loads all constraints and standard enumeration value sets to obtain field information, and accurately extracts field values from the original policy text based on this information. Subsequently, it examines whether these field values conform to preset rules and requirements, identifies and analyzes error details, generates reflective guidance prompts based on the policy text context to adjust the field value extraction strategy, and re-attempts extraction and review until convergence is achieved. The entire process constructs a self-verification and error-correction mechanism, ensuring the understanding and parsing of complex logical relationships in policy texts and the completeness and accuracy of field value extraction. This method not only guarantees high efficiency but also adapts to constantly changing policy environments, providing stable and reliable services to meet the needs of different application scenarios.
[0238] Figure 2 This is a schematic block diagram of an automated policy annotation device 300 based on an intelligent agent, provided in an embodiment of the present invention. Figure 2 As shown, corresponding to the above-described agent-based automated policy annotation method, the present invention also provides an agent-based automated policy annotation apparatus 300. This agent-based automated policy annotation apparatus 300 includes a unit for executing the above-described agent-based automated policy annotation method, and the apparatus can be configured in a server. Specifically, please refer to... Figure 2 The agent-based policy automated annotation device 300 includes a loading unit 301, a receiving and extraction unit 302, a review unit 303, an analysis unit 304, an adjustment unit 305, and a storage unit 306.
[0239] The loading unit 301 is used to select the field to be processed by the macro-field serial traversal scheduler, and load all constraints and standard enumeration value sets of the field to be processed to obtain field information; the receiving and extraction unit 302 is used to receive the original policy text and the name of the field to be extracted, and extract the field value from the original policy text based on the field information; the review unit 303 is used to review whether the field value meets the preset rules and requirements to determine the error details; the analysis unit 304 is used to analyze the error details, and generate reflection guidance prompts in combination with the context of the original policy text; the adjustment unit 305 is used to adjust the field value extraction strategy according to the reflection guidance prompts and extract the field value from the original policy text again, and perform the review of whether the field value meets the preset rules and requirements to determine the error details; the saving unit 306 is used to save the field value to a local file or database when the field extraction reaches the convergence condition.
[0240] In one embodiment, it further includes:
[0241] The next selection unit is used to select the next field to be processed by the macro field serial traversal scheduler, load all constraints and standard enumeration value sets of the field to be processed to obtain field information, and execute the received policy text and the name of the field to be extracted, and extract the field value from the policy text based on the field information until all fields to be processed have been processed.
[0242] In one embodiment, the review unit 303 is used to score the field value based on preset rules and requirements for each field, and determine the structured error location to obtain error details.
[0243] In one embodiment, the review unit 303 is used to score the field value using an indicative function based on preset rules and requirements for each field.
[0244] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned agent-based policy automated labeling device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0245] The aforementioned agent-based automated policy annotation device 300 can be implemented as a computer program, which can, for example... Figure 3 It runs on the computer device shown.
[0246] Please see Figure 3 , Figure 3This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0247] See Figure 3 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0248] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an agent-based automated policy labeling method.
[0249] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0250] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an agent-based automatic policy labeling method.
[0251] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0252] The processor 502 is used to run a computer program 5032 stored in a memory to implement all the steps of the agent-based policy automated labeling method.
[0253] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0254] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0255] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform all steps of an agent-based automated policy labeling method.
[0256] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0257] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0258] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0259] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0260] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0261] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An automated policy annotation method based on intelligent agents, characterized in that, include: The scheduler selects the field to be processed by serially traversing the macro fields, and loads all constraints and standard enumeration value sets of the field to be processed to obtain the field information; Receive the original policy text and the name of the field to be extracted, and extract the field value from the original policy text based on the field information; Review whether the field values conform to preset rules and requirements to determine the error details; Analyze the error details and, in conjunction with the context of the original policy text, generate reflection guidance prompts; Based on the reflection guidance, adjust the field value extraction strategy and extract the field values from the original policy text again, and perform the review to determine whether the field values meet the preset rules and requirements in order to determine the error details; When the field extraction meets the convergence criteria, the field values are saved to a local file or database.
2. The agent-based automated policy annotation method according to claim 1, characterized in that, After the field extraction reaches the convergence condition and the field value is saved to a local file or database, the process further includes: The scheduler selects the next field to be processed by serially traversing the macro fields, loads all constraints and standard enumeration value sets of the field to be processed to obtain field information, and executes the receiving policy text and the name of the field to be extracted. Based on the field information, the field value is extracted from the policy text until all fields to be processed are processed.
3. The agent-based automated policy annotation method according to claim 1, characterized in that, The fields to be processed include enumerated fields, absolute numerical fields, range numerical fields, and fuzzy semantic fields.
4. The agent-based automated policy annotation method according to claim 1, characterized in that, The review of whether the field value conforms to preset rules and requirements to determine error details includes: Based on the preset rules and requirements for each field, the field values are scored, and structured error localization is determined to obtain error details.
5. The agent-based automated policy annotation method according to claim 1, characterized in that, The scoring of the field values based on preset rules and requirements for each field includes: Indicative functions are used to score the field values based on preset rules and requirements for each field.
6. An automated policy annotation device based on intelligent agents, characterized in that, include: The loading unit is used to select the field to be processed by the macro-field serial traversal scheduler, and load all the constraints and standard enumeration value set of the field to be processed to obtain the field information; The receiving and extraction unit is used to receive the original policy text and the name of the field to be extracted, and extract the field value from the original policy text based on the field information; The review unit is used to review whether the field value conforms to preset rules and requirements in order to determine the error details; The analysis unit is used to analyze the error details and, in conjunction with the context of the original policy text, generate reflection guidance prompts; The adjustment unit is used to adjust the field value extraction strategy according to the reflection guidance prompts and extract the field values from the original policy text again, and perform the review to determine whether the field values meet the preset rules and requirements in order to determine the error details; The storage unit is used to save the field value to a local file or database when the field extraction reaches the convergence condition.
7. The agent-based automated policy annotation device according to claim 6, characterized in that, Also includes: The next selection unit is used to select the next field to be processed by the macro field serial traversal scheduler, load all constraints and standard enumeration value sets of the field to be processed to obtain field information, and execute the received policy text and the name of the field to be extracted, and extract the field value from the policy text based on the field information until all fields to be processed have been processed.
8. The agent-based automated policy annotation device according to claim 6, characterized in that, The review unit is used to score the field values based on the preset rules and requirements for each field, and to determine the structured error location in order to obtain error details.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 5.
10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 5.