Verification storage method and device for output content of intelligent agent, equipment and medium
Patent Information
- Application Number
- CN202611040229.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-14
AI Technical Summary
[0003]然而,不同任务类型对输出结构的要求存在差异,模型输出的随机性又导致字段缺失、来源标识错误、关键字段被改写等问题在实际部署中频繁出现
[0017]第四方面,本申请实施例还提供了一种计算机可读存储介质,存储有计算机可执行指令,所述计算机可执行指令用于执行如第一方面所述的对于智能体输出内容的校验存储方法。
Smart Images

Figure CN122570208B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of intelligent agent output verification technology, and in particular to a method, apparatus, device, and medium for verifying and storing the output content of an intelligent agent. Background Technology
[0002] As large language models (agents) expand from dialogue generation to engineering tasks such as document analysis, code review, and log processing, the consumers of model output have shifted from humans to downstream programs and persistent storage systems. In document compression, field names and value ranges need to be preserved; in code change summaries, module names, file paths, and items to be verified need to be recorded; and in log summarization, event levels, time windows, and location objects need to be preserved—in these scenarios, model output needs to be directly written to object storage or used as structured input for subsequent tasks.
[0003] However, different task types have different requirements for output structure, and the randomness of model output often leads to problems such as missing fields, incorrect source identification, and rewriting of key fields in actual deployments. Without a structured verification mechanism independent of the model, the stability of the engineering process and the traceability of the results are difficult to guarantee. In addition, existing solutions lack an independent pre-write verification step between the large language model and object storage, and do not take the integrity of references between output objects and object storage as a pre-write constraint. Summary of the Invention
[0004] This application provides a method, device, and medium for verifying and storing the output content of an intelligent agent, which can achieve accurate verification of the output of a large language model.
[0005] In a first aspect, embodiments of this application provide a method for verifying and storing the output content of an intelligent agent, applied to a verification middleware, wherein the verification middleware is a component independent of the large language model process, and the method includes: Receive the first output content of the large language model; The target JSON Schema template is determined based on the first output content, and the target JSON Schema template is loaded. The target JSON Schema template or equivalent constraint description information is sent to the large language model to constrain the output of the large language model; the equivalent constraint description information is generated based on the target JSON Schema template. Obtain the second output content of the large language model after constraint processing; The second output content is parsed using JSON. If the parsing is successful, the parsed second output content is then validated using JSON Schema, foreign key constraints, and field comparison. If the JSON Schema validation, the foreign key constraint validation, and the field comparison validation all pass, the parsed second output content is written to object storage.
[0006] In some embodiments, determining the target JSON Schema template based on the first output content includes: Based on the first output content, obtain the task intent, input source type, or upstream process identifier; Identify the target object type based on the task intent, the input source type, or the upstream process identifier; The target JSON Schema template is determined based on the target object type.
[0007] In some embodiments, the target object type includes at least one of a document summary object, a table field object, a code module object, and a log event object; determining the target JSONSchema template based on the target object type includes: When the target object type is the document summary object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a first required field, non-empty summary content, keywords and items to be processed being an array, and the existence of a source object identifier; When the target object type is the table field object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a second required field, the complete field name and value range, and the field name being consistent with the source table row. When the target object type is the code module object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a third required field, the correct module name, file path, and change point type, and consistency with the source source code object; When the target object type is the log event object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a fourth required field, the event level, time window, and summary meeting preset ranges, and the existence of a source object identifier.
[0008] In some embodiments, after the step of performing JSON parsing processing on the second output content, the method further includes: If parsing fails and the second output cannot be parsed into valid JSON, it is recorded as a failed attempt. The large language model is triggered to regenerate new second output content until the new second output content is parsed as valid JSON or the number of failed attempts reaches a preset value.
[0009] In some embodiments, the step of performing JSON Schema validation, foreign key constraint validation, and field comparison validation on the parsed second output content includes: If any of the checks fails, record the reason for the failure and increment the number of failed attempts by one. New constraints are generated based on the reasons for failure and sent to the large language model. The large language model is then triggered to generate new second output content based on the new constraints until the number of failed attempts reaches the preset value or the new second output content passes the JSON Schema validation, foreign key constraint validation, and field comparison validation.
[0010] In some embodiments, when the number of failed attempts reaches the preset value, the method further includes: Obtain the source object that has not been processed by the large language model; Based on preset rules, the corresponding fields are extracted from the source object to generate a structured object; Perform JSON Schema validation and foreign key constraint validation on the structured object. If the validation passes, write the structured object to the object storage.
[0011] In some embodiments, the JSON Schema validation includes: Load the target JSON Schema template, and use the canonical keywords of the target JSON Schema template to perform structural validation on the parsed second output content; If all the specified keywords pass the validation, the parsed second output content passes the JSONSchema validation. If any of the specified key validations fails, the parsed second output content will not pass the JSONSchema validation.
[0012] In some embodiments, the foreign key constraint verification includes: Obtain the source_oids field from the parsed second output content; the source_oids field is a set of foreign keys pointing to the source object. For each foreign key in the source_oids field, query the object storage to see if there is a corresponding object primary key; If each of the foreign keys has a corresponding object primary key in the object storage, then the parsed second output content passes the foreign key constraint verification. If any of the foreign keys cannot be found in the object storage as a corresponding primary key, then the parsed second output content fails the foreign key constraint validation.
[0013] In some embodiments, the field comparison and verification includes: Based on the pre-declared key field set corresponding to the target object type, the key field values in the parsed second output content are compared with the corresponding field values in the source object pointed to by each foreign key, or a normalized comparison is performed. If all key field values are successfully matched, the parsed second output content will pass the field comparison verification. If any of the key field values fails to match, the parsed second output content will not pass the field matching check.
[0014] In some embodiments, the method further includes: When writing the parsed second output content into the object storage, a reverse index is established between the source object identifier and the object primary key corresponding to the second output content. The reverse index is used to locate and trace the source object based on the object primary key.
[0015] Secondly, embodiments of this application provide a verification and storage device for the output content of an intelligent agent, comprising: The receiving module is used to receive the first output content of the large language model; A loading module is used to determine a target JSON Schema template based on the first output content and to load the target JSON Schema template. The constraint module is used to send the JSON Schema template or equivalent constraint description information to the large language model to perform constraint processing on the output of the large language model; the equivalent constraint description information is generated based on the JSON Schema template. The acquisition module is used to acquire the second output content of the large language model after constraint processing. The validation module is used to perform JSON parsing on the second output content. If the parsing is successful, the parsed second output content is validated for JSON Schema, foreign key constraints, and field comparison. The writing module is used to write the parsed second output content into object storage if the JSON Schema validation, the foreign key constraint validation, and the field comparison validation all pass.
[0016] Thirdly, embodiments of this application provide an electronic device, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the verification and storage method for the output content of an intelligent agent as described in the first aspect.
[0017] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for performing the verification and storage method for the output content of an intelligent agent as described in the first aspect.
[0018] This application provides a method, apparatus, system, and medium for verifying and storing the output content of an intelligent agent, which has at least the following beneficial effects: By setting up a verification middleware independent of the large language model and using the verification middleware to perform multiple verifications on the output of the large language model, the accuracy of the content written to the object storage is ensured, and the system rejects results that appear to be structurally correct but whose source is untraceable, whose fields are generalized, or inconsistent with the source object. Simultaneously, the verification middleware loads differentiated JSON Schema templates, so that different types of objects are subject to independent structural constraints, preventing different project objects from being written into the same loose structure. Furthermore, since the verification middleware is independent of the large language model, the judgment result of the output content of the large language model is directly given by the verification middleware based on the current state of the object storage, without relying on the large language model's self-verification, self-scoring, or any auxiliary signals generated by the large language model, thus achieving accurate verification of the output content of the large language model. Attached Figure Description
[0019] Figure 1 This is a flowchart of the steps of a method for verifying and storing the output content of an intelligent agent according to an embodiment of this application; Figure 2 This is a schematic diagram of the execution flow of a method for verifying and storing the output content of an intelligent agent, provided in another embodiment of this application; Figure 3 This is a flowchart of the steps following a JSON parsing failure, provided in another embodiment of this application; Figure 4 This is a flowchart of the steps following a verification failure provided in another embodiment of this application; Figure 5 This is a flowchart illustrating the specific steps of rule extraction rollback provided in another embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0021] It is understandable that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, or the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to be construed as... This application is restricted.
[0023] This application provides a method, device, and medium for verifying and storing the output content of an intelligent agent. The method is applied to a verification middleware, which is a component independent of the large language model process. The method includes: receiving first output content from the large language model; determining a target JSON Schema template based on the first output content and loading the target JSON Schema template; sending the target JSON Schema template or equivalent constraint description information to the large language model to constrain the output of the large language model; generating the equivalent constraint description information based on the target JSON Schema template; obtaining second output content from the constrained large language model; performing JSON parsing on the second output content; if the parsing is successful, performing JSON Schema verification, foreign key constraint verification, and field comparison verification on the parsed second output content; and writing the parsed second output content into object storage if all JSON Schema verification, foreign key constraint verification, and field comparison verification pass. According to the solution provided in this application, by setting up a validation middleware independent of the large language model and using the validation middleware to perform multiple validations on the output of the large language model, the accuracy of the content written to the object storage is ensured, and the system rejects results that appear to be correct in structure but have an untraceable source, whose fields are generalized, or inconsistent with the source object. Simultaneously, the validation middleware loads differentiated JSON Schema templates, ensuring that different types of objects are subject to independent structural constraints, preventing different project objects from being written into the same loose structure. Furthermore, since the validation middleware is independent of the large language model, the judgment result of the large language model's output content is directly given by the validation middleware based on the current state of the object storage, without relying on the large language model's self-validation, self-scoring, or any auxiliary signals generated by the large language model, thus achieving accurate validation of the large language model's output content.
[0024] The embodiments of this application will be further described below with reference to the accompanying drawings.
[0025] Firstly, this application proposes a method for verifying and storing the output content of an intelligent agent. This method is applied to a verification middleware. It should be noted that, in this application, the verification middleware is an independent component running outside the large language model process. Its loading, execution, and judgment do not go through the large language model's own generation path, nor can they be rewritten by the large language model's own output. The verification middleware can be independently replaced, upgraded, or formally verified, and is decoupled from the large language model through an interface. For example... Figure 1 As shown, the method for verifying and storing the output content of an intelligent agent according to the embodiments of this application includes, but is not limited to, steps S100 to S600: Step S100: Receive the first output content of the large language model; Step S200: Determine the target JSON Schema template based on the first output content, and load the target JSON Schema template; Step S300: Send the target JSON Schema template or equivalent constraint description information to the large language model to perform constraint processing on the output of the large language model; the equivalent constraint description information is generated based on the target JSON Schema template; Step S400: Obtain the second output content of the large language model after constraint processing; Step S500: Perform JSON parsing on the second output content. If the parsing is successful, perform JSON Schema validation, foreign key constraint validation, and field comparison validation on the parsed second output content. Step S600: If the JSON Schema validation, the foreign key constraint validation, and the field comparison validation all pass, write the parsed second output content into object storage.
[0026] It's important to note that JSON (JavaScript Object Notation) is a lightweight, text-based data exchange format, independent of programming languages, and widely used for API transmission, configuration files, and front-end / back-end data interaction. JSON Schema refers to a declarative specification for describing JSON data structures, allowing the definition of field names, field types, required fields, value ranges, and schema constraints, used for formal validation of JSON objects. After receiving the first output of the large language model, the validation middleware determines the relevant information that the large language model needs to output, determines the corresponding target JSON Schema template based on this information, and loads the target JSON Schema template. It's worth noting that there are various JSON Schemas suitable for different application scenarios; the target JSON Schema template is determined based on the first output of the large language model. After loading the target JSON Schema template, the template or its equivalent constraint description is sent to the large language model as part of the large language model's output constraints to improve the hit rate of the large language model's initial generation. The large language model, based on the constraints of the target JSON Schema template, returns a second output content to the validation middleware. The validation middleware first performs JSON parsing on this second output content. If parsing is successful, the second output content is interpreted as valid JSON. After successful JSON parsing, the parsed second output content undergoes three checks: JSON Schema validation, foreign key constraint validation (verifying that the primary key of the source object declared by `source_oids` actually exists in object storage), and field comparison validation (comparing the key fields of the second output content with the corresponding fields of the source object using equality or normalized equality comparison). These three checks are performed sequentially; if any check fails, the subsequent checks are not performed. Only when all three checks pass does the writing phase begin, where the parsed second output content is written to object storage.
[0027] According to the method for verifying and storing the output content of an intelligent agent according to embodiments of this application, a verification middleware independent of the large language model is set up. This middleware performs multiple verifications on the output of the large language model, thereby ensuring the accuracy of the content written to the object storage. This prevents the system from rejecting results that appear structurally correct but whose source is untraceable, whose fields are generalized, or inconsistent with the source object. Simultaneously, the verification middleware loads differentiated JSON Schema templates, subjecting different types of objects to independent structural constraints, preventing different project objects from being written into the same loose structure. Since the verification middleware is independent of the large language model, the judgment result of the large language model's output content is directly given by the verification middleware based on the current state of the object storage, without relying on the large language model's self-verification, self-scoring, or any auxiliary signals generated by the large language model, thus achieving accurate verification of the large language model's output content.
[0028] In some embodiments of this application, such as Figure 2 As shown, determining the target JSON Schema template based on the first output includes the following three steps: Based on the first output, obtain the task intent, input source type, or upstream process identifier; Identify the target object type based on task intent, input source type, or upstream process identifier; Determine the target JSON Schema template based on the target object type.
[0029] It should be noted that the task intent, input source type, and upstream process identifier are determined by the upstream task and do not require static pre-setting. The upstream task simply needs to record the definition and identification rules. After receiving the output of the large language model, the validation middleware obtains information such as the task intent, input source type, and upstream process identifier. The task intent refers to the core business action and processing goal that the current system / large language model needs to execute; the input source type refers to the data carrier / data form of the source object that the large language model needs to process; and the upstream process identifier is the identifier set by the upstream task for the source object. After receiving the output of the large language model and obtaining the task intent, input source type, and upstream process identifier, the validation middleware can determine the target object type based on this information. In this application, the target object type includes, but is not limited to, document summary objects, table field objects, code module objects, and log event objects. After determining the target object type, the validation middleware loads the corresponding target JSON Schema template and sends this template or its equivalent constraint description as part of the large language model's output constraints to the large language model, thereby improving the hit rate of the large language model's first generation. This application loads differentiated JSON Schema templates according to the target object type, ensuring that document summaries, table fields, code modules, and log events are subject to independent structural constraints, thus preventing different project objects from being written into the same loose structure. In some embodiments of this application, such as Figure 2 and Figure 3 As shown, in step S500 above, after the validation middleware parses the second output content into JSON, the following steps are also included: Step S510: If parsing fails and the second output content cannot be parsed into valid JSON, it is recorded as a failed attempt; Step S520: Trigger the large language model to regenerate new second output content until the new second output content is parsed into valid JSON or the number of failed attempts reaches a preset value.
[0030] It's important to note that if the second output of the large language model cannot be parsed into valid JSON, it indicates a problem with the model's output. In this case, the large language model needs to regenerate new second output content until it can be parsed into valid JSON. When the validation middleware fails to parse the second output into valid JSON, it records a failed attempt. Each failed attempt increments the count. When the count reaches a preset value, the large language model stops regenerating the second output. Instead of relying on the large language model, the middleware directly extracts summary fields, keywords, location identifiers, and items to be processed from the source object according to preset rules, generating a structured object. This structured object undergoes JSON Schema validation and foreign key constraint validation before being written to object storage.
[0031] After successfully parsing the second output of the large language model into valid JSON, the parsed second output is then subjected to JSON Schema validation, foreign key constraint validation, and field comparison validation in sequence. If a validation fails, the next validation is not performed. Figure 4 As shown, when performing the three verifications, the verification and storage method for the intelligent agent's output content in this embodiment of the application further includes the following two steps: Step S530: If any of the checks fails, record the reason for the failure and increment the number of failed attempts by one; Step S540: Generate new constraints based on the reason for failure and send them to the large language model. Re-trigger the large language model to generate new second output content based on the new constraints until the number of failed attempts reaches a preset value or the new second output content passes JSON Schema validation, foreign key constraint validation, and field comparison validation.
[0032] Specifically, when any validation fails, the reason for failure is recorded. If the number of failed attempts does not reach the preset value, the large language model is re-triggered to generate the second output content, with the reason for failure (such as missing field name, foreign key that does not hit the primary key, or inconsistent field items) as a new constraint condition, thereby constraining the second output content of the large language model so that the new second output content of the large language model passes all validations as much as possible.
[0033] During verification, the count of failed attempts is incremented by one for each failed verification. Once the count of failed attempts reaches a preset value, such as... Figure 5 As shown, the method for verifying and storing the output content of an intelligent agent in this embodiment of the application further includes the following three steps: Step S710: Obtain the source object that has not been processed by the large language model; Step S720: Extract the corresponding fields from the source object according to the preset rules to generate a structured object; Step S730: Perform JSON Schema validation and foreign key constraint validation on the structured object. If the validation passes, write the structured object to object storage.
[0034] It should be noted that when the second output of the large language model fails the three validation checks, the large language model is triggered to regenerate a new second output. Each failure increments the count of failed attempts. When the count reaches a preset value, the large language model stops regenerating the second output. At this point, instead of relying on the large language model for generation, the system directly extracts the summary field, keywords, location identifiers, and items to be processed from the source object according to preset rules, generating a structured object. This structured object is then validated using JSON Schema and foreign key constraints before being written to object storage.
[0035] To avoid sharing a loose structure among different target object types, this application loads different JSON Schema template constraints according to the target object type. The target object types include, but are not limited to, document digest objects, table field objects, code module objects, and log event objects. In some embodiments, referring to Table 1 below, the corresponding JSON Schema template is loaded into the validation middleware according to the target object type, including: When the target object type is a document summary object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a first required field (such as summary, key_terms, pending_items, source_oids, etc.), non-empty summary content, arrays of keywords and items to be processed, and the existence of a source object identifier; When the target object type is a table field object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a second required field (such as field_name, value_range, source_oids, etc.), complete field names and value ranges, and field names that are consistent with the source table rows; When the target object type is a code module object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a third required field (such as module_name, file_path, changed_points, source_oids, etc.), the correct module name, file path, and change point type, and consistency with the source source code object; When the target object type is a log event object, the constraints of the target JSON Schema template include, but are not limited to, the existence of a fourth required field (such as event_level, event_time_window, event_summary, source_oids, etc.), the event level, time window and summary meeting the preset range, and the existence of the source object identifier.
[0036] Specifically, the core configurations of JSON Schema templates for various target object types can be summarized in Table 1 below:
[0037] The above set of fields is used to limit the minimum writable structure. In practice, the JSON Schema template can further constrain field types, minimum number of array elements, string length, path patterns, and enumeration values, but it does not replace the differentiated constraints for different target object types with a uniform large template.
[0038] In this application, before writing the second output of the large language model to object storage, three checks are performed by a validation middleware to implement pre-write constraints. The validation middleware is an independent component running outside the large language model process. Its loading, execution, and judgment do not go through the large language model's own generation path, nor can they be rewritten by the large language model's own output. The validation middleware can be independently replaced, upgraded, or formally verified, and is decoupled from the large language model through an interface. Therefore, the judgment result of the pre-write constraints is directly given by the validation middleware based on the current state of the object storage, without relying on the large language model's self-validation, self-scoring, or any auxiliary signals generated by the large language model. The pre-write constraints are executed in a fixed order: JSON Schema validation, foreign key constraint validation, and field comparison validation. If the previous item fails, the next item is not performed. Each item is implemented by the validation middleware.
[0039] The first validation step is JSON Schema validation, which involves the following three steps: Load the target JSON Schema template and perform structural validation on the parsed second output content using the canonical keywords of the target JSON Schema template; If all specification keywords pass the validation, the parsed second output content will pass the JSON Schema validation. If any specification keyword fails validation, the parsed second output content will not pass JSON Schema validation.
[0040] Specifically, after the validation middleware loads the JSON Schema template corresponding to the target object type, it uses the JSON Schema template's specification keywords, such as required, type, pattern, enum, items, minItems, and additionalProperties, to perform structural validation on the parsed object. If any specification keyword fails validation, the item is considered to have failed. Only when all specification keywords pass validation is the JSON Schema validation considered successful.
[0041] The second check is the foreign key constraint check (verifying referential integrity), which includes the following four steps: Retrieve the source_oids field from the parsed second output; the source_oids field is a collection of foreign keys pointing to the source object. For each foreign key in the source_oids field, query the object storage to see if a corresponding object primary key exists; If each foreign key has a corresponding primary key in the object storage, the parsed second output content passes the foreign key constraint validation. If any foreign key cannot be found in the object storage as a corresponding primary key, the parsed second output content fails the foreign key constraint validation.
[0042] It should be noted that the `source_oids` field in the second output of the large language model is a set of foreign keys pointing to the source objects; each object in the object storage has a unique object primary key `oid`. The validation middleware performs an exact lookup of the object primary key in the object storage for each foreign key in `source_oids`. If no corresponding object primary key is found in the object storage for any foreign key, the referential integrity is considered violated, and this step fails. This validation is a deterministic determination based on the object primary key index and does not involve any similarity calculations or threshold determinations. Only when all foreign keys have corresponding object primary keys in the object storage will the foreign key constraint validation pass.
[0043] The third verification is a field comparison verification, which includes the following three steps: Based on the pre-declared key field set corresponding to the target object type, the key field values in the parsed second output content are compared with the corresponding field values in the source object pointed to by each foreign key, or a normalized comparison is performed. If all key field values match successfully, the parsed second output content will pass the field comparison verification. If any key field value fails to match, the parsed second output content will fail the field matching check.
[0044] Specifically, the verification middleware, based on the pre-declared set of key fields for the target object type, performs an equality or normalized equality comparison between the key field values in the second output and the corresponding field values in the source object pointed to by the foreign key (normalization includes reversible transformations such as case normalization, whitespace normalization, and hexadecimal bit width normalization). If any key field value does not match, the comparison fails. For example, if the field_name output of the table field object is VID, while the corresponding field in the source table row object is VLAN_ID, the comparison will fail.
[0045] It should be noted that, due to the randomness of the output of the large language model, this application includes a controlled retry mechanism. Preferably, the system sets a limited total number of attempts for different target object types. After each failure, only the failure category and necessary violation information are fed back to the large language model, such as missing fields, foreign keys that do not hit the primary key, or inconsistent key fields, thus subjecting subsequent generation to clearer constraints. If the large language model generates a result that satisfies all validations within the maximum number of attempts, the system records the successful rounds and writes them to object storage; if failures continue, the model retry is stopped and the rule extraction rollback process begins.
[0046] To ensure the system can continue to run even when the model output is unstable, this application introduces a rule extraction rollback mechanism. The rollback result must undergo the same JSON Schema validation and foreign key constraint validation as the normal path before it can be written to object storage, thereby ensuring that the rollback path does not violate the referential integrity of object storage.
[0047] The rule extraction fallback module no longer requires the large language model to be regenerated, but instead deterministically extracts the following fields from the source object:
[0048] The rollback result uses the same JSON Schema validation and foreign key constraint validation as the normal path as the write threshold, thereby avoiding interruption of the engineering process due to unstable model format.
[0049] In some embodiments of this application, the method for verifying and storing the output content of the intelligent agent further includes the following steps: When writing the parsed second output content to object storage, a reverse index is established between the source object identifier and the object primary key corresponding to the second output content. The reverse index is used to locate and trace the source object based on the object primary key.
[0050] Specifically, to ensure the connection between the structured results and the original evidence, this application requires that different target object types retain a set of source object identifiers, which is named `source_oids` in one implementation. This field is a necessary traceability field in the structured object. Preferably, the source object identifier can directly inherit the identifiers from the input object set, or it can come from chapter objects, table rows, code nodes, or log event identifiers generated by the slicing tool, or it can be assigned by the object storage system during the access phase. When writing the structured results, the system establishes a reverse index between the source object identifiers and the `oid` of the current result object to support result source location, change impact retrieval, and traceable citation. `source_oids` is the foreign key field constrained by the foreign key constraint validation; the reverse index established when writing to the object storage is the reverse index of the foreign key, used to retrieve all output objects that reference the primary key of the source object.
[0051] The following examples illustrate the specific application of the method for verifying and storing the output content of an intelligent agent. It should be noted that the following examples are merely illustrative and not intended to limit the scope of this application.
[0052] Example 1: Protocol Document Field Compression Scenario The system receives a table from a protocol document, with the goal of generating table field objects. After the validation middleware identifies that the task belongs to the `table_field` type (table field object), it loads the corresponding JSON Schema template, requiring the large language model output to contain at least `field_name`, `value_range`, and `source_oids`.
[0053] For example, the identifier of the source table row object is doc_eth_spec_v2_3_table5_row12, and its field name is VLAN_ID with a value range of 0x000-0xFFF. If the source_oids output by the large language model contains an identifier that does not exist in the object storage, the system rejects the result through foreign key constraint validation; if the large language model summarizes the field name as VID, the system rejects the result through field comparison validation. All of the above failure reasons are fed back to the large language model in controlled retries. The final result that can be written to the object storage should retain the field name VLAN_ID, the value range of 0x000-0xFFF, and the source object identifier ["doc_eth_spec_v2_3_table5_row12"].
[0054] Example 2: Code Change Summary Scenario The system performs a structured summary of a code modification result, with the target object type being `code_module` (code module object). The JSON Schema template requires the output of `module_name`, `file_path`, `changed_points`, and `source_oids`, and requires that the module name and file path be consistent with the source source code object. In a representative case, the source source code object displays the module name as `eth_mac_tx`, the file path as `src / mac / eth_mac_tx.v`, and the change locations corresponding to `code_node_eth_mac_tx_47` and `code_node_eth_mac_tx_103`. If the large language model summarizes the file path as `src / mac / tx.v`, the system determines that the field comparison and validation has failed and triggers a limited retries with a reason. If the required output is still not obtained, the system extracts the module name, file path, change points, and source object identifier from the original code object to form a structured code module object, and writes it to object storage after passing the same JSON Schema validation and foreign key constraint validation as the normal path.
[0055] According to the verification and storage method for intelligent agent output content according to the embodiments of this application, by loading differentiated JSON Schema templates according to the target object type, document summaries, table fields, code modules, and log events are subject to independent structural constraints, avoiding different engineering objects from being written into the same loose structure. After parsing the second output content of the large language model into JSON, JSON Schema verification, foreign key constraint verification, and field comparison verification are performed. This process enables the system to reject results that appear to be structurally correct but whose source is untraceable, whose fields are generalized, or inconsistent with the original object. When verification fails, this application adopts controlled retries instead of infinite retries and transforms the failure reason into constraint feedback; when retries still fail, the rule extraction rollback process produces writable structured results under the same JSON Schema verification and foreign key constraint verification, thereby maintaining the continuity of the engineering process. At the same time, the foreign key reverse index established when writing to object storage enables the structured results to locate the original evidence and supports subsequent review, citation, and change impact analysis.
[0056] This application models the output of a large language model as an object to be written in object storage, and formalizes `source_oids` as a foreign key field in the second output content pointing to the source object, thus transforming the verification task from a text consistency problem to a referential integrity verification problem in object storage. It should be noted that text consistency solutions address the text consistency problem: the objects being verified are whether the field text appears in the original text or a specialized thesaurus, and whether the sentence semantics match the original text. Its decision space is text-to-text similarity comparison, and its toolset includes thesaurus lookup, TF-IDF, and embedding vector similarity. This application, however, addresses the database referential integrity problem: the objects being verified are whether the foreign key of the output object exists in the primary key set of the object storage, and whether the key field is equal to the corresponding field of the source object pointed to by the foreign key. Its decision space is the reference relationship between the second output content and existing objects in the object storage, and its toolset includes primary key index exact lookup and field equality comparison. The two solutions belong to different problem domains, rely on different decision theories, and require different runtime conditions. Of the three checks, the foreign key constraint check is a deterministic determination based on the primary key index, and the field comparison check is an equality comparison based on direct location of the source object's primary key. Neither of these relies on any similarity or threshold. Rule extraction rollback can still produce writable objects under referential integrity constraints, ensuring that referential integrity remains intact even when the model is unstable. Therefore, the pass / fail of this application's checks depends solely on the current state of the object storage, and is independent of model training data, external vocabularies, accessibility of the original text, and any similarity threshold. Stable and replayable check conclusions can be obtained when the model and object storage remain unchanged. Furthermore, the foreign key reverse index allows for immediate lookup of all references to any source object's primary key, providing referential integrity backtracking capabilities beyond the write-before constraints. Together with the write-before constraints, these constitute the two pillars of this invention on the object storage side.
[0057] like Figure 6 As shown, Figure 6 This is a structural diagram of an electronic device provided in one embodiment of this application. The present invention also provides an electronic device 800, comprising: The processor 810 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 820 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 820 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 820, and the processor 810 calls and executes the verification and storage method for the intelligent agent's output content according to the embodiments of this application. The input / output interface 830 is used to implement information input and output; The communication interface 840 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 850 transmits information between various components of the device (e.g., processor 810, memory 820, input / output interface 830, and communication interface 840); The processor 810, memory 820, input / output interface 830 and communication interface 840 are connected to each other within the device via bus 850.
[0058] In addition, this application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program. When the computer program is executed by a processor, it implements the above-described method for verifying and storing the output content of the intelligent agent.
[0059] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate, and may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0060] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0061] The above provides a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.
Claims
1. A method for verifying and storing the output content of an intelligent agent, characterized in that, The method, applied to a validation middleware, which is a component independent of the large language model process, includes: Receive the first output content of the large language model; The target JSON Schema template is determined based on the first output content, and the target JSON Schema template is loaded. The target JSON Schema template or equivalent constraint description information is sent to the large language model to constrain the output of the large language model; the equivalent constraint description information is generated based on the target JSON Schema template. Obtain the second output content of the large language model after constraint processing; The second output content is parsed using JSON. If the parsing is successful, the parsed second output content is then validated using JSON Schema, foreign key constraints, and field comparison. If the JSON Schema validation, the foreign key constraint validation, and the field comparison validation all pass, the parsed second output content is written to object storage. The step of determining the target JSON Schema template based on the first output content includes: Based on the first output content, obtain the task intent, input source type, or upstream process identifier; Based on the task intent, the input source type, or the upstream process identifier, identify the target object type; the target object type includes at least one of document summary object, table field object, code module object, and log event object. The target JSON Schema template is determined based on the target object type.
2. The method for verifying and storing the output content of an intelligent agent according to claim 1, characterized in that, Determining the target JSON Schema template based on the target object type includes: When the target object type is the document summary object, the constraints of the target JSON Schema template include: the existence of a first required field, the summary content is not empty, the keywords and items to be processed are arrays, and the source object identifier exists; When the target object type is the table field object, the constraints of the target JSON Schema template include: the existence of a second required field, the complete field name and value range, and the field name being consistent with the source table row; When the target object type is the code module object, the constraints of the target JSON Schema template include: the third required field, the module name, file path and change point type are correct and consistent with the source source code object; When the target object type is the log event object, the constraints of the target JSON Schema template include: the existence of a fourth required field, the event level, time window, and summary meeting the preset range, and the existence of the source object identifier.
3. The method for verifying and storing the output content of an intelligent agent according to claim 1, characterized in that, After the step of performing JSON parsing on the second output content, the method further includes: If parsing fails and the second output cannot be parsed into valid JSON, it is recorded as a failed attempt. The large language model is triggered to regenerate new second output content until the new second output content is parsed as valid JSON or the number of failed attempts reaches a preset value.
4. The method for verifying and storing the output content of an intelligent agent according to claim 3, characterized in that, The step of performing JSON Schema validation, foreign key constraint validation, and field comparison validation on the parsed second output content includes: If any of the checks fails, record the reason for the failure and increment the number of failed attempts by one. New constraints are generated based on the reasons for failure and sent to the large language model. The large language model is then triggered to generate new second output content based on the new constraints until the number of failed attempts reaches the preset value or the new second output content passes the JSON Schema validation, foreign key constraint validation, and field comparison validation.
5. The method for verifying and storing the output content of an intelligent agent according to claim 3 or 4, characterized in that, When the number of failed attempts reaches the preset value, the method further includes: Obtain the source object that has not been processed by the large language model; Based on preset rules, the corresponding fields are extracted from the source object to generate a structured object; Perform JSON Schema validation and foreign key constraint validation on the structured object. If the validation passes, write the structured object to the object storage.
6. The method for verifying and storing the output content of an intelligent agent according to claim 1, characterized in that, The JSONSchema validation includes: Load the target JSON Schema template, and use the canonical keywords of the target JSON Schema template to perform structural validation on the parsed second output content; If all the specified keywords pass the validation, the parsed second output content passes the JSONSchema validation. If any of the specified key validations fails, the parsed second output content will not pass the JSONSchema validation.
7. The method for verifying and storing the output content of an intelligent agent according to claim 1, characterized in that, The foreign key constraint verification includes: Obtain the source_oids field from the parsed second output content; the source_oids field is a set of foreign keys pointing to the source object. For each foreign key in the source_oids field, query the object storage to see if there is a corresponding object primary key; If each of the foreign keys has a corresponding object primary key in the object storage, then the parsed second output content passes the foreign key constraint verification. If any of the foreign keys cannot be found in the object storage as a corresponding primary key, then the parsed second output content fails the foreign key constraint validation.
8. The method for verifying and storing the output content of an intelligent agent according to claim 7, characterized in that, The field comparison and verification includes: Based on the pre-declared key field set corresponding to the target object type, the key field values in the parsed second output content are compared with the corresponding field values in the source object pointed to by each foreign key, or a normalized comparison is performed. If all key field values are successfully matched, the parsed second output content will pass the field comparison verification. If any of the key field values fails to match, the parsed second output content will not pass the field matching verification.
9. The method for verifying and storing the output content of an intelligent agent according to claim 1, characterized in that, The method further includes: When writing the parsed second output content into the object storage, a reverse index is established between the source object identifier and the object primary key corresponding to the second output content. The reverse index is used to locate and trace the source object based on the object primary key.
10. A verification and storage device for the output content of an intelligent agent, characterized in that, include: The receiving module is used to receive the first output of the large language model; A loading module is used to determine a target JSON Schema template based on the first output content and to load the target JSON Schema template. The constraint module is used to send the target JSON Schema template or equivalent constraint description information to the large language model to perform constraint processing on the output of the large language model; the equivalent constraint description information is generated based on the target JSON Schema template. The acquisition module is used to acquire the second output content of the large language model after constraint processing. The validation module is used to perform JSON parsing on the second output content. If the parsing is successful, the parsed second output content is validated for JSON Schema, foreign key constraints, and field comparison. The writing module is used to write the parsed second output content into the object storage if the JSON Schema validation, the foreign key constraint validation, and the field comparison validation all pass. The step of determining the target JSON Schema template based on the first output content includes: Based on the first output content, obtain the task intent, input source type, or upstream process identifier; Based on the task intent, the input source type, or the upstream process identifier, identify the target object type; the target object type includes at least one of document summary object, table field object, code module object, and log event object. The target JSON Schema template is determined based on the target object type.
11. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the verification and storage method for the output content of an agent as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the verification and storage method for the output content of an intelligent agent as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Content generation method and device, agent, equipment, medium and product
CN122334498A
Multi-agent-based result verification method and device, storage medium and equipment
CN122346663A