Man-machine interaction language model output post-processing system supporting field-level score comparison and rule engine and optimization method

By supporting a human-computer interaction language model output post-processing system that supports field-level scoring comparison and rule engine, the problems of single function and lack of systematization in large model post-processing modules are solved. The system achieves standardization, format consistency and traceability of large model output content, improves the quantitative evaluation and automatic optimization capabilities of generation quality, and ensures that the content meets the seamless integration of industrial systems.

CN121598935APending Publication Date: 2026-03-03YUNNAN KUNMING SHIPBUILDING DESIGN & RESEARCH INSTITUTE

Patent Information

Application Number
CN202510947534.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing large model post-processing modules are limited in function, lack versatility and modularity, cannot adapt to the diverse needs of different tasks and fields, cannot meet the requirements of accurate evaluation and format consistency of structured data in industrial scenarios, lack a collaborative mechanism for manual verification, and lack a systematic post-processing solution.

Method used

This paper presents a human-computer interaction language model output post-processing system that supports field-level scoring comparison and rule engine. Through information recording, scoring comparison, modification, manual verification and confirmation, and data output modules, it achieves precise management of structured generated content, quantitative evaluation and automatic optimization of generated quality, provides auditability and reliability of generated content, supports structured template and task type recognition, introduces BLEU and ROUGE scoring mechanisms, and performs multi-stage text correction and confirmation to ensure that the content meets the requirements of industrial information systems.

Benefits of technology

It achieves standardization, format consistency, and traceability of large model output content, improves the quantitative evaluation and automatic optimization capabilities of generated quality, ensures seamless integration of content with industrial systems, solves the output quality assurance problems when generated content is uncontrollable, quality is unquantifiable, and reference samples are lacking, and provides a full-process recording and verification mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598935A_ABST
    Figure CN121598935A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction language model output post-processing system supporting field-level score comparison and a rule engine and an optimization method. And the post-processing system cooperatively works through an information recording module, a score comparison module, a modification module, a manual check confirmation module and a data output module. According to the post-processing optimization method, a framework and a prototype software product are built according to the post-processing logic of the test stage and the post-processing logic of the working stage. According to the post-processing optimization method, firstly, a generation result and a structured demand are collected and marked, then the text quality is evaluated through BLEU and ROUGE, optimization and modification are carried out according to a scoring result, finally, the text is output according to a standard format and is in butt joint with an industrial system, and it is ensured that the text is accurate and available. According to the method, traceable, intervening and integratable high-quality post-processing of the output result of the language model is realized, and the requirements of complex task scenes in industrial application on the content, structure and quality control of the intelligently generated text are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to large model post-processing technology, belonging to the field of artificial intelligence and natural language processing technology. Specifically, it is a human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine. Background Technology

[0002] Existing large-scale model post-processing modules are typically composed of one or a few specific modules, offering limited functionality, high project adaptability, and a lack of a general workflow architecture. Currently, large-scale model post-processing solutions on the market are often tailored to a specific project or scenario, using one or two techniques (such as duplication penalties, temperature adjustment, sampling strategies, and text cleaning). Their evaluation methods often directly utilize flexible evaluation metrics originally used to determine the grammar of model text, failing to meet the need for complete consistency in evaluating structured data in industrial scenarios.

[0003] (1) Shortcomings and defects of existing large models

[0004] Existing large-scale model post-processing modules are often designed for specific tasks or application scenarios, with relatively singular functionality and high adaptability to specific projects. This means that many modules can only solve specific types of text generation problems, lacking flexibility and unable to cope with the diverse needs of different tasks or domains. Specifically, the lack of a general workflow architecture means that current post-processing methods have not formed a universal and reusable workflow architecture, often performing targeted processing for a specific task, lacking good modularity and versatility. As the application of large-scale models expands, the lack of a unified post-processing framework makes the migration and adaptation costs between different tasks high.

[0005] Most existing post-processing methods lack a systematic approach, relying solely on scoring functions to evaluate text's grammatical and fluency characteristics. While these scoring functions are primarily used for assessing grammatical and fluency features and are suitable for Natural Language Generation (NLG) tasks, they are not true post-processing methods. In many cases, post-processing either depends solely on scoring or is entirely performed by another large model. However, these metrics cannot fully meet the needs of evaluating structured data in industrial scenarios. The industrial sector often requires more precise and standardized evaluation methods, especially for structured text and data. Existing evaluation methods cannot assess the accuracy, data integrity, and format consistency of the output text. Specifically, in industrial scenarios, generated content often needs to precisely conform to specific formats and requirements, such as product lists, reports, and technical documents. Existing evaluation methods fail to adequately consider how to assess whether the generated content conforms to preset structured template requirements, and they also lack effective personalized scoring for different task types.

[0006] Existing post-processing methods for large models, such as repetition penalties, while effective in natural language generation, primarily optimize for fluency and diversity in text generation, lacking adaptability to complex industrial tasks. Specifically, in industrial scenarios, generated content not only requires grammatical correctness and fluency but also ensures data format consistency, content accuracy, and compliance with industry standards. However, existing post-processing methods are mostly limited to text-level processing, failing to meet the demands of structured data, precise formatting requirements, and contextual consistency. Content filtering, concatenation, and formatting adjustments are mainly used to remove irrelevant content and adjust text formatting, but in complex industrial applications, text generation often involves the integration of multiple data sources, and single text adjustment and filtering methods cannot effectively handle the diversity and complexity of content.

[0007] Existing post-processing modules rely heavily on automated correction and adjustment, such as duplicate penalties, lacking an effective collaborative mechanism with human review. Especially in industrial applications, text generation requires not only grammatical and format accuracy but also compliance with business logic, industry standards, and database requirements. In complex tasks and demanding scenarios, complete reliance on automation often fails to ensure content completeness and data unbiasedness; the collaborative role of human review becomes paramount. The absence of a combined human and automated review mechanism poses certain risks to existing technologies in mission-critical scenarios, particularly those with extremely high accuracy requirements.

[0008] The limitation of existing technologies lies in the fact that most existing technologies fail to consider the diversity and complexity of the output results of large models. Different tasks and different fields (such as the coordination and cooperation of different workshops and factories in an enterprise) have their own business requirements, data specifications and content formats, while existing post-processing technologies mostly process text in a "general" way, lacking targeted business customization support. For example, in the manufacturing scenario, text generation not only requires generating grammatically correct reports, but also requires embedding specific product data, production plans, equipment information, etc. Existing post-processing technologies have failed to provide flexible solutions that meet industry needs. (2) The current status and shortcomings of large model post-processing technologies

[0009] Currently, most optimization methods for large language models (such as model fine-tuning, architecture modification, pruning, and quantization) focus primarily on the preprocessing stage (such as data cleaning, knowledge base construction, and prompt word optimization) and the model adaptation stage (such as improving model performance through fine-tuning and distillation), while neglecting the role of post-processing techniques in the model's output. Specifically, preprocessing and model optimization typically target the model's performance, failing to adequately address issues such as standardization, format consistency, and accuracy of the output content. This is why post-processing techniques are crucial.

[0010]

[0011] The post-processing stage described in Table 1 is a crucial step implemented after the large model output content is generated. Its purpose is to further adjust the generated content to meet application requirements, standard formats, or structured requirements. In many industrial applications, post-processing is indispensable, especially when the model's generated content directly impacts business decisions; post-processing technology ensures the standardization and consistency of the content.

[0012] Existing post-processing techniques still have many limitations, especially when processing large model outputs in the absence of reference samples.

[0013] Regarding scoring, most current post-processing scoring methods rely on flexible evaluation metrics originally used for semantic detection, such as BLEU or ROUGE. These metrics typically depend on reference templates, requiring comparison of the model output with existing standard answers. One dilemma arises from this inability to adapt to the demands of structured output. In industrial scenarios, content conforming to specific formats or structures needs to be generated. Existing scoring metrics focus more on the fluency and grammatical accuracy of natural language generation, failing to measure the completeness and accuracy of structured data. A second dilemma is the lack of support for unsampled conditions. In the absence of reference samples, existing scoring methods are ineffective, resulting in a failure to effectively evaluate the quality of the model's output.

[0014] From the current state of research on post-processing technology, most existing post-processing logics are limited to single functions, such as deletion and filtering, and are only used to remove redundant or inappropriate content. Research on how to effectively add or modify generated content is neither systematic nor in-depth, resulting in significant functional limitations in existing technologies. Furthermore, many post-processing methods lack integration, failing to fully integrate multiple processing techniques. They are often separate, independent functional modules, unable to provide complete and systematic solutions, especially when facing tasks requiring complex modifications, where the processing results are often unsatisfactory.

[0015] Regarding the standardization of post-processing, existing post-processing technologies have many shortcomings in large-scale model applications, particularly lacking support for handling structured data and ensuring format consistency. Existing technologies tend to focus on optimizing specific aspects, failing to provide a systematic and standardized post-processing solution. They often simply rely on large models and post-process their output, which is essentially equivalent to multi-agent collaborative operations. However, due to the lack of systematic and standardized design in existing post-processing modules, collaboration between modules is inefficient and difficult to achieve effective linkage. This results in the final output failing to meet the high standards required for industrial scenarios and hindering the efficient and accurate completion of complex tasks. Existing technologies rely on multiple independent modules to process model output (such as scoring, modification, and cleaning). These modules lack unified standards in processing standards, methods, and data formats, leading to inconsistent overall post-processing results that fail to meet industry standards or specific task requirements. Existing post-processing methods often optimize specific aspects, such as text fluency and grammatical accuracy, but rarely consider how to ensure consistency and format standardization of output content across diverse tasks, meeting the requirements of industrial applications. This lack of a unified framework makes it difficult for the synergy between various processing methods to improve the standardization of the entire process. Especially in practical applications, the model output often contains various types of structured information and text content, and existing technologies fail to provide a comprehensive correction mechanism that can bridge text generation, structured data, and formatting specifications. This lack of global evaluation often results in processing outcomes that fail to meet the expected high standards, particularly in industrial applications where precise formatting and content consistency are required.

[0016] (3) Current Status of Typical Existing Literature

[0017] Regarding the use of large language models as post-processing tools, Chinese patent (CN119003742 B A Digital Intelligent Power Control Question Answering Method and System Based on a Large Language Model) employs a large language model to post-process texts that meet the comprehensive similarity requirements and generate corresponding answers. However, this technical solution has shortcomings: the method mainly relies on a large language model to post-process texts that meet the comprehensive similarity requirements and generate answers. Although effective for power control question answering systems, its versatility in diverse and complex application scenarios is poor. The method relies on similarity scores to filter and generate answers, but this method encounters the problem of generating incorrect answers from highly similar texts. If the model only relies on surface similarity and ignores deeper semantic understanding, the accuracy of the output answers will be insufficient. Furthermore, the method does not address how to handle dynamic changes in the power control environment, such as changes in equipment status and the influence of external environmental factors, and the large model fails to effectively adapt to these changes, resulting in the invalidation of the obtained answers in practical applications.

[0018] Regarding the use of post-processing code to score matching sequences, Chinese patent (CN 119377434 A Method and Apparatus for Fast Image Data Filtering Based on Large Models) calculates matching sequences of image and text data based on visual features and semantic information using image-text matching technology. Post-processing code then processes these sequences, analyzing the matching scores and confidence scores between the image data and prompt words. However, this approach has shortcomings: the method uses a preset dynamic threshold to filter and determine the matching degree of image data. Static thresholds are unsuitable for different datasets or scenarios and cannot be dynamically adjusted, resulting in unsatisfactory filtering performance in complex scenarios. For example, when the content or semantic information of an image changes significantly, the preset threshold cannot accurately determine its matching degree. Secondly, while the method calculates similarity scores and filters by matching images with text, the accuracy of image matching decreases under high-noise data or complex background information. Due to the lack of a more complex adaptation mechanism, the method is insufficient in handling subtle differences or images of different scales. In addition, the method relies on the matching score and confidence score between the image and the prompt word, which is insufficient to fully reflect the true matching relationship between the image and the text in some scenarios, especially when the image features are not completely related to the prompt word, resulting in misjudgment.

[0019] In the field of discriminative model-based post-processing for text translation, the paper "Research on Multi-Task Learning and Post-Processing Methods for Text Grammatical Error Detection and Correction" (Pan Fayu, 2023) uses a newly proposed text correctness discriminator trained on a modification acceptability discrimination task to compare the correctness of the text before and after the error correction model's modification. It retains correct modifications to the original text and filters out erroneous modifications, thereby improving the accuracy and F-square of the error correction model. 0.5 However, the shortcomings of this technical solution are as follows: the method uses a text correctness discriminator to post-process the error correction model, but the effectiveness of error correction heavily depends on the quality and diversity of the training data. Although the method uses a discrimination task to filter error corrections, it does not address how to handle ambiguous or polysemous errors. The method mainly focuses on the detection and correction of grammatical errors in text and cannot effectively handle semantic understanding problems. In natural language processing, in addition to grammatical errors, semantic errors and contextual inconsistencies also cause problems, and the processing of these aspects is not in-depth enough.

[0020] Therefore, the motivation for improving the technical solution described in this invention is to evolve from "fragmented, aesthetically pleasing" post-processing to a "structured, auditable, and fully closed-loop" intelligent post-processing platform, truly bringing large models to rigorous industrial applications. Summary of the Invention

[0021] (1) Purpose of the invention: In view of the problem that the output content of the model is difficult to achieve high standard, automation and structured processing in the current industrial post-processing process, the present invention proposes a human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine, thereby constructing an algorithm system that can adapt to and process the output results of various large models, generate preset structured data / text, improve the standardization of model output, and enhance the practical role of large models in industrial scenarios.

[0022] The core technical problem that this invention aims to solve is the lack of format standardization, semantic consistency, and traceability in content generated by Large Language Models (LLM) in industrial scenarios.

[0023] The core technical problems to be solved by this invention can be subdivided into the following five issues: First, the problem of uncontrollable generated content, difficulty in standardizing output format, and lack of structured control leading to the inability to automatically transfer or execute output results; Second, the problem of unquantifiable generated quality, lack of effective evaluation mechanisms to measure whether the generated content meets business objectives and semantic accuracy, and difficulty in objectively scoring and comparing the quality of output text; Third, the problem of lack of output quality assurance mechanisms in automatic post-processing in scenarios lacking reference samples or with complex tasks; Fourth, the problem of untraceable output content, inability to reconstruct the large model generation and modification process, and lack of a full-process recording and verification mechanism; Fifth, the problem of lack of standard encapsulation of model output and cross-system output interface support, and lack of systematic structural support for format conversion and system integration.

[0024] Therefore, the primary objective of this invention is to provide a post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine. By constructing a complete post-processing system architecture that includes recording, scoring, modification, verification, and output, it achieves precise management of structured generated content, quantitative evaluation and automatic optimization of generated quality, and improved auditability and reliability of output data. Ultimately, it enables seamless integration and practical application with the enterprise's industrial system.

[0025] The second objective of this invention is to provide a post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine. This system can provide a content generation control mechanism based on structured templates and task type recognition. By introducing structured template management, task type classification, and field-level control mechanisms, the system standardizes the output format of large models, making them conform to the data structure requirements of industrial information systems (such as ERP / MES / APS), and achieves structured and standardized output of model-generated content.

[0026] The third objective of this invention is to provide a post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine. This system can provide a quantitative evaluation mechanism for generation quality in structured task scenarios. By introducing automated scoring mechanisms such as BLEU and ROUGE, combined with field-level comparison strategies, a unified quantitative standard for generation quality is established, and a quantitative content quality evaluation system is built.

[0027] The fourth objective of this invention is to provide a post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine. This system can provide a multi-stage text correction and confirmation mechanism for complex task scenarios, which can solve the problem that the current model-generated content has insufficient automatic correction capability and difficulty in ensuring quality when there is a lack of reference samples or complex structure requirements.

[0028] The fifth objective of this invention is to provide a human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine. It can provide a full-process information traceability and control mechanism. Through the information recording module, the entire process of model input, output, intermediate results, modification suggestions and manual verification operations is recorded to ensure that data processing is traceable and verifiable.

[0029] The sixth objective of this invention is to provide a post-processing system and optimization method for the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine. This system can provide a multi-format data conversion and industrial protocol adaptation mechanism, and can solve the problem that the content generated by large models cannot be directly embedded into industrial system processes, thus achieving seamless connection from intelligent generation to industrial execution. (2) Summary of the invention: This invention provides a post-processing system and optimization method for the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine. It mainly includes a post-processing system for the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine, and an optimization method for the post-processing of the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine.

[0030] Firstly, a post-processing system for human-computer interaction language model output that supports field-level scoring comparison and rule engine includes a complete post-processing optimization method, as well as five modules corresponding to sending and receiving methods.

[0031] First, a method for optimizing the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine, comprising the following steps:

[0032] S1. Recording the operation of the information module: By recording key data in the generation process and combining it with structured management, the text generation process can be traced, and the generated text conforms to the preset structure and classification standards;

[0033] The S1.1 Result Recording Unit operates by fully recording the model output, field templates, task identifiers, and intermediate content, enabling full-process traceability of the generated data.

[0034] S1.11 Capture Model Generation Results: Completely record the final text content and related field information generated from the model, ensuring that all output content has a traceable data foundation;

[0035] S1.12 Capture intermediate content generated during post-processing: Record and capture intermediate text and data generated during post-processing;

[0036] S1.2 Operation of the structured template management unit: Classifies and identifies text according to preset structured templates, so that the generated text conforms to specific structural specifications;

[0037] S1.21 Structured Classification: Based on the preset template, the generated text and fields are classified and marked to adapt to the dynamic management of multiple types of tasks, and then recorded and collected according to the structural requirements;

[0038] S1.22 Structured storage: Storing recorded content according to a unified format standard;

[0039] S1.3 Task Type Identifier Classification Unit Operation: Refined classification management of content, marking explicit and implicit information;

[0040] S1.31 Explicit Information Labeling: The basic information and structural hierarchy of user input are labeled in a hierarchical manner, and the explicit fields of the generated content are clearly identified and managed;

[0041] S1.32 Labeling Implicit Information: Extract and label implicit information in the original input to support task understanding and model reasoning;

[0042] S1.33 Annotation level information: Used to annotate the structural hierarchy and attribute information of the generated content, enhancing the system's ability to understand the multidimensional logical structure of the text;

[0043] S2. Operation of the scoring and comparison module: Similarity assessment and field-level analysis are performed by using BLEU and ROUGE indicator units to quantify the matching degree and coverage of the generated text with the reference content, and then the text quality is scored.

[0044] The S2.1BLEU indicator unit operates by calculating similarity through n-gram matching to evaluate the consistency of the generated text in terms of vocabulary and word order.

[0045] S2.11 Input Data Processing

[0046] S2.111 Model Output and Reference Sample Preparation: Collect the model-generated text and corresponding reference samples, and compare and analyze the similarity between the generated content and the standard answer;

[0047] S2.112 Field-level extraction: Extract key field information from the generated text and the reference text to perform refined field-level difference analysis;

[0048] S2.12 Calculation of word overlap rate

[0049] S2.121 n-gram matching: The overlap at the word order level is evaluated by performing n-gram segmentation matching between the model output and the reference text.

[0050] S2.122n-gram accuracy P n Calculation: Calculate the matching ratio of each level of n-gram in the generated text and the reference text to measure its lexical-level consistency.

[0051] The accuracy formula is as follows:

[0052]

[0053] S2.131 Calculate BLEU value: based on n-gram accuracy P n The BLEU score is calculated by combining the lengths of the reference and generated text, and is used to measure overall text similarity.

[0054] The specific formula is as follows:

[0055]

[0056] Where r represents the length of the reference test case; c represents the length of the model output; P n Represents the precision of the n-gram; ω n Weights representing the precision of n-grams;

[0057] S2.132 Calculate overall and field-level BLEU scores: Calculate the overall BLEU value of the text and perform field-level scoring as needed to assess the generation quality of specific structured fields;

[0058] S2.14 Results Output and Visualization

[0059] S2.141 Results Presentation: The overall BLEU score and field-level scores are visualized in the form of charts or tables, showing the matching situation of different n-grams;

[0060] The S2.2ROUGE indicator unit operates by calculating an n-gram to measure the coverage of the reference content and identifying the degree of similarity in the generated text through field-level analysis.

[0061] S2.21 Input Data Processing

[0062] S2.211 Model Output and Reference Sample Preparation: The model-generated text Y and reference text X undergo unified preprocessing.

[0063] The calculation is as follows:

[0064] ROUGE-1: Based on 1-gram

[0065] ROUGE-2: Based on 2-gram

[0066] S2.212 Field-level data extraction: Field-level analysis is performed on some data to prepare for subsequent field-level ROUGE calculations;

[0067] S2.22n-gram extraction and matching

[0068] S2.221 n-gram generation: Extract 1-grams and 2-grams from the reference text and the generated text for ROUGE-1 and ROUGE-2 matching analysis.

[0069] The ROUGE-N formula is as follows:

[0070]

[0071] Recall, or recall, is the percentage of n-grams in the reference text that appear in the model's output.

[0072]

[0073] S2.222 Matching Statistics: Statistically count the total number of n-grams in the reference text and the number of n-grams that match in the generated text, providing basic data for similarity calculation.

[0074] Extraction of the Longest Common Subsequence (LCS) in S2.23

[0075] S2.231 LCS Calculation: Calculate the longest common subsequence between the generated text Y and the reference text X, which is used to measure the degree of retention of reference information;

[0076] S2.232 Recall Calculation: ROUGE-L recall is calculated based on LCS length to evaluate the proportion of reference text successfully covered in the generated text.

[0077] The formula for the ROUGE-L index is:

[0078]

[0079] Where LCS stands for Longest Common Subsequence, X represents the use case reference result, and Y represents the model output;

[0080] S2.24 Field-level ROUGE Calculation

[0081] S2.241 Independent ROUGE Calculation for Each Field: ROUGE-1, ROUGE-2, and ROUGE-L are calculated separately for structured fields to evaluate the output quality of the model under different semantic dimensions;

[0082] S2.25 Results Summary and Output

[0083] S2.251 generates global and field-level reports: outputs overall and field-level ROUGE scoring results and visualizes them in the form of tables or charts;

[0084] S3. Modify module operation: Optimize text quality based on scoring results, automatically correct text that does not meet standards, and output standardized format data;

[0085] S3.1 The operation of the reference comparison and modification unit: Based on the scoring results, the deviation is identified and the consistency and expression quality of the generated text are automatically optimized;

[0086] S3.1 Operation of the reference comparison and modification unit: When a reference sample exists, the text is automatically optimized and generated based on the scoring results;

[0087] S3.11 Receiving the scoring comparison results: Receive the results from scoring modules such as BLEU or ROUGE, and identify the deviations between the generated text and the reference sample;

[0088] S3.12 Identify Deviations: Automatically locate and identify segments in the generated text that differ from the reference sample;

[0089] S3.13 Perform automatic correction: perform structural corrections on the text;

[0090] S3.14 Output modified text: Output the optimized text, ensuring that the difference from the reference sample is minimized;

[0091] S3.2 Operation of the No-Reference Comparison Modification Unit: In the absence of reference samples, the text content is independently verified and corrected based on the template and rules;

[0092] S3.21 Receive generated text: Receive uncompared model-generated text as input for subsequent verification and correction;

[0093] S3.22 Structured Template and Syntax Rule Validation: Performs compliance checks on text structure and format based on preset templates and rules;

[0094] S3.23 Field Integrity Check: Identify and complete any missing necessary fields in the generated text;

[0095] S3.24 Illegal Character Filtering: Detects and removes certain characters from the text;

[0096] S3.25 Semantic Logic Verification: Verify the logical and semantic rationality of the text, and correct the expression after discovering problems;

[0097] S3.26 Output the corrected text: Output the standard text corrected by the template rules;

[0098] S3.3 Automatic modification mechanism: When the text does not meet the quality standards, the system automatically starts the modification process for correction and optimization;

[0099] S3.31 Determine if the generated text meets the standard: Determine if the text meets the standard through field integrity and semantic logic checks;

[0100] S3.32 Automatic Editing: Automatically completes fields, corrects expressions, and cleans up content for text that does not meet the standards;

[0101] S3.33 Output modified text: Outputs the automatically corrected standard text;

[0102] S3.4 Standardized Content Output: Outputs the modified text in a standard structured format and pushes it to downstream systems;

[0103] S3.41 Standardized Format Output: Converts the corrected content into a unified structured data format;

[0104] S3.42 Push to downstream modules: Send the corrected text results to the business system or database to achieve data integration and final delivery;

[0105] S4. Operation of the manual verification and confirmation module: When the text quality does not meet the standard, the manual verification mechanism is triggered to confirm or modify the text through manual intervention, and the system operation process is recorded;

[0106] S4.1 Triggering the manual verification process: When the scoring result is below the threshold or automatic modification fails, the system automatically triggers the manual verification mechanism;

[0107] S4.11 Scoring Threshold Judgment: The system determines whether the BLEU or ROUGE score is lower than the preset threshold in order to decide whether to initiate manual verification;

[0108] S4.12 Modification module not meeting standards: When the automatic modification results do not meet the structured requirements, the system automatically enters the manual verification process;

[0109] S4.2 Automatic Data Extraction: The system automatically extracts key information related to the task from the input content or data source according to preset rules;

[0110] S4.21 Original Generated Text: Extract the unmodified original generated text as a basis for manual judgment;

[0111] S4.22 Modification Suggestion: Display the modified version generated by the rule engine for comparison and evaluation reference;

[0112] S4.23 Reference Sample: If a standard reference text exists, the system will extract and display it synchronously to assist manual comparison;

[0113] S4.24 Structured Comparison View: The original text, suggested modifications, and reference samples are displayed side by side in a comparative format, allowing for rapid human identification of differences and intervention in decision-making;

[0114] S4.3 Manual Intervention and Confirmation: Confirm, cancel, or manually modify the text according to task requirements to ensure output quality;

[0115] S4.31 Confirm Modification: Confirm the automatic modification suggestion and use it directly as the final text output;

[0116] S4.32 Undo Modification: If automatic modification does not meet expectations, the modification can be undone, and the original generated text can be retained;

[0117] S4.33 Manual Adjustment: Manually edit the original text or suggested modifications to ensure they conform to formatting specifications and semantic requirements;

[0118] S4.4 Feedback and Interaction Operation: The system provides feedback on modification results through the interactive interface, and synchronously records operation information to drive the optimization of subsequent processes;

[0119] S4.41 Interactive Interface Feedback: The text after confirmation, cancellation, or manual adjustment is fed back to the system through the interface;

[0120] S4.42 Record Operation History: The system records feedback content, operation time, and personnel information, making the entire process traceable;

[0121] S4.5 Verification Trajectory Recording and Auditing: The system fully records the entire verification trajectory process, enabling auditable operations and verifiable and traceable results;

[0122] S4.51 Complete Verification Track: The system records verification information in detail;

[0123] S4.52 Audit Function: Saves all operation records and supports auditing and process quality assessment;

[0124] S5. Operation of the data output module: Evaluate and format the generated text and its processing, ensure that the data conforms to the system interface standard and is reliably transmitted to the target system, and finally output the results and provide feedback;

[0125] S5.1 Data Processing and Scoring Comparison: The generated text is preprocessed and quantitatively evaluated by scoring, modification, and verification.

[0126] S5.1 Data Processing and Scoring Comparison Operation: Before output, the content quality is ensured to meet the standards through a triple guarantee of scoring, automatic modification and manual verification;

[0127] S5.11 Scoring Comparison: The system automatically scores the generated content to measure its degree of matching with standard templates or rules;

[0128] S5.12 Automatic Correction: Automatically corrects some content based on the scoring results;

[0129] S5.13 Manual review: Automatic correction of semantic or structural issues not covered by the system is done manually.

[0130] S5.2 Data Format Encapsulation Operation: The generated content is encapsulated into a standardized format according to preset standards;

[0131] S5.21 Selecting the Packaging Format: Select a suitable packaging format based on the requirements of the target system;

[0132] S5.22 Data Structuring: Organize data according to format specifications and construct a standard structure;

[0133] S5.23 Encapsulate Data: Encapsulate structured data according to the selected format to generate the final data file or data stream;

[0134] S5.3 Interface Protocol Adaptation Operation: Adapt the encapsulated data according to the interface protocol of the target system;

[0135] S5.31 Protocol Parsing: Identify and analyze the interface protocol type used by the target system;

[0136] S5.32 Protocol Conversion: Converts encapsulated data into the format and structure required by the target protocol;

[0137] S5.33 Interface Configuration: Configure the protocol adapter or middleware to ensure that data is transmitted correctly according to the protocol specifications;

[0138] S5.4 Data Transmission and Integration Operation: Reliably transmit and integrate the encapsulated and adapted data into industrial systems;

[0139] S5.41 Establish Connection: Establish a stable communication channel with the target system through network protocols;

[0140] S5.42 Data Transmission: Accurately transmits the encapsulated and adapted data to the target system;

[0141] S5.43 Integration into the Target System: Data integration is performed within the target system, and the data can be invoked and applied.

[0142] S5.5 Final Output and Feedback Operation: As the end of the human-computer interaction closed loop, it outputs the final result and provides feedback after the task is completed;

[0143] S5.51 Output Data Confirmation: Perform a quality review on the final data to ensure that the data meets the standards and industry requirements of the target system;

[0144] S5.52 Feedback Mechanism: When the output is abnormal or non-standard, the feedback process is triggered in a timely manner to notify for correction;

[0145] S5.53 Error Correction and Re-output: Correct the problematic data based on feedback and re-output it to the target system.

[0146] Secondly, a post-processing system for the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine includes the following modules and units:

[0147] The post-processing system includes an information recording module, a scoring comparison module, a modification module, a manual verification and confirmation module, and a data output module;

[0148] The information recording module is responsible for recording and storing model generation results, structured templates, and task type identifiers to enable data traceability.

[0149] The structured template management unit in the recorded information module categorizes and records text according to a preset structured set, ensuring that the generated text meets the predetermined structured requirements.

[0150] The task type identifier classification unit in the recorded information module is used to hierarchically label the explicit and implicit information of the model output. The explicit information includes basic information and hierarchical information, while the implicit information is extracted from the labeled content of the original input.

[0151] The scoring and comparison module evaluates the similarity between the generated text and the reference text using the BLEU and ROUGE indices, providing quantitative and visual analysis of the generated results.

[0152] The scoring comparison module includes a BLEU indicator unit and a ROUGE indicator unit;

[0153] The BLEU and ROUGE index units in the scoring comparison module are used to evaluate the similarity between the generated text and the reference text through n-gram matching, and support the analysis of the overall text and key fields.

[0154] The modification module performs format correction and semantic optimization based on the scoring comparison results to improve structural consistency and expression accuracy.

[0155] The modification module includes a modification unit with reference comparison and a modification unit without reference comparison;

[0156] The modification module includes a reference comparison modification unit that can automatically adjust the generated text by comparing BLEU and ROUGE indices.

[0157] The referenceless comparison modification unit in the modification module corrects the text based on structured templates and grammatical rules;

[0158] The manual verification module can intervene manually when the automatic score is below the threshold.

[0159] The data output module outputs the processed content to the industrial software system in a preset format, supporting multiple data format conversions and interface adaptations.

[0160] The data output module is responsible for encapsulating and converting the processed generated content according to a preset format (such as JSON, XML, CSV, etc.) to ensure that it meets the requirements of industrial applications and achieves seamless connection between different industrial application systems.

[0161] Thirdly, the post-processing system and optimization method of the human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention are applicable to the entire process of standardized output system in high-end intelligent equipment industrial scenarios.

[0162] High-end intelligent equipment industrial scenarios require equipment with complex requirements such as high precision, high automation, real-time monitoring, and feedback, such as cigarette manufacturing equipment and industrial automation equipment. This invention utilizes a human-computer interaction language model output post-processing system that supports field-level scoring comparison and rule engine to process and output data. This process goes beyond simply processing data from a single stage; it encompasses standardized processing throughout the entire process, from data acquisition to output. Standardized output ensures that data or information is output in a fixed, repeatable format to meet the needs of different systems or applications.

[0163] Fourthly, the post-processing system and optimization method for the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine, as described in this invention, are suitable for building intelligent agents and encapsulating the functional modules of the intelligent agent into independent components with interface definitions.

[0164] The present invention discloses a human-computer interaction language model output post-processing system and its optimization method that supports field-level scoring comparison and rule engine. This system not only improves the standardization, accuracy and operability of data processing, but also provides support for the construction of intelligent agents. That is, by encapsulating the functional modules of intelligent agents into independent components with interface definitions, this system achieves an efficient, flexible and scalable design that can cope with the needs of complex industrial environments.

[0165] (3) Inventive principle

[0166] The core principle of this invention, which supports field-level scoring comparison and rule engine-based human-computer interaction language model output post-processing system and its optimization method, is based on a multi-stage automated process combined with standardized data formats and interface protocols. This enables accurate generation, optimization, correction, and quality assessment of the output content of large language models. Human assistance is only introduced when the automated process cannot meet special needs, ensuring that the generated results meet the requirements of industrial applications and can be seamlessly integrated with the enterprise's existing information systems.

[0167] Firstly, in terms of data traceability and recording mechanisms, the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention adopts a structured recording and hierarchical annotation strategy to ensure that detailed information of each processing step, such as task type and output results, is fully recorded, which facilitates subsequent tracking and verification.

[0168] Specifically, the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention first records the entire process of input, output, intermediate generation results, scoring and modification suggestions of the large language model through the information recording module, thereby ensuring the traceability and verifiability of the data processing process, supporting full-link auditing from generation to output, and effectively ensuring the compliance of model results.

[0169] Secondly, regarding the automated scoring and comparison mechanism, the BLEU and ROUGE scoring mechanism of this invention, which supports field-level scoring comparison and rule engine-based human-computer interaction language model output post-processing system and optimization method, quantifies the similarity between the model-generated content and the reference sample by comparing n-gram precision, recall, and longest common subsequence, thereby providing a quantitative basis for subsequent modifications and ensuring that the quality of the generated text meets the requirements.

[0170] Thirdly, regarding the modification mechanism where automation is the primary driver and human intervention is secondary, the human-computer interaction language model output post-processing system and optimization method described in this invention, which supports field-level scoring comparison and rule engine, enables the automatic modification unit to perform structural adjustments and redundant information deletion based on hard-coded rules when the scoring result is below the threshold. Meanwhile, the manual verification module manually confirms and finally corrects the modification suggestions through a structured comparison view, thus forming a quality assurance mechanism that combines automation and human intervention.

[0171] Specifically, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention performs intelligent correction through reference-based comparison modification units and non-reference-based comparison modification units. At the same time, when the automatic correction fails to meet the standard, the system will trigger a manual verification and confirmation module to perform manual intervention, ensuring accuracy in complex semantic understanding and format calibration.

[0172] Fourth, in terms of result encapsulation and efficient integration, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention uses standardized data format encapsulation and protocol adaptation to ensure that the generated content can be seamlessly integrated into existing industrial systems, thus bridging the "last mile" from data generation to execution.

[0173] Specifically, the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention encapsulates the model-generated content in a preset format, such as JSON, XML, CSV, etc., through the data output module, and adapts and outputs it by supporting multiple industrial protocols, thereby achieving efficient integration with existing enterprise information systems (such as ERP, MES, APS, etc.).

[0174] (4) Effects of the invention: The present invention provides a human-computer interaction language model output post-processing system and its optimization method that supports field-level scoring comparison and rule engine. It proposes a system architecture and prototype software product as a whole, as well as a two-stage post-processing system architecture and logic.

[0175] This invention discloses a human-computer interaction language model output post-processing system and its optimization method that supports field-level scoring comparison and rule engine. It integrates and modifies the original methods used for grammar testing (ROUGE, BLEU metrics) for scoring across two dimensions: overall text and key metrics. This enables scoring and program setup for the "test phase post-processing logic" during the debugging phase. Furthermore, it provides the system architecture and logic for the "work phase post-processing logic" during the actual operation phase, proposing an integrated setup method. Additionally, the system's post-processing encoding is used for front-end application protection in industrial software.

[0176] First, it improves the quality and consistency of content generated by large language models, reducing errors, redundancy, and formatting issues in the output results, thus making the content of large models more reliable in industrial applications. The human-computer interactive language model output post-processing system and optimization method described in this invention, which supports field-level scoring comparison and rule engine, effectively improves the quality and structural consistency of the model output by introducing multi-level scoring comparison modules (such as BLEU and ROUGE metrics) and intelligent correction modules (including comparisons with and without references). The combination of automated correction and manual verification ensures no errors occur in complex semantic understanding and format constraints, thereby ensuring more accurate and precise generated results.

[0177] Secondly, it enhances the traceability and compliance of generated content, ensuring the system's usability in industries with high compliance requirements and strengthening the monitoring of model output. The human-computer interaction language model output post-processing system and optimization method described in this invention, which supports field-level scoring comparison and rule engine, comprehensively records the input, output, modification, and manual verification operations during the model generation process through a recording information module, ensuring data traceability at each processing stage; all generated content and related operations can be audited through log recording.

[0178] Third, efficient system integration and data flow improve the integration efficiency of model-generated content with enterprise information systems, reduce manual intervention, enhance enterprise automation and digitalization levels, and improve the compatibility and real-time performance of cross-platform data interaction. The human-computer interaction language model output post-processing system and optimization method described in this invention, which supports field-level scoring comparison and rule engine, utilizes standardized format encapsulation of the data output module and adapts to various industrial protocols (such as RESTful API, MQTT, OPC UA, etc.). The generated content can be directly integrated with existing industrial software systems (such as APS, MES, ERP).

[0179] Fourth, it supports diverse business scenarios and task requirements, enhancing the universality and flexibility of the post-processing mechanism. It can be widely applied to various business and industry scenarios (such as production scheduling and equipment status reports), strengthening the system's adaptability and scalability. The post-processing system and optimization method for human-computer interaction language model output, which supports field-level scoring comparison and rule engine, described in this invention, utilizes structured template management and task type identifier classification to flexibly configure and adapt to different types of tasks according to different business needs.

[0180] Specifically, in terms of model adaptation, the post-processing logic system described in this invention is independent of the language model itself and can flexibly adapt to language models with different parameter scales and types, without relying on a specific or existing model architecture; the post-processing logic module can also optimize the output of language models with various parameter volumes, and has broad adaptability and versatility.

[0181] Specifically, in the application of the two-stage post-processing logic, the post-processing logic system described in this invention uses "test stage post-processing logic" for scoring during the debugging phase to evaluate the rationality of the test programming. During the actual system operation phase, "working stage post-processing logic" is used to add, delete, and modify the results, including several stages such as recording, scoring, adding, deleting, and modifying, inputting into the subsequent system, and manual confirmation.

[0182] Fifth, the mechanism of automation and human collaboration ensures that the generated content still meets high standards even under complex tasks or abnormal data input, avoiding potential errors in fully automated systems and improving system robustness. The post-processing system and optimization method for human-computer interaction language model output, which supports field-level scoring comparison and rule engine, combines automated correction with human verification mechanisms. When automated correction fails to achieve ideal results, human intervention is effectively introduced to ensure accuracy and reliability in critical task scenarios. Attached Figure Description

[0183] Figure 1 This is a technical architecture diagram of the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention.

[0184] Figure 2 This diagram illustrates the differences before and after the use of "manual input + post-processing logic" in the industrial software processing flow related to the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention.

[0185] Figure 3 A bar chart showing an empirical multivariate comparison of three mainstream distillation models regarding the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine as described in this invention.

[0186] Figure 4 This diagram compares the effectiveness of the post-processing mechanism in a human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine, as described in this invention, across models of different scales. Detailed Implementation

[0187] The technical field of this invention is large model post-processing technology, which performs a series of adjustments and optimizations on the results generated by large models to ensure that the output text has preset standard structured features, make up for the problems existing in the model itself during the generation process, and improve the usability of the output results.

[0188] The technical feature of this invention lies in the secondary processing of the output results of a large model using pre-defined logic, achieving controllability and standardization of the model output results, particularly tailored to the needs of specific application scenarios. This invention adds multiple cache verification modules to enable operations such as editing, deleting, modifying, and replacing the model output results. This post-processing logic system can further optimize and personalize the response after the LLM generates an initial response, based on various factors such as contextual information, user preferences, and real-time data, thereby making the output more in line with the user's actual needs and expectations.

[0189] Figure 1 This is a technical roadmap diagram of a post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention. It shows a post-processing system for optimizing the human-computer interaction output of large language models (LLMs), specifically including the logical flow of multiple nodes and modules. The technical roadmap diagram includes a start node, a judgment node (diamond box), operation steps (rectangle box), a loop or repetition node, and an end node.

[0190] The starting node represents the beginning of the process, the starting point of the entire workflow. The process begins here, triggered by a specific event or data input. The system starts upon receiving user input or other external signals, beginning information collection or executing corresponding processing tasks. The information recording module, located at the beginning of the process, is responsible for receiving data input from the front-end user. Through this module, the user's raw natural language input, structured field requirement templates, and task type identifiers are captured and transformed into structured data. The information recording module is the foundation of the entire system; all subsequent processing modules rely on the structured data provided by this module. The information recording module transmits structured data to the scoring and comparison module, modification module, etc., ensuring that these modules can obtain correct input information and process it further.

[0191] The decision nodes (diamond-shaped boxes) are used to evaluate certain conditions. The process determines the next execution path based on the evaluation results; for example, checking if an input meets specific conditions, checking if the text similarity meets predetermined standards (such as BLEU and ROUGE scores), and determining if manual intervention or verification is needed. If the conditions are met, the process continues in a certain direction; if the conditions are not met, it jumps to another step or ends the process. Decision nodes are used to check conditions and determine whether the process continues or redirects to another path. In this flowchart, decision nodes are used to evaluate whether the similarity of the generated results (such as BLEU and ROUGE scores) meets preset standards. The scoring comparison module is responsible for calculating the similarity between the model-generated results and samples in the reference case library; the decision nodes make judgments based on the scoring results. If the score meets the standard, the process continues; if it does not, it triggers the modification or manual verification module.

[0192] The operation steps (rectangles) in the process represent specific operations or tasks. These steps involve actual work or processing, including information recording, scoring comparison, modification and optimization, and manual verification. Information recording involves collecting user input, structured field requirements, and task identifiers from the front-end interface. Scoring comparison involves scoring the similarity between the model-generated results and reference samples. Modification and optimization involve automatic modification, grammar correction, keyword insertion, and other operations based on the scoring results. Manual verification involves generating a comparison view for manual review when automatic correction fails to meet quality standards.

[0193] The recurring or repeating nodes are used when multiple similar tasks need to be performed repeatedly in the process, such as processing multiple data items or performing scoring comparisons multiple times. These steps are set as recurring nodes. Recurring or repeating nodes ensure that the same operations can be efficiently repeated when processing large amounts of data until all tasks are completed. In some steps, multiple similar operations or tasks may need to be performed repeatedly. For example, multiple documents or data samples may need to undergo the same scoring comparison, modification, and verification process. Recurring nodes allow these steps to be repeated until all data has been processed, ensuring that the system can complete tasks efficiently when processing large amounts of data.

[0194] The end node signifies the completion of the entire process, outputting the final result or sending a completion signal. At the end node, all data has been processed, modified, and verified, and the final result is pushed to downstream systems or a report is generated. Judgment nodes determine the branching or redirection of the process by checking conditions; if the conditions are met, subsequent steps continue; otherwise, the process jumps to the end node. The end node marks the termination of the entire process. Once the process reaches the termination conditions, the system outputs the final result or a completion signal, ending the entire workflow. The data output module is responsible for outputting the post-processed text results in a standardized format (such as JSON or XML) and pushing the results to downstream application systems or business modules. Simultaneously, the data output module records complete processing logs, scoring records, and modification paths for subsequent quality assessment and optimization. The data output module is the endpoint of the process; all corrected and confirmed results are output to external systems through this module. Throughout the process, the data output module receives the final data processed by other modules (such as a manual verification module) and encapsulates it into a structured format for output.

[0195] The core concept of the human-computer interaction language model output post-processing optimization system that supports field-level scoring comparison and rule engine, as described in this invention, is not to rely on modifying model parameters. Instead, it introduces post-processing logic such as field-level matching verification, syntax structure regularization, forced keyword insertion, and field integrity checks after the model output by designing configurable hard-coded rule modules. This effectively limits the format space of the generated text and improves the standardization and consistency of the output.

[0196] This invention constructs a system architecture guided by system coding rules for an innovative implementation method of human-computer interaction for large language models (LLMs) based on post-processing logic optimization. It includes a recording information module, a scoring comparison module, a modification module, a manual verification and confirmation module, and a data output module, which are used to perform structural constraints and format correction on the generated results of the large model.

[0197] like Figure 1As shown, this invention first uses an information recording module to connect to the front-end user interface, collecting user input and its corresponding contextual information in real time, including raw natural language input, structured field requirement templates, and task type identifiers. This information is then uniformly stored in a structured data format as the basic input for subsequent processing. Secondly, through a scoring and comparison module, the system uses text similarity evaluation metrics such as BLEU and ROUGE to perform multi-dimensional matching analysis between the model-generated results and historical samples in the reference case library. It calculates the similarity score between the field-level and overall text, and combines this with a preset threshold to determine whether the current output meets the consistency requirements of format and semantics. In the modification module, a built-in extensible encoding rule engine is included, containing functions such as field integrity checks, forced keyword insertion, syntax structure regularization, and illegal character filtering. Based on the scoring and comparison results, the system automatically triggers the corresponding rule set to correct the format and optimize the semantics of the generated text, ensuring that the output content meets the target structural requirements. To ensure the accuracy of outputs under critical tasks, the system incorporates a manual verification module. When the automatically modified output still fails to meet the set quality standards, or when sensitive operations are involved, the system generates a comparison view showing the differences between the original output, suggested modifications, and reference examples, assisting in rapid human decision-making and intervention. In the data output module, the system encapsulates the final post-processed output into a standard structured data format (such as JSON or XML) according to a preset interface protocol and simultaneously pushes it to the business system or other downstream application modules. At the same time, it stores complete processing logs, scoring records, and modification history in the database for subsequent quality assessment and system optimization.

[0198] This invention does not rely on modifying or retraining model parameters. Instead, it introduces post-processing logic such as field-level matching verification, syntax structure regularization, forced keyword insertion, and field integrity checks after the model output by designing a configurable hard-coded rule module. This effectively limits the format space of the generated text and improves the standardization and consistency of the output.

[0199] Figure 2 This is a diagram showing the differences before and after the use of "manual input + post-processing logic" in the industrial software processing flow related to the post-processing system and optimization method of human-computer interaction language model output supporting field-level scoring comparison and rule engine described in this invention. It includes two parts, which describe different process steps and decision nodes respectively.

[0200] like Figure 2As shown, the left-hand diagram illustrates the result of "manual input + post-processing logic" in an industrial software processing flow related to the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine as described in this invention. The "Input" box here represents the starting point of the data or task received at the beginning of the system. Looking at the "Manual Operation" section in the left-hand diagram, the human-computer interaction node indicates that a user or human operator will participate in this process, meaning the system relies on human judgment or input. The judgment node (diamond-shaped box) contains a decision point: "Should manual confirmation be performed?" It determines whether manual confirmation is required based on certain conditions. If the result is "yes," the manual confirmation operation will be triggered; if the result is "no," this step can be skipped and the process can continue to the next step. The end node indicates that the final process will enter the end state.

[0201] like Figure 2 As shown, the right-hand diagram represents the result of the "manual input + post-processing logic" step in the industrial software processing flow related to the human-computer interaction language model output post-processing system and optimization method for supporting field-level scoring comparison and rule engine as described in this invention. It contrasts this with another task processing flow related to the left-hand diagram. The difference between this and the left-hand flow is that this involves more systematic processing, does not involve manual confirmation, and directly continues to subsequent steps. The right-hand diagram appears more advanced than the left-hand diagram, mainly because it demonstrates the automation and efficiency of the process.

[0202] Figure 3 The bar chart presents a comparison of binary classification tasks for three mainstream distillation models, illustrating the post-processing system and optimization method for human-computer interaction language models that support field-level scoring comparison and rule engine as described in this invention. It shows the evaluation results of different datasets under several categories. The x-axis represents different categories or groups, and each group represents different comparison dimensions, such as different test conditions, models, algorithms, etc. The y-axis represents the metrics under each category, such as accuracy, precision, recall, etc. Figure 3 The bar chart clearly shows the performance comparison of multiple datasets or methods under different categories. The y-axis shows the evaluation results of each method in each category, while the x-axis represents different test categories or conditions. The bar chart can intuitively show the advantages and disadvantages of each method and its performance under different categories.

[0203] like Figure 3 As shown, Deepseek 7B, Deepseek 14B, and Deepseek 32B represent three orders of magnitude distillation models; t-bleu represents the overall BLEU index, p-bleu represents the key field BLEU index, t-rouge represents the overall ROUGE index, and p-rouge represents the key field ROUGE index.

[0204] Figure 4 This image shows a comparison of the content quality scoring performance of the post-processing mechanism in a human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine, as described in this invention, across models of different sizes. It illustrates the performance comparison of large distillation models of different sizes (Deepseek7B, Deepseek14B, Deepseek32B) before and after applying the hard-coded post-processing mechanism. The evaluation results for each model include multiple metrics (such as BLEU and ROUGE), and the change in each metric after post-processing ("-h" version) is clearly indicated. Firstly, these are models of different sizes; a larger number indicates a larger model size, with correspondingly greater computational power and a greater number of parameters. For example, 7B represents a model with approximately 7 billion parameters, 14B represents approximately 14 billion parameters, and 32B represents approximately 32 billion parameters. Secondly, these are the model versions after applying the post-processing mechanism, corresponding to the original models (7B, 14B, 32B) after hard-coded post-processing logic; the post-processing logic can improve the generation quality of the model, especially in the processing of text structure and key fields.

[0205] The percentages following each metric (e.g., +5.42%, +12.75%) indicate the improvement after applying post-processing. For example, Deepseek7B-h shows a +5.42% improvement on t-BLEU, meaning that the model-generated text showed a significant increase in similarity to the reference text after applying post-processing. Larger models (such as Deepseek32B), while having higher raw performance, still benefit from post-processing, particularly in t-ROUGE and p-ROUGE scores, which helps improve the recall and structural canonicity of text generation.

[0206] like Figure 4 As shown, the post-processing mechanism demonstrated positive effects on models of different sizes, with the greatest improvement observed in medium-sized models (Deepseek14B). Smaller models (Deepseek7B) also benefited significantly, showing marked improvements, particularly in the processing of key fields. The post-processing logic's enhancement of the generated text quality and structural consistency makes these models more suitable for industrial applications (such as order entry, production planning, and equipment status report generation), ensuring that the generated text is not only semantically accurate but also conforms to specific formatting requirements.

[0207] Example 1: A post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention.

[0208] The present invention describes a human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine. It constructs a system architecture guided by coding rules for a human-computer interaction innovation implementation method of large language models (LLMs) based on post-processing logic optimization. The system includes a recording information module, a scoring comparison module, a modification module, a manual verification and confirmation module, and a data output module, which are used to perform structural constraints and format correction on the generated results of the large model.

[0209] The technical approach of this invention—a post-processing system and optimization method for human-computer interaction language model output supporting field-level scoring comparison and rule engine—is as follows: First, an information recording module is used to connect to the front-end user interface, collecting user input content and its corresponding contextual information in real time, including raw natural language input, structured field requirement templates, and task type identifiers. This information is then uniformly stored in a structured data format as the basic input for subsequent processing. Second, through the scoring comparison module, the system performs multi-dimensional matching analysis between the model-generated results and historical samples in the reference case library based on text similarity evaluation indicators such as BLEU and ROUGE. It calculates the similarity score between the field-level and overall text, and combines this with a preset threshold to determine whether the current output meets the consistency requirements of format and semantics. In the modification module, a built-in extensible encoding rule engine is incorporated, including functions such as field integrity checking, forced keyword insertion, grammatical structure regularization, and illegal character filtering. Based on the scoring comparison results, the system automatically triggers the corresponding rule set to correct the format and optimize the semantics of the generated text, ensuring that the output content meets the target structural requirements. To ensure the accuracy of outputs under critical tasks, the system introduces a manual verification module. In the data output module, the system encapsulates the final output results after post-processing into a standard structured data format (such as JSON and XML) according to a preset interface protocol, and pushes it synchronously to the business system or other downstream application modules. At the same time, the complete processing logs, scoring records and modification history are stored in the database for subsequent quality assessment and system optimization.

[0210] Firstly, the human-computer interaction language model output post-processing system described in this invention, which supports field-level scoring comparison and rule engine, has each module responsible for different functions to ensure the accuracy, consistency, and structural standardization of the generated content. Especially in industrial applications, it ensures efficient integration and data flow with existing information systems. Specifically, it is divided into the following modules:

[0211] Module 1. Record Information Module: The information recording module is mainly responsible for recording and storing model generation results, structured field requirement templates, task type identifiers, and post-processing intermediate content, thereby enabling data traceability.

[0212] Module 2. Scoring Comparison Module:The scoring and comparison module is mainly responsible for evaluating and comparing the overall content and key fields of the generated content, so as to quantify and visualize the generated results.

[0213] Module 3. Modify the module: The modification module, based on the output of the scoring comparison module, performs format correction and semantic optimization on the generated content to ensure that the generated text meets standardized output requirements and improves structural consistency and expression accuracy.

[0214] Module 4. Manual Verification and Confirmation Module: The manual verification module is mainly responsible for manually intervening when the automatic score is lower than the set threshold to ensure the accuracy and reliability of the output results.

[0215] Module 5. Data Output Module: The data output module is mainly responsible for packaging the processed generated content into a preset format and outputting it to the subsequent industrial software system. This ensures that the generated content meets the requirements of industrial applications, supports the conversion of various standard data formats such as JSON, XML, and CSV, and can be adapted to output according to the target system's interface protocol (such as RESTful API, OPCUA, MQTT, etc.).

[0216] Secondly, this invention presents a post-processing system for human-computer interaction language model output that supports field-level scoring comparison and rule engine. It proposes a system architecture and prototype software product, and details a two-stage post-processing system architecture design. The specific architecture and logic of the two-stage post-processing system are as follows:

[0217] Phase 1: Debugging Phase (Post-testing logic) The debugging phase focuses on using scoring comparison modules (such as BLEU and ROUGE metrics) to evaluate the generated content, help identify differences and quality issues in the generated text, and ensure that the generated text meets the predetermined format requirements.

[0218] During the debugging phase, the core task is to evaluate and optimize the quality of the text generated by the model, and to ensure that the text meets the preset standards through scoring, modification and manual intervention.

[0219] The debugging phase mainly focuses on the quality assessment of the model-generated content and the debugging and optimization of the post-processing logic. The goal of this phase is to conduct systematic testing through the model generation results to ensure that the generated text conforms to standards in terms of syntax, structure, and semantics.

[0220] The debugging phase mainly includes modules such as an information recording module, a scoring comparison module, a modification module, and a manual verification and confirmation module, which are used for evaluation, correction, and manual intervention.

[0221] Phase Two: Actual Operation Phase (Post-Processing Logic of the Working Phase)In the actual operation phase, the post-processing system will enter the actual working environment and be applied to specific industrial tasks, such as production planning and equipment status reporting. This requires that the content generated by the text not only conforms to the syntax and structure, but also meets specific format and semantic requirements so that it can be correctly used by downstream industrial systems (such as MES, APS, etc.).

[0222] In the actual operation phase, the post-processing logic enters the actual application scenario, processes the text content generated by the model, and ensures that it meets the actual industrial application requirements. This phase focuses on the formatting, semantic accuracy, and controllability of the generated content, and ensures that the text conforms to industry standards.

[0223] In the actual operation phase, the focus of the post-processing logic is to ensure that the generated text content meets the format, structure, and semantic requirements of industrial applications, and can support actual industrial production and management tasks.

[0224] The actual operation phase mainly includes modules such as an information recording module, a scoring and comparison module, a modification module, a manual verification and confirmation module, and a data output module, which are used to ensure that the generated content meets the requirements of industrial-grade applications and is output to the system.

[0225] Thirdly, the post-processing optimization method for the human-computer interaction language model output that supports field-level scoring comparison and rule engine according to the present invention specifically includes the following steps:

[0226] S1. Operation of the recording information module S1.1 Result Recording Unit Operation: Records model generation results, structured field requirement templates, task type identifiers, and post-processing intermediate generated content to ensure data traceability; then, the generated text results are recorded in detail through the result recording unit.

[0227] S1.2 Operation of the Structured Template Management Unit: Using the structured template management unit, text is classified and identified according to the preset structured templates to ensure that the text meets specific structural requirements;

[0228] S1.3 Operation of the Task Type Identifier Classification Unit: The explicit and implicit information of the generated content is marked through the Task Type Identifier Classification Unit. The explicit information includes basic information such as classification and number, while the implicit information includes internal system marking content such as the number of input characters.

[0229] S2. Operation of the scoring comparison module: S2.1 Operation of BLEU index unit: Using BLEU index unit, the similarity between generated text and reference text is evaluated through n-gram matching to ensure that the generated text meets the expected requirements in terms of vocabulary usage and word order structure;

[0230] S2.2 Operation of the ROUGE indicator unit: The ROUGE indicator unit is used to calculate the recall rate of the generated text and the reference text, evaluate the coverage of the generated text to the reference content, especially the recall ability of the key information; then, the overall quality score of the generated text is performed, and field-level analysis of key fields (such as "product name", "delivery time" etc.) is supported to identify the parts that perform poorly in key fields.

[0231] S3. Modify the module's operation: S3.1 Operation of the Reference Comparison and Modification Unit: Based on the results of the scoring comparison module, the reference comparison and modification unit is started to identify the differences between the generated text and the reference text, and automatically perform operations such as keyword insertion, redundant information deletion, and sentence structure adjustment to optimize the consistency of the text.

[0232] S3.2 No-reference Comparison Modification Unit Operation: In the no-reference comparison modification unit, if there is no reference sample, the system will independently verify and correct the generated content according to the predefined structured template and syntax rules; these corrections include field integrity checks, illegal character filtering, semantic logic verification, content completion, error correction, and invalid content removal, etc.

[0233] S3.3 Automatic modification mechanism: In the system, based on specific conditions, such as if the quality assessment of the generated text does not meet the preset standards, the modification mechanism will be automatically activated to correct and optimize the relevant content to ensure that the generated content meets the predetermined quality standards and format requirements.

[0234] S4. Operation of the manual verification and confirmation module: S4.1 Triggering the manual verification process: When the BLEU or ROUGE index output by the scoring comparison module is lower than the set threshold, the manual verification confirmation module is triggered; at this time, the system will automatically extract the original generated text, modification suggestions and reference samples, and display them to the operator in a structured view;

[0235] S4.2 Automatic Data Extraction: The system automatically extracts key information related to the current task from the input content or data source according to set rules or conditions;

[0236] S4.3 Manual Intervention and Confirmation Operation: The system records the complete verification process, including modification comments, confirmation time, and operator identity information, to ensure that the operation is traceable and the results are auditable;

[0237] S4.4 Feedback and Interaction Operation: During system execution, there is a two-way interaction between the user and the system. The system provides feedback information based on previous operations, and the user can make decisions or adjust inputs based on the feedback, triggering subsequent operations.

[0238] S4.5 Verification Trajectory Recording and Auditing: The system records the complete verification trajectory, including modification comments, confirmation time, and operator identity information, ensuring that the operation is traceable and the result is auditable;

[0239] S5. Operation of the data output module: S5.1 Data Processing and Scoring Comparison: In the system, the generated text and reference text are preprocessed and quantitatively evaluated;

[0240] S5.2 Data Format Encapsulation Operation: After the generated content has been scored, compared, modified, and manually verified, the system will encapsulate the text that meets the requirements according to the preset format and output it to the industrial software system. The data output module supports the encapsulation and conversion of various standardized data formats.

[0241] S5.3 Interface Protocol Adaptation Operation: Adapt according to the target system's interface protocol, such as RESTful API, OPC UA, MQTT, etc.

[0242] S5.4 Data Transmission and Integration Operation: Ensure that the generated content can be seamlessly integrated with the enterprise's existing Advanced Planning and Scheduling System (APS), Manufacturing Execution System (MES), and Enterprise Resource Planning System (ERP);

[0243] S5.5 Final Output and Feedback Operation: After completing all processing steps, the system generates the final result and outputs it to the target system or presents it to the user; at the same time, the system will also provide feedback information to help users understand the quality of the result or problems in the generation process, and adjust subsequent steps based on the feedback.

[0244] This invention discloses a post-processing optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine. It achieves post-processing optimization of model output through multiple steps, each playing a crucial role in ensuring text generation quality and format consistency. Especially in industrial applications, the accuracy of text content, structural standardization, and processing of key information are paramount. Therefore, each step combines automation and manual intervention to ensure that the generated text meets the requirements of practical applications.

[0245] Example 2: Specific Implementation of the "Record Information Module" in the Human-Computer Interaction Language Model Output Post-Processing System and Optimization Method Supporting Field-Level Scoring Comparison and Rule Engine of the Present Invention.

[0246] The "record information module" in the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention, provides a basic framework for subsequent processing and tracing.

[0247] Specifically, the "recording information module" of this invention aims to record model generation results, structured field requirement templates, task type identifiers, and post-processing intermediate generated content, thereby achieving data traceability. The "recording information module" of this invention mainly includes a result recording unit, a structured template management unit, and a task type identifier classification unit.

[0248] The result recording unit is able to completely record relevant information from model-generated results and post-processing intermediate content;

[0249] The task type identifier classification unit can use a hierarchical annotation strategy to annotate the explicit and implicit information of the model respectively.

[0250] The explicit information includes the basic information and hierarchical information entered; the basic information includes information such as preset categories and log numbers; the hierarchical information includes information such as object text structure and object attributes.

[0251] The implicit information is extracted from the basic marker content of the original input, such as the number of characters at the input end of the large model.

[0252] First, the operation of the "recording information module" in the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention can be completed in the following specific steps:

[0253] S1.1 Operation of the result recording unit: The purpose of operating the result recording unit is to completely record the model generation results and intermediate content of the post-processing stage;

[0254] S1.11 Capture Model Generation Results: Record the final text content generated from the model, including all text results and field information generated by the model.

[0255] The capture model's generated results need to ensure that every piece of information output by the model is recorded, and all generated text has a complete storage record to provide basic data for subsequent tracing and analysis;

[0256] S1.12 Capture intermediate content generated during post-processing: Record intermediate text and data generated during post-processing, such as corrected text, scoring results, and generated temporary fields, to ensure that intermediate data is not lost and to facilitate later debugging and optimization;

[0257] S1.2 Operation of the Structured Template Management Unit: S1.21 Structured Classification of Generated Content: Based on the preset structured set, the generated text, fields and data are classified, managed and identified. Different types of generated content and fields are collected and stored according to the preset structure and recorded.

[0258] S1.22 Structured storage: Ensures that each record conforms to a predetermined format and standard, facilitating querying and comparison;

[0259] The structured storage is intended to facilitate subsequent queries and tracing, especially when dealing with complex data, where maintaining format consistency is crucial;

[0260] S1.3 Operation of the Task Type Identifier Classification Unit

[0261] The purpose of the hierarchical annotation of explicit and implicit information is to extract and distinguish different types of information by hierarchically annotating explicit and implicit information, and to provide rich context for subsequent processing.

[0262] The explicit information annotation includes basic information annotation and hierarchical information annotation. The basic information includes preset categories, log numbers, and other information, which directly reflect the explicit content of the model's input and output. The hierarchical information includes object text structure, object attributes, and other information, which helps to understand the logical framework of the generated text. The implicit information annotation is extracted from the original input content, such as the number of characters at the large model input end and the temporal information during input, and then this information is not usually directly present in the model's explicit results. The implicit information annotation helps to understand the background of the generated results and is crucial to the model's generation process.

[0263] S1.31 Labeling Explicit Information: Based on explicit information, including basic information input by the user, structural hierarchy, classification number, etc., a hierarchical labeling method is adopted;

[0264] S1.32 Annotate implicit information: Extract implicit information from the original input data, such as the number of characters and the input length;

[0265] Although the implicit information is not intuitive, it is crucial for the analysis task and can help the system analyze and predict the task type. For example, changes in input length are related to certain processing requirements or affect subsequent generation.

[0266] S1.33 Annotation level information: Annotation level information includes text structure, attributes and other information to help the system understand the generated content from multiple dimensions;

[0267] The results obtained through the "recording information module" in the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine, as described in this invention, include complete text generation data, intermediate data in the post-processing process, structured fields, classification and storage of task identifiers, etc., providing a solid foundation for subsequent data tracing, analysis, optimization, quality control, etc.

[0268] Example 3: Specific Implementation of the "Scoring Comparison Module" in the Human-Computer Interaction Language Model Output Post-Processing System and Optimization Method Supporting Field-Level Scoring Comparison and Rule Engine as described in this invention.

[0269] The "scoring comparison module" in the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine, as described in this invention, aims to evaluate and compare the overall generated content and key fields. This quantifies and visualizes the differences between the generated text and the reference text, facilitating accurate evaluation of the quality of the model's generated results. The "scoring comparison module" of this invention mainly includes BLEU indicator units and ROUGE indicator units.

[0270] The BLEU metric unit mainly measures the similarity between the generated content and the ideal output in terms of vocabulary usage and word order structure by performing n-gram matching between the model-generated text and the standard output in the reference sample, calculating the word overlap rate between the two.

[0271] The BLEU indicator unit can support scoring the generated text as a whole, and can extract key fields such as "product name", "planned quantity" and "delivery time" separately to perform field-level BLEU analysis in order to identify the performance differences of each field in the structured output of the model.

[0272] The main tasks of the BLEU index unit include: firstly, measuring the similarity between the generated text and the reference answer in terms of vocabulary and word order structure; secondly, quantifying the consistency between the model output and the ideal answer in terms of content expression through n-gram matching; and thirdly, supporting overall scoring and detailed field-level analysis to discover differences in the model output in different structural parts.

[0273] The ROUGE metric unit mainly considers the recall rate of the generated results, focusing on evaluating the overlapping n-gram units, n-gram pairs, and longest common subsequence between the model-generated text and the reference samples, covering multiple evaluation modes such as ROUGE-1, ROUGE-2, and ROUGE-L.

[0274] The main tasks of the ROUGE indicator unit include: firstly, measuring the content coverage (recall) and word order similarity between the generated text and the reference answer; secondly, focusing on evaluating whether the generated text contains key information from the reference answer, especially the overlap of n-grams and the longest common subsequence; and thirdly, supporting multidimensional analysis at the overall and field levels.

[0275] The operation of the "scoring comparison module" in the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention can be completed in the following specific steps:

[0276] S2.1 Operation of the BLEU indicator unit: S2.11 Input data processing

[0277] S2.111 Model Output and Reference Sample Preparation: Collect the output text generated by the model and the corresponding reference samples for standard output;

[0278] The purpose of the model output and reference sample preparation is to compare the generated content with the standard answer and calculate their similarity.

[0279] S2.112 Field-level extraction: Extract key fields such as "product name", "planned quantity" and "delivery time" from the generated text and reference text for specialized field-level analysis;

[0280] The purpose of the field-level extraction is to conduct in-depth analysis of the performance differences of key fields;

[0281] S2.12 Lexical overlap rate calculation: S2.121 n-gram matching: Perform n-gram matching between the model-generated text and the reference text, and calculate the accuracy (i.e., overlap) of each n-gram;

[0282] The n-gram is a series of n consecutive lexical units, commonly 1-gram (a single word) and 2-gram (two consecutive words);

[0283] S2.122n-gram accuracy P n Calculation: Calculate the accuracy of different n-grams (such as 1-gram, 2-gram), that is, the proportion of the same n-gram in the model output and the reference text;

[0284] The accuracy formula is as follows:

[0285]

[0286] S2.123n-gram weight ω n Settings: Assign weights ω to different n-grams based on their precision. n That is, the weight is set according to the contribution of different n-grams. A common practice is to give higher weights to low n values, such as 1-gram and 2-gram.

[0287] The weight settings for the four n-grams are ω1 = 0.25, ω2 = 0.25, ω3 = 0.25, and ω4 = 0.25, respectively.

[0288] S2.13 BLEU Score Calculation

[0289] S2.131 Calculate the BLEU value: based on n-gram precision Pn The BLEU value is calculated using the weight ωn, and the specific formula is as follows:

[0290]

[0291] Where r represents the length of the reference test case; c represents the length of the model output; P n Represents the precision of the n-gram; ω n Weights representing the precision of n-grams;

[0292] S2.132 Calculate overall and field-level BLEU scores: Calculate the overall BLEU score of the generated text according to the BLEU formula, and perform field-level BLEU scoring as needed;

[0293] The field-level scoring is intended to identify performance differences on specific structured information, such as "product name".

[0294] S2.14 Results Output and Visualization

[0295] S2.141 Results Presentation: The calculated overall BLEU score and field-level scores are output and visualized to help users analyze and understand the differences between the generated text and the reference text;

[0296] The visualization can be presented in the form of charts, tables, etc., highlighting the matching of different n-grams.

[0297] S2.2 Operation of the ROUGE index unit: The ROUGE index unit not only supports overall evaluation, but also field-level analysis. It focuses on recall and emphasizes the situation where information in the reference sample is "covered". That is, the ROUGE index unit is not just a simple recall calculation, but a comprehensive analysis system that includes "multi-granularity (n-gram) + sequence integrity (LCS) + field-level robustness".

[0298] S2.21 Input Data Processing: S2.211 Model Output and Reference Sample Preparation: Prepare the model-generated text Y to be evaluated and the corresponding reference text X. The text is preprocessed, including word segmentation or character segmentation, to maintain a consistent standard and avoid spaces and punctuation affecting the matching.

[0299] S2.212 Field-level data extraction: Some data requires field-level analysis. Specific fields, such as title, abstract, and date, are extracted from the text in advance.

[0300] The purpose of the field-level data extraction is to facilitate the subsequent calculation of the ROUGE index for different fields and the analysis of semantic dimension stability.

[0301] S2.22n-gram extraction and matching

[0302] The ROUGE-N formula is as follows:

[0303]

[0304] Extraction of the Longest Common Subsequence (LCS) in S2.23

[0305] S2.231 LCS Calculation: Calculate the longest common subsequence (LCS) between the reference text X and the generated text Y;

[0306] The length of the longest common subsequence (LCS) indicates the degree to which reference text information is "preserved" in the generated text;

[0307] S2.232 Recall Calculation: Perform LCS recall calculation;

[0308] The specific formula for the LCS recall rate (ROUGE-LRecall) is as follows:

[0309]

[0310] That is, what proportion of sequence information was recalled by model Y from reference text X;

[0311] The specific formula for the ROUGE-L index is as follows:

[0312]

[0313] Where LCS stands for Longest Common Subsequence, X represents the use case reference result, and Y represents the model output;

[0314] S2.24 Field-level ROUGE Calculation

[0315] S2.241 Independent ROUGE Calculation for Each Field: For structured fields (such as abstracts and key sentences), ROUGE-1, ROUGE-2, and ROUGE-L are calculated separately.

[0316] The independent ROUGE calculation for each field allows the ROUGE scores of different fields to reveal the output quality of the model under different semantic dimensions, reflecting consistency and stability in key details;

[0317] S2.25 Results Summary and Output

[0318] S2.251 generates global and field-level reports: it outputs the overall ROUGE-1, ROUGE-2, and ROUGE-L scores; if field-level analysis is performed, it also outputs the ROUGE index for each field. It also supports visualization, using tables, bar charts, etc. to display the index scores of different granularities and fields, helping developers quickly identify the model's strengths and weaknesses.

[0319] The "scoring comparison module" in this invention, which supports field-level scoring comparison and rule engine-based human-computer interaction language model output post-processing system and optimization method, not only provides an overall score for the generated content but also encompasses multiple detailed analyses at the field level and semantic dimensions. This provides quantitative and visualized evaluation data, supporting developers and business personnel in making scientific decisions regarding the optimization of generated text quality, quality control, and task feedback. These results provide necessary data support for the continuous optimization of generative models, system feedback, and version comparison.

[0320] Example 4: Specific Implementation of the "Modification Module" in the Human-Computer Interaction Language Model Output Post-Processing System and Optimization Method Supporting Field-Level Scoring Comparison and Rule Engine of the Present Invention.

[0321] The "modification module" in the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention aims to perform format correction and semantic optimization on the content generated by large-scale language models based on the scoring comparison results, thereby improving the structural consistency and expressive accuracy of the output text. The "modification module" of this invention, by making targeted adjustments to the overall structure and key fields of the generated content, ensures that it meets preset standardized output requirements, thereby enhancing the practicality and controllability of the model in industrial scenarios.

[0322] The operation of the "modification module" in the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention can be completed in the following specific steps:

[0323] S3.1 Operation with reference comparison and modification unit

[0324] The purpose of the reference comparison modification unit is to automatically optimize the generated text based on the results of the scoring comparison module (such as BLEU, ROUGE) when a reference sample is available.

[0325] The reference comparison and modification unit is used to identify deviations in the generated text based on the scoring comparison results and to perform correction operations such as keyword insertion, redundant information deletion, and sentence structure adjustment.

[0326] S3.11 Receiving the scoring comparison results: Receive the similarity scores between the model-generated text and the reference sample from the scoring comparison module (such as BLEU and ROUGE scores), and then evaluate the differences between the text generation results and the reference sample to find the parts of the text that have deviations.

[0327] S3.12 Identify Deviations: Based on the scoring results, automatically locate the parts that deviate significantly from the reference sample, and identify inconsistencies, redundant content, grammatical errors, and other issues in the generated text;

[0328] S3.13 Perform automatic correction operations: insert keywords, delete redundant information, and adjust sentence structure;

[0329] S3.14 Output the modified text: Output the optimized text and ensure that the difference between it and the reference sample is minimized, thus ensuring structural consistency and expression accuracy;

[0330] S3.2 Operation of the Modified Unit without Reference Comparison

[0331] The purpose of the operation of the referenceless comparison modification unit is to perform text verification and correction based on predefined templates and syntax rules in the absence of reference samples.

[0332] The no-reference comparison modification unit is used to perform field integrity checks, illegal character filtering, and structural verification in the absence of reference samples.

[0333] S3.21 Receive the generated text: The no-reference comparison modification unit directly receives the generated text, without a reference sample for comparison, but the text still needs to be optimized;

[0334] S3.22 Perform structured template and syntax rule validation: Use predefined structured templates and syntax constraint rules to validate the generated text to ensure that it meets the target format requirements;

[0335] The structured template and syntax rule validation is performed, for example, checking whether the text contains the required fields and whether it conforms to specific format requirements, such as date, number, etc.

[0336] S3.23 Field Integrity Check: Verify whether there are any missing fields in the generated text, such as missing necessary category fields or data fields, and complete these fields;

[0337] S3.24 Illegal Character Filtering: Check whether the generated text contains illegal characters, such as garbled text, incorrect symbols, etc., and filter and correct them.

[0338] S3.25 Semantic Logic Verification: Verify the semantics of the text to ensure that the logic and structure of the text are correct; if logical inconsistencies or errors are found, correct them.

[0339] S3.26 Output the corrected text: Return the corrected text, ensuring that its format meets the requirements and its expression is accurate;

[0340] S3.3 Automatic Triggering of Modification Mechanism

[0341] The purpose of the automatic modification mechanism is to trigger the automatic modification mechanism when the text format does not meet the requirements, and to correct missing fields, incorrect expressions or invalid content.

[0342] The automatic modification mechanism automatically completes missing fields, corrects erroneous expressions, and deletes invalid content when the generated text does not meet the requirements.

[0343] S3.31 Determine whether the generated text conforms to the standard: In scenarios where there is no reference sample, determine whether the generated text conforms to the preset standard through field integrity checks, semantic verification, etc.

[0344] S3.32 Automatic Modification: When the text does not meet the target format or quality standards, the modification mechanism is automatically triggered;

[0345] The modification mechanism includes: completing missing fields; correcting incorrect expressions or inaccurate content; and removing invalid content, such as duplicate or meaningless parts.

[0346] S3.33 Output modified text: Ensure that the text content meets the target format requirements after correction, and output the final result;

[0347] S3.4 Output Standardized Content

[0348] The purpose of standardizing the output content is to ensure that the modified content meets the standard format requirements and can be seamlessly output to downstream systems.

[0349] The standardized output content refers to outputting the corrected content in a standard format and pushing it to the downstream system to complete the entire modification process.

[0350] S3.41 Standardized Format Output: Outputs the corrected text in a standardized structured data format for subsequent processing and application;

[0351] S3.42 Push to downstream modules: Push the processed text to other systems or downstream modules, such as business systems or databases, to complete the terminal delivery of the text.

[0352] The post-processing system and optimization method for human-computer interaction language model output, which supports field-level scoring comparison and rule engine, described in this invention, achieves optimized generated text through the operation of the "modification module." This optimized text undergoes meticulous corrections in structure, content, format, and syntax, conforming to predetermined standardization requirements, and can be seamlessly integrated into downstream systems. Field-level optimization and multi-dimensional scoring feedback make text generation more accurate and reliable, ensuring high-quality output. Finally, all correction operations and generated content are traceable, providing a solid foundation for data traceability and subsequent optimization.

[0353] Example 5: Specific Implementation of the "Manual Verification and Confirmation Module" in the Human-Computer Interaction Language Model Output Post-Processing System and Optimization Method Supporting Field-Level Scoring Comparison and Rule Engine as described in this invention.

[0354] The "manual verification and confirmation module" in the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine, as described in this invention, aims to manually intervene and finally confirm generated content whose automatic scoring results are below a set threshold, ensuring the accuracy and reliability of output results in critical task scenarios. This module, as a quality assurance link in the post-processing workflow, effectively compensates for the limitations of automated modification strategies in complex semantic understanding and format constraints. The "manual verification and confirmation module" of this invention mainly includes a verification triggering unit, an information extraction and display unit, a manual intervention operation unit, and a verification result recording unit.

[0355] The verification triggering unit is mainly used to monitor the BLEU or ROUGE index output by the scoring comparison module, thereby determining whether it is below the preset threshold or whether the modification module does not meet the structured output requirements, and further automatically triggering the manual verification process.

[0356] The main function of the manual intervention unit is to provide an interactive interface that allows for "confirmation", "cancellation" or "manual adjustment" functions, namely, "confirmation" to accept the suggestion, "cancellation" to retain the original text, and "manual adjustment" to modify the text manually.

[0357] The operation of the "manual verification and confirmation module" in the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention can be completed in the following specific steps:

[0358] S4.1 Triggering the manual verification process: The manual verification process is triggered automatically when the scoring result is lower than the threshold or the modification module fails to meet the standard.

[0359] The purpose of triggering the manual verification process is to trigger the manual verification process when the automatic scoring result is lower than the set threshold, such as BLEU or ROUGE being lower than 0.6 to 0.99, or when the modified module fails to fully meet the structured output requirements.

[0360] S4.11 Scoring Threshold Judgment: The system first determines whether the scoring comparison result is lower than the preset threshold;

[0361] S4.12 Modification module not meeting the requirements: If the automatic correction results of the modification module cannot meet the target format requirements, the system will automatically enter the manual verification process;

[0362] S4.2 Automatic Extraction of Relevant Data: The automatic extraction of relevant data involves extracting the original generated text, modification suggestions, and reference samples, and displaying them in a structured comparison view.

[0363] The purpose of automatically extracting relevant data is to provide operators with necessary comparison information during the manual verification process so that effective intervention can be carried out.

[0364] S4.21 Original Generated Text: The system extracts the original generated text without modification;

[0365] S4.22 Modification Suggestions: The system displays modification suggestions generated by the rules engine, i.e., the automatically modified text;

[0366] S4.23 Reference Sample: If a reference sample is available, the system will simultaneously display the corresponding reference text for operators to compare.

[0367] S4.24 Structured Comparison View: The original text, modification suggestions, and reference samples are displayed in a structured comparison view, which makes it easier for operators to quickly identify problems and make decisions.

[0368] S4.3 Operation of manual intervention and confirmation: The manual intervention and confirmation refers to the operator confirming, canceling or manually adjusting the generated text;

[0369] The purpose of the manual intervention and confirmation is for operators to manually intervene in the modification suggestions, and to confirm, withdraw or adjust them according to the task requirements, so as to ensure the accuracy and consistency of the text.

[0370] S4.31 Confirm Modification: Operators can choose to confirm the automatic modification suggestion and use it as the final output;

[0371] S4.32 Undo Modification: If the operator believes that the automatic modification does not meet the requirements, the modification can be undone, and the original text can be retained;

[0372] S4.33 Manual Adjustment: Operators can manually adjust the generated text or modification suggestions to ensure that the text conforms to the target format and semantic accuracy;

[0373] S4.4 Feedback and Interaction Operation: The feedback and interaction refers to the operator providing feedback on modification results and recording modification opinions through the interactive interface;

[0374] The purpose of the feedback and interaction is for operators to send their modifications and confirmations back to the system database so as to record and further optimize the system's modification strategy;

[0375] S4.41 Interactive Interface Feedback: Modified text, undoed modifications, or manually adjusted content are fed back to the system through the interactive interface;

[0376] S4.42 Record Operation History: Operator feedback will be synchronized to the system database, recording modification opinions, confirmation time, and operator identity to ensure process traceability;

[0377] S4.5 Verification Trajectory Recording and Auditing: The verification trajectory recording and auditing system records the complete verification history to ensure that the process is traceable and the results are auditable.

[0378] The purpose of recording and auditing the verification process is to ensure the transparency and auditability of the manual verification process, facilitating subsequent inspection and optimization.

[0379] S4.51 Complete Verification Track: The system records the complete verification process, including information such as operator identity, modified content, and confirmation time;

[0380] S4.52 Audit Function: All operation traces and modification records will be saved to ensure that auditing and backtracking are possible, facilitating quality assessment and process optimization.

[0381] The "manual verification and confirmation module" in the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention mainly yields the optimized final text, confirming or adjusting its accuracy, structural consistency, and semantic logic. All modifications, reversals, and manual adjustments are recorded, providing support for data traceability and quality auditing.

[0382] Example 6: Specific Implementation of the "Data Output Module" in the Human-Computer Interaction Language Model Output Post-Processing System and Optimization Method Supporting Field-Level Scoring Comparison and Rule Engine of the Present Invention.

[0383] The "data output module" in the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention aims to encapsulate the generated content after scoring comparison, automatic modification and manual verification in a preset format and output it to the subsequent industrial software system, so as to achieve efficient docking and data flow with the enterprise's existing information system.

[0384] The "data output module" in this invention, which supports field-level scoring comparison and rule engine-based human-computer interaction language model output post-processing system and optimization method, supports the encapsulation and conversion of various standardized data formats, such as JSON, XML, and CSV. It can also adapt its output according to the target system's interface protocols (such as RESTful API, OPCUA, and MQTT), and can be directly integrated into typical industrial software platforms such as Advanced Planning and Scheduling Systems (APS), Manufacturing Execution Systems (MES), and Enterprise Resource Planning Systems (ERP). The "data output module" mainly includes a data encapsulation unit, a format conversion unit, an interface protocol adaptation unit, a system integration and data flow unit, and a verification and validation unit.

[0385] The data encapsulation unit is responsible for encapsulating the generated content, which has been scored, compared, automatically modified, and manually verified, according to a preset format (such as JSON, XML, CSV).

[0386] The format conversion unit is responsible for converting and adapting data between different standardized formats (such as JSON, XML, and CSV) to ensure that the data format meets the requirements of the target system.

[0387] The interface protocol adaptation unit adapts and outputs data according to the target system's interface protocol (such as RESTfulAPI, OPCUA, MQTT, etc.) to ensure efficient integration with the enterprise's existing information systems.

[0388] The system integration and data transfer unit is responsible for outputting the packaged data and integrating it into the target industrial software system (such as APS, MES, ERP), realizing data transfer and ensuring seamless integration with subsequent systems.

[0389] The operation of the "data output module" in the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine described in this invention can be completed in the following specific steps:

[0390] S5.1 Data Processing and Scoring Comparison Operation

[0391] The key to the data processing and scoring comparison is to ensure that the data has been scored, compared and automatically modified before entering the output stage, and then manually verified to meet the high standards of industrial applications.

[0392] The main task of the data processing and scoring comparison is for the system to score and compare the generated content, automatically correct any parts that do not meet expectations, and for manual verification to ensure semantic and structural accuracy.

[0393] S5.11 Scoring Comparison: Automatically performs data scoring comparison, compares the generated content, and judges its matching degree with standard templates or preset rules;

[0394] S5.12 Automatic Modification: Based on the results of the scoring comparison, automatically correct any errors or non-standard parts of the data;

[0395] S5.13 Manual Verification: Manual review and correction of errors that could not be resolved by automatic modification to ensure the semantic accuracy and structural standardization of the data;

[0396] S5.2 data format encapsulation operation

[0397] The key point of the data format encapsulation is to encapsulate the processed data according to the preset format requirements, usually using standardized formats such as JSON, XML, CSV, etc.

[0398] The main task of the data format encapsulation is to encapsulate the data according to different standardized formats to ensure the standardization of the data structure, so that the data can be correctly parsed and used in subsequent systems. The choice of encapsulation format depends on the requirements of the target system and the usage scenario.

[0399] S5.21 Select Packaging Format: Select the appropriate data format based on the requirements of the target system, such as JSON, XML, CSV, etc.

[0400] S5.22 Data Structuring: Organize data into a standard structure according to the target format specification, such as JSON object key-value pairs, XML tags, etc.

[0401] S5.23 Encapsulate Data: Encapsulate structured data according to the selected format to generate the final data file or data stream;

[0402] S5.3 Interface Protocol Adaptation Operation: The key point of interface protocol adaptation is to adapt the output according to the interface protocol of the target system, such as RESTfulAPI, OPCUA, MQTT, etc., to ensure that data can flow between different systems without transmission failure due to protocol incompatibility.

[0403] The main task of the interface protocol adaptation is that the data output module needs to adapt the encapsulated data according to the interface protocol adopted by the target system, such as RESTfulAPI, OPCUA, MQTT, etc.

[0404] S5.31 Protocol Parsing: Analyze the interface protocols used by the target system, such as RESTfulAPI, OPCUA, MQTT, etc., to determine the data transmission methods and requirements;

[0405] S5.32 Protocol Conversion: Convert data according to the requirements of the target protocol, such as converting JSON data into the POST request body of a RESTful API, or converting the data format to the OPCUA standard;

[0406] S5.33 Interface Configuration: Configure the adapter or middleware to ensure that data can be transmitted according to the target protocol and process specific fields and header information in the protocol;

[0407] S5.4 Data Transmission and Integration Operation

[0408] The key point of the data transmission and integration is to transmit the data from the data output module and integrate it into subsequent industrial software systems, such as APS, MES, ERP, etc.

[0409] The main task of the data transmission and integration step is to transmit the encapsulated and protocol-adapted data to the target system and integrate it.

[0410] The target systems include Advanced Planning and Scheduling System (APS), Manufacturing Execution System (MES), and Enterprise Resource Planning System (ERP), and the integrated data can be directly provided to other systems or modules for further processing or analysis;

[0411] S5.41 Establish a connection: Establish a connection with the target system through a suitable network protocol, such as connecting to a RESTful API via HTTP protocol, or connecting to an industrial control system via OPCUA;

[0412] S5.42 Data Transmission: The encapsulated and adapted data is transmitted to the target system through the established connection, ensuring accurate and error-free data transmission;

[0413] S5.43 Integration into the target system: Data integration is performed in the target system (such as APS, MES, ERP) to ensure that the data can be correctly received and applied by the system;

[0414] S5.5 Final Output and Feedback Operation

[0415] The key point of the final output and feedback is that, as the final link in the human-computer interaction process, it ensures the accuracy and structural standardization of the data output and provides feedback on the output results.

[0416] The final output and feedback are primarily responsible for ensuring that the generated data meets industrial application requirements, is semantically accurate and structurally standardized. Furthermore, if any errors or non-standard outputs occur, feedback and necessary corrections will be provided to ensure the high quality and reliability of the output results.

[0417] S5.51 Output Data Confirmation: Perform a final confirmation of the final output data to ensure that the data meets the standards and industry requirements of the target system;

[0418] S5.52 Feedback Mechanism: If any errors or non-standard outputs occur during the output process, the feedback mechanism will notify the relevant personnel for correction.

[0419] S5.53 Error Correction and Re-output: Based on feedback, correct errors in the data output and re-output it to the target system to ensure that the system is not affected.

[0420] The results obtained by the "data output module" in the post-processing system and optimization method of the human-computer interaction language model output supporting field-level scoring comparison and rule engine of this invention are high-quality standardized data. After complete format encapsulation, interface adaptation, field-level verification, rule engine optimization and data integration, the data is ensured to meet the requirements of industrial applications in terms of semantics and structure, and can be seamlessly connected to the enterprise's existing information system to realize data flow and business process optimization.

[0421] Comparative Example 1: Measurement results of the use and non-use of the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention.

[0422] Comparative Example 1 highlights the differences before and after the use of "manual input + post-processing logic" in the post-processing system and optimization method of the human-computer interaction language model output of the present invention, which supports field-level scoring comparison and rule engine, in the industrial software processing flow. It specifically demonstrates the key role of post-processing logic in data quality control.

[0423] like Figure 2 As shown, the flowchart on the left illustrates the scenario without post-processing logic. Data flows directly from "text input" to the "manual entry" stage, where manual input is the primary method. This results in quality issues such as erroneous input, incorrect spelling, and duplicates. Furthermore, the input results are directly used for "importing into the system" without system validation, posing a high risk. The characteristics of not using post-processing logic are a large workload, low error tolerance, and a lack of intelligent validation, leading to erroneous information flowing into the business system.

[0424] like Figure 2 As shown in the flowchart on the right, the case where post-processing logic is introduced involves inserting an "automatic post-processing judgment" module between "manual input" and "import into the system." In this case, the post-processing logic performs operations such as structural standardization verification, spell correction, and redundant deletion. Only the judgment result is imported into the system, improving accuracy and standardization. The key feature of introducing post-processing logic is the introduction of automated judgment, which significantly reduces manual costs while improving data consistency and ensuring the stable operation of the software.

[0425]

[0426] Furthermore, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention, are presented in two comparative embodiments and detailed descriptions around the two scenarios of "using post-processing" and "not using post-processing". These embodiments particularly highlight the role of the post-processing system in industrial software for quality control of manually entered data and adaptability of language models.

[0427] Example 1: Smart Factory Information Input Process Using a Post-Processing System

[0428] In a factory situational awareness system, operators need to input information such as equipment status, maintenance logs, and anomaly descriptions daily for subsequent analysis and scheduling. The key processes and modules are as follows:

[0429] S1. Input Stage: Manual text input + LLM-generated auxiliary description

[0430] Operators enter the equipment number, operating status, and abnormal phenomena into the system; and complete the description by combining the RAG or LLM modules, such as GPT generating a natural language description "This equipment experienced abnormal shaking after running for 3 hours today".

[0431] S2. Post-processing: Field-level score comparison + rule engine

[0432] The system scores each field (such as device number, status, and fault description) for whether it is standardized, whether there is redundancy or spelling errors, and whether it is a duplicate input.

[0433] The application rules engine is manifested as follows:

[0434] First, the device number "A1234" was found to be repeated multiple times → marked as redundant;

[0435] Secondly, the misspelling of "runing" will be automatically corrected to "running".

[0436] Third, the "fault description" contains expressions that do not match the database rules → prompts for modification or suggests replacing the terms;

[0437] Fourth, output the final result and score it to track the optimization effect.

[0438] S3. Manual confirmation + business system entry

[0439] If the system determines that the output quality is below the set threshold (e.g., BLEU / ROUGE < 0.6), it will proceed to manual review; after review, the information will be formatted and imported into the MES or factory ERP system.

[0440] S4. Separation of the operation phase and the testing phase

[0441] During the testing phase, a scoring module is used to evaluate the LLM adaptation effect; during the working phase, add, delete, and modify logic is used in conjunction with a manual confirmation mechanism to ensure a closed loop of business quality.

[0442] The post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention, have the following effects: automatically identifying low-quality input and correcting spelling or structural problems; reducing manual review costs and improving the stability of industrial software; adapting to various large models and improving the consistency and reliability of RAG / LLM output; and supporting audit trails: each modification and confirmation has a timestamp and a responsible person identifier.

[0443] Example 2: Smart Factory Data Entry Process Without Using a Post-Processing System

[0444] While both systems involve entering factory equipment operation logs, this one lacks a post-processing module, relying solely on manual input and model output. The specific operation process is as follows:

[0445] S1. The operator manually enters text, such as "Equipment No.: A1234, Status: running, Phenomenon: Severe shaking";

[0446] S2. Generate auxiliary descriptions for large models;

[0447] S3. The system does not perform any field scoring, semantic comparison, or spell checking; it directly imports data into the business system.

[0448] As a result, potential problems include unrecognized spelling errors, unoptimized redundant descriptions, inconsistent business rules, and excessive manual workload.

[0449] The spelling error was not recognized because "runing" was misspelled as "ruing," which was mistakenly interpreted by the system as a new state, causing data clustering or statistics to fail.

[0450] The description of redundancy is not optimized, resulting in repeated recording of "Device A1234 is abnormal" without redundancy filtering, leading to database bloat.

[0451] The inconsistency in the business rules is due to the use of non-standard terms in the description, such as "severe vibration," which is not the "abnormal vibration" accepted by the system and cannot be automatically identified.

[0452] The excessive workload mentioned above stems from the fact that the subsequent data analysis department needs to manually clean up a large amount of redundant and non-standard data, which is extremely inefficient.

[0453] Furthermore, the two-stage application scenario of the post-processing logic of the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention is specifically manifested as follows:

[0454]

[0455] This demonstrates that the post-processing logic of the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine, as described in this invention, is not a "retraining" of the model, but rather provides a non-intrusive enhancement mechanism, establishing an "intelligent buffer layer" between model output and system use. The post-processing optimization mechanism can improve the structuring and consistency of input and output data, enhance the overall robustness and scalability of the system, reduce the burden of manual verification, and improve overall efficiency and user experience.

[0456] In summary, the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention effectively solves the problems of low quality, redundancy, and unstructured data entered manually in industrial software by introducing a two-stage post-processing logic that combines field-level scoring comparison and rule engine. This improves the data consistency and operating efficiency of the system, supports automatic verification and structural transformation of large language model output results, and provides key support for the intelligent transformation of the manufacturing industry.

[0457] Comparative Example 2: Empirical test results of the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine as described in this invention.

[0458] The present invention relates to a human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine, and its application of hard-coded post-processing logic in a large distillation model and its effectiveness verification in industrial scenarios.

[0459] To verify the effectiveness of this mechanism, a test set for industrial scenarios was constructed, and three mainstream distillation models (Deepseek 7B, Deepseek 14B, and Deepseek 32B) were systematically evaluated. Evaluation dimensions included binary confusion matrix analysis and text generation quality metrics such as BLEU and ROUGE. Experimental results show that introducing hard-coded post-processing logic significantly improves the model's output performance in terms of precision, recall, accuracy, ROUGE, and BLEU metrics (e.g., Figure 3 As shown), this demonstrates its enhanced capabilities in content generation accuracy and structural matching (e.g. Figure 4 (As shown). This confirms that the method described in this invention can significantly improve efficiency in large-scale distillation models.

[0460] Firstly, the application of hard-coded post-processing logic in a large distillation model and its evaluation methods and test set construction in verifying its effectiveness in industrial scenarios.

[0461] Firstly, the evaluation method and test set construction: Specifically, the test set construction was conducted to verify the effectiveness of the proposed post-processing mechanism. An industrial-scenario-oriented test set was used in the experiments. This test set mainly evaluated three mainstream large-scale distillation models: Deepseek 7B, Deepseek 14B, and Deepseek 32B, representing models of different scales. The purpose of test set construction was to ensure the effectiveness of the proposed post-processing mechanism in practical industrial applications, and an industrial-scenario-oriented test set was constructed.

[0462] Specifically, the evaluation dimensions include binary confusion matrix analysis and text generation quality metrics (such as BLEU and ROUGE); the binary confusion matrix analysis is used to evaluate the accuracy of the model output; the text generation quality metrics (such as BLEU and ROUGE) are used to evaluate the quality of the generated text. BLEU measures the similarity between the generated text and the reference text, and ROUGE measures the overlap between the generated text and the reference text.

[0463] The text generation quality metrics (BLEU and ROUGE) are used to measure the quality of the text generated by the model, ensuring that the generated content is not only semantically accurate but also conforms to formatting requirements. The two text generation quality metrics, BLEU and ROUGE, evaluate the quality of text generation from different perspectives.

[0464] The BLEU (Bilingual Evaluation Understudy) is mainly used to evaluate the similarity between generated text and reference text, and is particularly suitable for quality evaluation of machine translation tasks.

[0465] The t-BLEU refers to the overall BLEU index, which measures the similarity between the generated text and the reference text.

[0466] p-BLEU refers to the BLEU metric for key fields, measuring the similarity of key fields in text, such as product name, date, and quantity. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) primarily assesses the overlap between generated text and reference text, focusing particularly on recall, i.e., how much information from the reference text is included in the generated text.

[0467] The ROUGE algorithm calculates n-gram overlaps (such as 1-gram, 2-gram, etc.) and can be extended to LCS (Longest Common Subsequence), etc.

[0468] ROUGE-N is used to evaluate the overlap between the generated text and the reference text at the n-gram level.

[0469] ROUGE-L is used to evaluate the matching degree between the generated text and the reference text in the longest common subsequence.

[0470] The t-ROUGE mentioned above is a metric representing the overall ROUGE category, which measures the overlap of the entire text.

[0471] The p-ROUGE mentioned above is a ROUGE-type metric that represents key fields and measures the overlap of key fields in text.

[0472] Furthermore, experimental results show that the BLEU and ROUGE indices of the model are significantly improved after adopting hard-coded post-processing logic, indicating that the post-processing mechanism has a significant effect on improving the quality and structural consistency of the generated text.

[0473] Specifically, in the evaluation method and test set construction section of this invention, the experiment used a test set tailored to industrial scenarios and systematically evaluated three large-scale distillation models (Deepseek 7B, Deepseek 14B, and Deepseek 32B). The purpose of the evaluation was to verify the improvement in the quality of the generated text achieved by the post-processing mechanism.

[0474] The test set includes test data for three different-scale distillation models: Deepseek 7B, Deepseek 14B, and Deepseek 32B. These three models represent language models with different computational resources and scales, and are used to evaluate the effectiveness of post-processing mechanisms on models of different sizes.

[0475] Experiments show that introducing hard-coded post-processing logic significantly improves the model's precision, recall, accuracy, ROUGE, and BLEU scores. This indicates that the mechanism enhances the model's performance in terms of content generation accuracy and structural matching.

[0476] like Figure 3 As shown, this paper evaluates the impact of post-processing mechanisms on model performance by comparing the precision, recall, and accuracy of three large distillation models (Deepseek7B, Deepseek14B, and Deepseek32B). The comparison of precision, recall, and accuracy demonstrates that the post-processing logic improves these key metrics. The "hard-coded post-processing logic" significantly improves the model's output accuracy, specifically in the following aspects: precision, the proportion of correctly outputted data, reflects the model's accuracy in generating content; recall, the proportion of relevant information retrieved by the model, emphasizes the model's coverage; and accuracy, the degree to which the model's predictions match the true values. These metrics can measure whether the hard-coded post-processing mechanism has successfully improved the model's effectiveness in industrial applications, especially in tasks with format requirements such as order entry and production planning.

[0477] Specifically, such as Figure 3 As shown, each group of bars represents the performance of different models under different evaluation metrics; each group has four bars, corresponding to the four metrics: t-bleu, p-bleu, t-rouge, and p-rouge. Each bar represents a model's score under a specific metric, with colors distinguishing different models: green bars represent the score of the Deepseek7B model, orange bars represent the score of the Deepseek14B model, and blue bars represent the score of the Deepseek32B model.

[0478] Empirical results show that the high-parameter Deepseek32B model scores highly on all evaluation metrics, including overall BLEU, key field BLEU, overall ROUGE, and key field ROUGE. This indicates that the Deepseek32B model is not only fluent and accurate in generating text, but also ensures that the key content and information of the generated text are more closely matched with the reference text. In contrast, the low-parameter LLM (Deepseek7B, Deepseek14B) models have lower fluency, accuracy, and similarity to the reference text when generating text, and may lose some key information.

[0479] like Figure 4 The diagram shows a comparison of the ROUGE and BLEU metrics, further demonstrating the effectiveness of the hard-coded post-processing mechanism in improving text generation quality. ROUGE measures the similarity between the generated text and the reference text, and is often used to evaluate the recall capability of the generated text. BLEU measures the accuracy between the generated text and the reference text, focusing particularly on the matching of key fields and words in the generated text. These metrics reflect how the post-processing logic improves the quality of the model during the generation process, especially how it enhances semantic accuracy and structural regularity when generating text. Experiments using these metrics demonstrate that introducing hard-coded post-processing logic can effectively improve the overall quality of the generated text.

[0480] Figure 4 This paper presents a comparison of the performance changes of three large models of different sizes (Deepseek7B, Deepseek14B, and Deepseek32B) on multiple evaluation metrics before and after post-processing (-h). Figure 4 The upper part (blue-green) represents the performance of the original model (Deepseek7B / Deepseek14B / Deepseek32B); Figure 4 The lower half (yellow) represents the performance of the model (7B-h / 14B-h / 32B-h) after adding the post-processing mechanism and the relative percentage improvement; Figure 4 Each model is evaluated using the following seven metrics: “t-bleu” represents the text-level BLEU score; “p-bleu” represents the paragraph-level BLEU score; “t-rouge1 / t-rouge2 / t-rougeL” represents the text-level ROUGE-1 / 2 / L; and “p-rouge1 / p-rougeL” represents the paragraph-level ROUGE-1 / L. These metrics comprehensively evaluate the accuracy and coverage of the text generated by the model.

[0481] Overall, Figure 4The performance of large distillation models of different sizes (Deepseek7B, Deepseek14B, Deepseek32B) under various metrics (BLEU and ROUGE) is shown, as well as the performance improvement after applying the post-processing mechanism. Figure 4 On the left are models of different sizes: Deepseek7B, Deepseek14B, and Deepseek32B, representing large distillation models with varying parameter sets. Each model has two versions: a standard version (Deepseek7B, Deepseek14B, Deepseek32B) and a version with post-processing (7B-h, 14B-h, 32B-h). The post-processing mechanism (h) improves the quality of the model output through certain rules or techniques. Here, t-bleu is the overall BLEU score, while p-bleu focuses on the BLEU scores of key fields; higher scores indicate a higher degree of linguistic structural matching between the generated and reference texts, especially in key information sections. The OUGE metric measures the recall between the generated and reference texts; t-rouge represents the overall recall, while p-rouge focuses on the recall of key fields.

[0482] As shown in the figure, Deepseek7B shows significant improvement for small models, indicating that structured output constraints are more beneficial in small models, especially with a noticeable increase in BLEU scores. Deepseek14B medium-sized models benefit the most, with the post-processing mechanism showing the most significant effect on their structure optimization, resulting in improvements exceeding 8% in almost all metrics. Deepseek32B large models already have very high raw performance with limited room for further improvement, but the post-processing mechanism still provides stable improvements, particularly in BLEU scores, resulting in more standardized output.

[0483] Empirical evidence shows that, among the Deepseek7B, Deepseek14B, and Deepseek32B models, the Deepseek7B model scores relatively low, indicating that smaller-scale models perform worse than larger models when handling complex language generation tasks, especially in key field processing. The Deepseek14B and Deepseek32B models show significant improvement in overall performance with increasing scale, particularly in t-bleu and p-bleu (BLEU scores for the whole and key fields) and t-rouge and p-rouge (ROUGE scores for the whole and key fields), demonstrating better text generation capabilities and higher matching accuracy.

[0484] Empirical evidence shows that, regarding the impact of post-processing mechanisms on Deepseek7B, Deepseek14B, and Deepseek32B models (h version), the models (7B-h, 14B-h, and 32B-h) exhibit significant performance improvements after applying post-processing mechanisms. Specifically, post-processing mechanisms improve the model scores on various evaluation metrics, particularly in key field matching and recall. Specifically, Deepseek7B-h shows improvements of 5.32% to 8.85% across all metrics compared to Deepseek7B; Deepseek14B-h shows even greater improvements compared to Deepseek14B, especially exceeding 12% in p-bleu and p-rouge; while Deepseek32B-h shows the largest improvement, particularly a 6.76% increase in t-rouge and p-rouge, indicating that the role of post-processing mechanisms is more prominent in larger models.

[0485] The conclusions are as follows: Regarding the applicability of the post-processing mechanism, the human-computer interaction language model output post-processing system and optimization method described in this invention, which supports field-level scoring comparison and rule engine, improves not only the overall generation quality of the model but also enhances the controllability of its structure and its ability to process key fields. This mechanism enables the model to better meet format requirements and key content generation needs in industrial application scenarios such as order entry, production planning, and equipment status report generation. The flexibility and scalability of the post-processing mechanism allow the system to adapt to the application needs of different fields; thus, through rule base updates, the structure and quality of generated content can be optimized according to the characteristics of different industries or business scenarios.

[0486] Therefore, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention, provides a feasible technical path for the implementation of large-scale language models in industrial environments based on hard-coded rule-guided post-processing logic. It not only enhances the structural consistency of model output but also provides strong support for subsequent data processing and automated process integration, demonstrating good engineering application prospects and patent protection value.

[0487] Application Example 1: A specific implementation method of the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention, to achieve standardized output in industrial scenarios.

[0488] Application Example 1 focuses on how to achieve standardized output in industrial scenarios using the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine, as described in this invention. While standalone LLMs are powerful in text generation, their output contains data that does not conform to standard formats and cannot be directly used in information systems. Therefore, post-processing of LLM output is necessary to ensure that the data meets preset format requirements and to prevent invalid or non-standard data from being imported into the system. This is the practical motivation behind the improvement of the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine, as described in this invention.

[0489] First, the production status of our company's factories is reported via voice feedback as follows:

[0490] In factories, operators use voice input to provide feedback on production conditions, such as machine malfunctions and downtime. However, operator voice input does not directly conform to a preset standard format. LLMs can understand speech and generate relevant content, but this only generates a description in natural language. The post-processing logic unit then formats the generated text, converting it into a standardized structure (such as JSON) so that it can be correctly recognized and processed by information systems (such as MES / APS).

[0491] Large Language Models (LLMs) alone cannot achieve standardized output; however, standardization can be achieved by optimizing the output through post-processing logic units.

[0492] First specific implementation method: Standardized input of production data

[0493] In manufacturing plants, operators provide feedback on production equipment malfunctions via voice input. Operators verbally describe the production situation using smart headsets or microphones, and the system uses speech recognition technology to convert this into text. The text is processed by a large language model, and the generated output contains key information such as machine status, fault description, and downtime. At this point, the post-processing module parses the generated text and converts it into a standardized JSON format for use by MES (Manufacturing Execution System) or APS (Advanced Planning and Scheduling System).

[0494] For example, if operator A uploads a voice message saying "Machine group 1 on line A is malfunctioning and is estimated to need to be shut down for 3 hours...", the large language models (LLMs) optimized by the post-processing logic can convert it into a preset standard format and feed it back to MES / APS directly in formats such as JSON.

[0495] Therefore, the reference output is:

[0496] {

[0497] "line":"A",

[0498] "group":1,

[0499] "status":false,

[0500] "downtime":180,

[0501] "remark":"-"

[0502] }

[0503] Therefore, the output of large-scale language models (LLMs) is:

[0504] 1. 90% of the outputs will contain the following unbiased content:

[0505] {

[0506] "line":"A",

[0507] "group":1,

[0508] "status":false,

[0509] "downtime":180,

[0510] "remark":"-"

[0511] }

[0512] As can be seen from the above input, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention can standardize the output of LLMs into a data format that the system can recognize, while filtering out invalid or non-standard outputs (e.g., time format does not meet expectations).

[0513] Second specific implementation method: Invalid information filtering

[0514] When the factory system receives information uploaded by operators or other systems, it may encounter some invalid, noisy data. The post-processing logic analyzes the input to determine if it meets preset standards, such as whether it is empty or contains meaningless characters. Once invalid data is confirmed, such as the input "ABC", the system will filter it directly and not import it into any information system, avoiding unnecessary errors.

[0515] For example: 10% of the output results are biased / in standard format, causing them to be unrecognizable by the system.

[0516] {

[0517] "line":"A",

[0518] "group":"1",

[0519] "status":false,

[0520] "downtime":"180",

[0521] "remark":"-"

[0522] "time":"20250508103411"

[0523] }

[0524] Therefore, using this invention will only output the following fields, filtering out invalid outputs. That is, after post-processing logic optimization, the output is standardized as follows:

[0525] {

[0526] "line":"A",

[0527] "group":1,

[0528] "status":false,

[0529] "downtime":180,

[0530] "remark":"-"

[0531] }

[0532] As can be seen from the above input, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention ensures that the information system only receives valid standardized data and avoids invalid data interfering with system operation; and improves the data quality of the information system by filtering noise information through post-processing logic.

[0533] For example, when importing invalid noise information into an information system, it is desirable for the system not to accept the input.

[0534] Therefore, the input is: ABC

[0535] Reference output: {}

[0536] Therefore, the output of large-scale language models (LLMs) is:

[0537] The system does not return an empty value after accepting input, but instead outputs the following content, causing it to be unrecognizable by the system:

[0538] Sorry, please enter your details.

[0539] The system identifies this type of invalid information through a post-processing mechanism, filters it directly, and outputs an empty string: {}

[0540] As can be seen from the above input, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention ensures that the information system only receives valid standardized data and avoids invalid data interfering with system operation; and improves the data quality of the information system by filtering noise information through post-processing logic.

[0541] Second, the order information is transferred to the contract manufacturer.

[0542] When a factory is unable to handle large orders, it outsources the excess orders to a contract manufacturer. The factory sends order information to the contract manufacturer via email or other communication methods. This information includes product name, quantity, delivery date, etc. To ensure that the contract manufacturer can process orders in a timely manner, the system uses LLMs to parse this text information and format it into a standardized JSON structure, ensuring that the contract manufacturer's information system can correctly receive and process orders.

[0543] For example, when importing invalid noise information into an information system, it is desirable for the system not to accept the input.

[0544] Therefore, the input is: ABC

[0545] Reference output: {}

[0546] Therefore, the output of large-scale language models (LLMs) is:

[0547] The system does not return an empty value after accepting input, but instead outputs the following content, causing it to be unrecognizable by the system:

[0548] "Sorry, please enter your details."

[0549] The system identifies this type of invalid information through a post-processing mechanism, filters it directly, and outputs an empty string: {}

[0550] As can be seen from the above input, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention ensures that the information system only receives valid standardized data and avoids invalid data interfering with system operation; and improves the data quality of the information system by filtering noise information through post-processing logic.

[0551] Third, scenario 2: This company's subordinate factory systems import standardized information into the information system.

[0552] When a factory faces a large volume of orders, it needs to transfer orders exceeding its production capacity to contract manufacturers. The order information sent by the upstream factory includes product specifications, delivery quantity, delivery deadline, and other information. This information needs to be imported into the information system in strict accordance with a standardized format.

[0553] The fourth specific implementation method: Industrial software input data control

[0554] Factory managers upload equipment maintenance reports, which are in PDF format and contain detailed equipment status information and maintenance dates. The system uses a post-processing mechanism to parse the content of the PDF, extract key information, and generate standardized output according to a preset format. The system ensures that all output conforms to the input specifications of industrial software, preventing non-compliant data from affecting subsequent processing.

[0555] For example, when a manufacturing plant faces large or urgent orders, it may subcontract the excess capacity to a contract manufacturer. In this case, the contract manufacturer must organize production according to the client's technical requirements, considering factors such as product specifications, delivery quantity, and delivery deadline. The parent factory then sends a PDF document describing the product specifications, delivery quantity, and delivery deadline.

[0556] Using this invention will strictly control the output format, ensuring it meets the standardized input requirements of industrial software.

[0557] {

[0558] "product":"A",

[0559] "number":1000,

[0560] "status":001,

[0561] "deliveryDate":"2025-05-01",

[0562] "destination":"Kunming,CN"

[0563] "remark":"-"

[0564] }

[0565] As can be seen from the above input, the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine described in this invention can ensure that the generated output format meets the standard requirements of industrial software systems, improve the efficiency of cross-factory collaboration, and enable information to be seamlessly integrated into various industrial software.

[0566] Therefore, the human-computer interaction language model output post-processing system and optimization method supporting field-level scoring comparison and rule engine described in this invention, through the optimization of this post-processing logic unit, enables large language models (LLMs) to process natural language input and output standardized data, ensuring that the data meets system requirements and preventing invalid or non-standard data from being imported into industrial software. This mechanism has broad application prospects in industrial applications, especially in scenarios requiring standardized data input and output.

[0567] Application Example 2: Specific Implementation of the Post-Processing System and Optimization Method for Human-Computer Interaction Language Model Output Supporting Field-Level Scoring Comparison and Rule Engine as described in this invention in Agent Construction.

[0568] Application Example 2 highlights the application of the post-processing system and optimization method for human-computer interaction language model output supporting field-level scoring comparison and rule engine, as described in this invention, in the construction of intelligent agents, particularly for built-in encapsulated clauses. In the system and implementation architecture, the post-processing system and optimization method for human-computer interaction language model output supporting field-level scoring comparison and rule engine, as described in this invention, are applied to the construction of intelligent agents, encapsulating the functional modules of the intelligent agent as independent components with interface definitions. Called as a basic tool for the intelligent agent, the post-processing optimization technology described in this invention can serve as a necessary component for building intelligent agents. For example, in creating a factory situational awareness intelligent agent, the post-processing method mentioned in this invention can be directly used to complete the task of inputting standardized information.

[0569] The present invention discloses a post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, which is a post-processing optimization method for intelligent agent systems. Its core innovation is that it supports field-level scoring comparison, supports automatic judgment by rule engine, and is geared towards the output optimization of human-computer interaction language models. It is especially suitable for tasks such as structuring, verifying, and standardizing the input of model output within intelligent agents.

[0570] First, regarding the construction and functional encapsulation of the agent, an agent is composed of multiple independent functional modules, each capable of independently completing a specific task. These modules interact through interface definitions, ensuring the system can flexibly adapt to various needs during operation. In this embodiment, the factory situational awareness agent is constructed as a system containing multiple functional modules. Each module is responsible for different tasks, such as data acquisition, data analysis, model prediction, and result display.

[0571] Second, the application of the post-processing logic module. The post-processing logic module described in this invention is mainly used for optimizing manually input data. Specifically, this module processes the data output by the model to eliminate spelling errors, duplicate inputs, and other problems, and to ensure data quality.

[0572] Specifically, the post-processing logic module of the present invention includes field-level scoring comparison and rule engine stage;

[0573] The field-level scoring and comparison refers to the post-processing logic system scoring and comparing each field output by the model to ensure that the field content conforms to predetermined specifications. For example, when inputting the status of factory equipment, the system can automatically check whether the equipment ID is valid and whether the equipment status meets preset standards (such as "running" or "stopped").

[0574] The rule engine processes data based on preset rules, automatically identifying and correcting erroneous or rule-inconsistent data. For example, if a non-existent device ID is entered during data entry, the rule engine will correct it or prompt the user based on standard data in the database.

[0575] Third, the application of "post-processing logic" in the testing and operational phases of the post-processing system.

[0576] The post-processing system and optimization method of the human-computer interaction language model output supporting field-level scoring comparison and rule engine described in this invention are divided into two stages: the testing stage and the working stage.

[0577] The post-processing logic in the testing phase is used to score the model's output during the testing phase. At this time, the system evaluates the quality of preprocessing, model adaptation, and the generated output, and adjusts the processing methods based on the score. Through multiple iterations, the system can ensure the quality of the model's output.

[0578] The post-processing logic for the working stage involves adding, deleting, and modifying data in the model's output during actual system runtime. The focus of this stage is correcting erroneous data and ensuring all records match reality. Manual verification also plays a crucial role in the post-processing logic stage, allowing users to review the automated processing results and make necessary modifications.

[0579] Intelligent agent application example: factory situational awareness

[0580] Its main application scenario is to process tasks such as structuring, verifying, and standardizing the input of model output within the intelligent agent.

[0581] In a factory under this company, the status and operational data of production equipment need to be collected and analyzed. Data entry, post-processing optimization, and output result confirmation are completed through the post-processing logic of an intelligent agent. Data entry involves operators inputting equipment data into the system, such as equipment status and running time. The system automatically performs quality checks on each input field to ensure that each field conforms to preset rules, such as valid equipment IDs and equipment status within predetermined ranges. Post-processing optimization involves the system detecting invalid equipment IDs or non-compliant status inputs, such as incorrect equipment IDs or non-standard status descriptions, correcting them through a rule engine, and generating a scoring report. Output result confirmation occurs during the operational phase when the system automatically corrects the data and transmits it to the factory's situational awareness system, helping monitoring personnel to understand the equipment's operational status in a timely manner and make adjustments or troubleshoot when necessary.

[0582] The goal of the factory situation awareness is to construct a factory situation-aware agent, whose responsibilities include: collecting information reported by various personnel or sensors, standardizing and processing it before inputting it into the MES system and connecting it with alarm, scheduling and other systems.

[0583] Typical task flow: Intelligent entry of abnormal operating condition reports

[0584] Thus, the original input (workshop operator dialogue or model natural language output) is: "Around 10 o'clock today, the temperature of heat treatment furnace No. 5 rose too quickly, reaching nearly 700°C, and some white smoke was also emitted."

[0585] Therefore, the goal of post-processing is to "transform unstructured natural language information into structured fields + verification and comparison into standard input format".

[0586] Therefore, the following table content appears in the field-level scoring and rule engine processing:

[0587]

[0588] Therefore, the final standardized output (structured record) is as follows:

[0589] {

[0590] Equipment: "Heat Treatment Furnace No. 3"

[0591] Time: "2024-06-09 22:00",

[0592] "Abnormality Type": "Temperature rising too rapidly"

[0593] "Outlier": "800℃",

[0594] Accompanying phenomenon: "Appearance of white smoke"

[0595] "Event Level": "Level 1"

[0596] "Entry Person":"Operator A",

[0597] "Rating Summary":{

[0598] Average field confidence level: 0.93

[0599] Manual review suggestion: "No need".

[0600] }

[0601] }

[0602] Therefore, the system behaviors (post-processing of the work phase) are as follows: first, automatically generating structured records; second, automatically scoring credibility (field level + overall); third, if there is a low credibility or poor matching, the post-processing mechanism can mark it as requiring manual review; otherwise, the post-processing mechanism directly writes to the "event logging system" and triggers the alarm subsystem.

[0603] Furthermore, the advantages of using the human-computer interaction language model output post-processing system and optimization method that supports field-level scoring comparison and rule engine as described in this invention are reflected in the following aspects:

[0604]

[0605] This demonstrates that Application Example 2 clearly shows the application value of the technology of the present invention in the construction of actual industrial agents, which can improve the level of automation and standardization of data processing, enhance the practicality of human-computer interaction model output, and facilitate modular encapsulation for multi-scenario portability.

[0606] In conclusion, post-processing logic plays a crucial role in such intelligent agent applications, effectively improving data accuracy and quality and optimizing the overall system's operational efficiency.

[0607] The following four pseudocode snippets further illustrate the post-processing system and optimization method for human-computer interaction language model output that supports field-level scoring comparison and rule engine, as described in this invention.

[0608] (1) First pseudocode segment:

[0609]

[0610] Add avg_bleu to all_bleu_scores

[0611] Calculate the average BLEU score of all objects: total_avg

[0612] return total_avg

[0613] #2, Overall Unbiased Score

[0614] The function OverallBLEU(references_json, candidates_json):

[0615] Define the function json_to_tokens(json_str):

[0616] Parsing a JSON string into a list of objects

[0617] Convert each object into a list of tokens in key=value format.

[0618] reference_tokens=json_to_tokens(references_json)

[0619] candidate_tokens=json_to_tokens(candidates_json)

[0620] Calculate the overall BLEU score using sentence_bleu.

[0621] return overall_score

[0622] #3. Field-Level ROUGE as a Supplementary Indicator

[0623] Function MultiFieldLevelROUGE(ref_list,cand_list):

[0624] Initialize the result dictionary result = {'rouge1':[], 'rougeL':[]}

[0625] Iterate through each pair of references and candidate objects (ref_item, cand_item) in ref_list and cand_list:

[0626] For each field key in ref_item:

[0627] Retrieve the reference value r_val = ref_item[key]

[0628] Get candidate value c_val = cand_item.get(key,"")

[0629] If r_val == c_val: assign it the full score; otherwise: use RougeScorer to calculate the F1 scores of ROUGE-1 and ROUGE-L, add rogue1_f1 to result['rouge1'], and add rogueL_f1 to result['rougeL'].

[0630] Calculate the average F1 score (avg_scores) for ROUGE-1 and ROUGE-L.

[0631] return avg_scores

[0632] #4. Overall ROUGE calculation, as a supplementary indicator.

[0633] Function MultiItemROUGE(references_json,candidates_json):

[0634] Parse references_json and candidates_json into a list of objects.

[0635] Define the function merge_all_fields(items):

[0636] Merge all fields into strings in key=value format.

[0637] ref_text=merge_all_fields(references)

[0638] cand_text=merge_all_fields(candidates)

[0639] Use RougeScorer to calculate the ROUGE scores for ref_text and cand_text.

[0640] return scores

[0641] Output summary results:

[0642] Multi-object field-level average BLEU, overall BLEU, field-level average ROUGE-1, field-level average ROUGE-L, overall ROUGE-1 / 2 / L

[0643]

[0644]

[0645] (3) The third pseudocode segment:

[0646]

[0647] (4) Fourth pseudocode segment:

[0648]

[0649]

Claims

1. A method for optimizing the output of a human-computer interaction language model that supports field-level scoring comparison and rule engine, characterized in that, The post-processing optimization method includes: S1. Recording the operation of the information module: By recording key data in the generation process and combining it with structured management, the text generation process can be traced, and the generated text conforms to the preset structure and classification standards; S2. Operation of the scoring and comparison module: By using indicator units to perform similarity assessment and field-level analysis, the matching degree and coverage of the generated text with the reference content are quantified, and then the text quality is scored; S3. Modify module operation: Optimize text quality based on scoring results, automatically correct text that does not meet standards, and output standardized format data; S4. Operation of the manual verification and confirmation module: When the text quality does not meet the standard, the manual verification mechanism is triggered to confirm or modify the text through manual intervention, and the system operation process is recorded; S5. Operation of the data output module: Evaluate and format the generated text and its processing, ensure that the data conforms to the system interface standard and is reliably transmitted to the target system, and finally output the results and provide feedback.

2. The method for post-processing optimization of human-computer interaction language model output supporting field-level scoring comparison and rule engine as described in claim 1, characterized in that: The operation of the S1. information recording module includes: The S1.1 Result Recording Unit operates by fully recording the model output, field templates, task identifiers, and intermediate content, enabling full-process traceability of the generated data. S1.2 Operation of the structured template management unit: Classifies and identifies text according to preset structured templates, so that the generated text conforms to specific structural specifications; S1.3 Operation of the Task Type Identifier Classification Unit: Refined classification management of content, labeling explicit and implicit information; The operation of the S2. scoring comparison module includes: The S2.1BLEU indicator unit operates by calculating similarity through n-gram matching to evaluate the consistency of the generated text in terms of vocabulary and word order. The S2.2ROUGE indicator unit operates by calculating similarity through n-gram matching to measure the coverage of reference content and identifying the degree of similarity in the generated text through field-level analysis. The operation of the S3 modification module includes: S3.1 The operation of the reference comparison and modification unit: Based on the scoring results, the deviation is identified and the consistency and expression quality of the generated text are automatically optimized; S3.2 Operation of the No-Reference Comparison Modification Unit: In the absence of reference samples, the text content is independently verified and corrected based on the template and rules; S3.3 Automatic modification mechanism: When the text does not meet the quality standards, the system automatically starts the modification process for correction and optimization; S3.4 Output of Standardized Content: Output the modified text in a standard structured format and push it to the downstream system; The operation of the S4. manual verification and confirmation module includes: S4.1 Triggering the manual verification process: When the scoring result is below the threshold or automatic modification fails, the system automatically triggers the manual verification mechanism; S4.2 Automatic Data Extraction: The system automatically extracts key information related to the task from the input content or data source according to preset rules; S4.3 Manual Intervention and Confirmation: Confirm, cancel, or manually modify the text according to task requirements to ensure output quality; S4.4 Feedback and Interaction Operation: The system provides feedback on modification results through the interactive interface, and synchronously records operation information to drive the optimization of subsequent processes; S4.5 Verification Trajectory Recording and Auditing: The system fully records the entire verification trajectory process, enabling auditable operations and verifiable and traceable results; The operation of the S5 data output module includes: S5.1 Data Processing and Scoring Comparison: Preprocessing and quantitative evaluation of generated text, modifications, and verification; S5.2 Data Format Encapsulation Operation: The generated content is encapsulated into a standardized format according to preset standards; S5.3 Interface Protocol Adaptation Operation: Adapt the encapsulated data according to the interface protocol of the target system; S5.4 Data Transmission and Integration Operation: Reliably transmit and integrate the encapsulated and adapted data into industrial systems; S5.5 Final Output and Feedback Operation: As the end of the human-computer interaction closed loop, it outputs the final result and provides feedback after the task is completed.

3. The method for post-processing optimization of human-computer interaction language model output supporting field-level scoring comparison and rule engine as described in claim 1 or 2, characterized in that: The operation of the S1.1 result recording unit includes: S1.11 Capture Model Generation Results: Completely record the final text content and related field information generated from the model, ensuring that all output content has a traceable data foundation. S1.12 Capture Intermediate Content Generated During Post-Processing: Records and captures intermediate text and data generated during post-processing. The operation of the S1.2 structured template management unit includes: S1.21 Structured Classification: Based on preset templates, the generated text and fields are classified and labeled to adapt to dynamic management of various task types, and then recorded and stored according to structural requirements. S1.22 Structured storage: Storing recorded content according to a unified format standard; The operation of the task type identifier classification unit in S1.3 includes: S1.31 Explicit Information Labeling: A hierarchical labeling system is used to label the basic information and structural levels entered by the user. Explicit fields in the generated content are clearly identified and managed. S1.32 Labeling Implicit Information: Extract and label implicit information in the original input to support task understanding and model reasoning; S1.33 Annotation level information: Used to annotate the structural hierarchy and attribute information of the generated content, enhancing the system's ability to understand the multidimensional logical structure of the text; The operation of the S2.1BLEU indicator unit includes: S2.11 Input Data Processing S2.111 Model Output and Reference Sample Preparation: Collect the model-generated text and corresponding reference samples, and compare and analyze the similarity between the generated content and the standard answer. S2.112 Field-level Extraction: Extract key field information from the generated text and the reference text, and perform refined field-level difference analysis. S2.12 Calculation of word overlap rate S2.121 n-gram matching: By performing n-gram segmentation matching between the model output and the reference text, the overlap at the word order level is evaluated. S2.122 n-gram precision Pn calculation: Calculate the matching ratio of each level of n-gram in the generated text and the reference text to measure its word-level consistency. The accuracy formula is as follows: P_(n=) (Number of matched n-grams) / (Total number of n-grams in the model output) S2.131 Calculate the BLEU score: Based on n-gram precision Pn and weight ωn, and combined with the lengths of the reference and generated text, the BLEU score is calculated to measure overall text similarity. The specific formula is as follows: , Where r represents the length of the reference test case; c represents the length of the model output; Pn represents the accuracy of the n-gram; and ωn represents the weight of the accuracy of the n-gram. S2.132 Calculate overall and field-level BLEU scores: Calculate the overall BLEU value of the text and perform field-level scoring as needed to assess the generation quality of specific structured fields; S2.14 Results Output and Visualization S2.141 Results Presentation: The overall BLEU score and field-level scores are visualized in the form of charts or tables, showing the matching status of different n-grams; The operation of the S2.2ROUGE indicator unit includes: S2.21 Input Data Processing S2.211 Model Output and Reference Sample Preparation: Perform unified preprocessing on the model-generated text Y and the reference text X. The calculation is as follows: ROUGE-1: Based on 1-gram ROUGE-2: Based on 2-gram S2.212 Field-level data extraction: Field-level analysis is performed on some data to prepare for subsequent field-level ROUGE calculations; S2.22n-gram extraction and matching S2.221 n-gram generation: Extract 1-grams and 2-grams from the reference text and the generated text for ROUGE-1 and ROUGE-2 matching analysis. The ROUGE-N formula is as follows: ROUGE-N = Countmatch(〖gram 〗_N ) / Count(〖gram 〗_N ) Recall, or recall, is the percentage of n-grams in the reference text that appear in the model's output. ROUGE-N = (Number of matched n-grams) / (Total number of n-grams in the reference text) S2.222 Matching Statistics: Statistically count the total number of n-grams in the reference text and the number of n-grams that match in the generated text, providing basic data for similarity calculation; Extraction of the Longest Common Subsequence (LCS) in S2.23 S2.231 LCS Calculation: Calculate the longest common subsequence between the generated text Y and the reference text X, which is used to measure the degree of retention of reference information; S2.232 Recall Calculation: ROUGE-L recall is calculated based on LCS length to evaluate the proportion of reference text successfully covered in the generated text. The formula for the ROUGE-L index is: ROUGE-L = (LCS(X,Y)) / (max(|X|,|Y|)) Where LCS stands for Longest Common Subsequence, X represents the use case reference result, and Y represents the model output; S2.24 Field-level ROUGE Calculation S2.241 Independent ROUGE Calculation for Each Field: Calculate ROUGE-1, ROUGE-2, and ROUGE-L for each structured field to evaluate the output quality of the model under different semantic dimensions; S2.25 Results Summary and Output S2.251 Generate global and field-level reports: Output overall and field-level ROUGE scoring results and visualize them in the form of tables or charts; The operation of the reference comparison and modification unit in S3.1 includes: S3.1 Operation of the modified unit with reference comparison: When a reference sample exists, the generated text is automatically optimized based on the scoring results; S3.11 Receiving the scoring comparison results: Receive the results from scoring modules such as BLEU or ROUGE, and identify the deviation between the generated text and the reference sample; S3.12 Identify Deviations: Automatically locate and identify segments in the generated text that differ from the reference sample; S3.13 Perform automatic correction: perform structural corrections on the text; S3.14 Output modified text: Output the optimized text, ensuring that the difference from the reference sample is minimized; The operation of the S3.2 no-reference comparison modification unit includes: S3.21 Receive generated text: Receive uncompared model-generated text as input for subsequent verification and correction; S3.22 Structured Template and Syntax Rule Validation: Performs compliance checks on text structure and format based on preset templates and rules; S3.23 Field Integrity Check: Identify and complete any missing necessary fields in the generated text; S3.24 Illegal Character Filtering: Detects and removes certain characters from the text; S3.25 Semantic Logic Verification: Verify the logical and semantic rationality of the text, and correct the expression after discovering problems; S3.26 Output the corrected text: Output the standard text corrected by the template rules; The automatic triggering of the modification mechanism in S3.3 includes: S3.31 Determine if the generated text meets the standard: Determine if the text meets the standard through field integrity and semantic logic checks; S3.32 Automatic Editing: Automatically completes fields, corrects expressions, and cleans up content for text that does not meet the standards; S3.33 Output modified text: Outputs the automatically corrected standard text; The execution of the standardized output content in S3.4 includes: S3.41 Standardized output format: Converts the corrected content into a unified structured data format; S3.42 Push to downstream modules: Send the corrected text results to the business system or database to achieve data integration and final delivery; S4.1 triggers the manual verification process, including: S4.11 Scoring Threshold Judgment: The system determines whether the BLEU or ROUGE score is lower than the preset threshold in order to decide whether to initiate manual verification; S4.12 Modification module not meeting standards: When the automatic modification results do not meet the structured requirements, the system automatically enters the manual verification process; The automatic extraction of relevant data in S4.2 includes: S4.21 Original generated text: Extract the unmodified original generated text as a basis for manual judgment; S4.22 Modification suggestion: Display the modified version generated by the rule engine for comparison and evaluation reference; S4.23 Reference Sample: If a standard reference text exists, the system will extract and display it synchronously to assist manual comparison; S4.24 Structured Comparison View: The original text, suggested modifications, and reference samples are displayed side by side in a comparative format, allowing for rapid human identification of differences and intervention in decision-making; The operation of S4.3 manual intervention and confirmation includes: S4.31 Confirm Modification: Confirm the automatic modification suggestions and output them directly as the final text. S4.32 Undo Modification: If automatic modifications do not meet expectations, the modifications can be undone, retaining the original generated text; S4.33 Manual Adjustment: Manually edit the original text or suggested modifications to ensure they conform to formatting and semantic requirements; The operation of S4.4 feedback and interaction includes: S4.41 Interactive Interface Feedback: The text after confirmation, cancellation, or manual adjustment is fed back to the system through the interface; S4.42 Record Operation History: The system records feedback content, operation time, and personnel information, making the entire process traceable; The operation of the S4.5 verification trajectory recording and auditing includes: S4.51 Complete Verification Track: The system records verification information in detail; S4.52 Audit Function: Saves all operation records and supports auditing and process quality assessment; The operation of S5.1 data processing and scoring comparison includes: S5.1 Data Processing and Scoring Comparison Operation: Before output, the content quality is ensured to meet the standards through scoring, automatic modification and manual verification; S5.11 Scoring Comparison: The system automatically scores the generated content to measure its degree of matching with standard templates or rules; S5.12 Automatic Correction: Automatically corrects some content based on the scoring results; S5.13 Manual review: Automatic correction of semantic or structural issues not covered by the system is done manually. The operation of the S5.2 data format encapsulation includes: S5.21 Selecting the Packaging Format: Select a suitable packaging format based on the requirements of the target system; S5.22 Data Structuring: Organize data according to format specifications and construct a standard structure; S5.23 Encapsulate Data: Encapsulate structured data according to the selected format to generate the final data file or data stream; The operation of the S5.3 interface protocol adaptation includes: S5.31 Protocol Parsing: Identify and analyze the interface protocol type used by the target system; S5.32 Protocol Conversion: Converts encapsulated data into the format and structure required by the target protocol; S5.33 Interface Configuration: Configure the protocol adapter or middleware to ensure that data is transmitted correctly according to the protocol specifications; The operation of S5.4 data transmission and integration includes: S5.41 Establish Connection: Establish a stable communication channel with the target system through network protocols; S5.42 Data Transmission: Accurately transmits the encapsulated and adapted data to the target system; S5.43 Integration into the Target System: Data integration is performed within the target system, and the data can be invoked and applied. The operation of the final output and feedback of S5.5 includes: S5.51 Output Data Confirmation: Perform a quality review on the final data to ensure that the data meets the standards and industry requirements of the target system; S5.52 Feedback Mechanism: When the output is abnormal or non-standard, the feedback process is triggered in a timely manner to notify for correction; S5.53 Error Correction and Re-output: Correct the problematic data based on feedback and re-output it to the target system; The method for generating the n-gram in S2.221, which includes the ROUGE-N formula, can be either a cumulative statistical method or a fusion method after independent calculation. The first method, the cumulative statistical method, calculates the ROUGE-N recall rate by summing the n-gram match counts and total counts of all reference answers in a multi-reference scenario, thus evaluating the overall coverage of the model output for multiple references. The ROUGE-N index uses a cumulative statistical formula as follows: , Where gramn represents N consecutive words or characters, Count(gramn) represents the number of n-grams in the reference result, and Countmatch(gramn) represents the number of n-grams that are the same in the reference result and the model output; Wherein, the numerator is the sum of the number of n-grams that match the generated result in all reference texts; the denominator is the sum of the total number of n-grams in all reference texts. The second approach, independent calculation followed by fusion, calculates the ROUGE-N score independently for each reference in a multi-reference scenario. The results are then fused by taking the maximum or average value to fairly evaluate the model output's matching degree to diverse expressions. The formula for determining the maximum value of the ROUGE-N index is: , Where K is the number of reference answers; Countmatch and Count need to be based on the reference text. The formula for averaging the ROUGE-N index is as follows: , Where K is the number of reference answers; Countmatch and Count need to be based on the reference text. The optimal value of the BLEU index approaches 1 infinitely; The optimal value of the ROUGE metric approaches 1 infinitely; The ROUGE-N index can be applied using either the cumulative statistical method or the fusion method after independent calculation.

4. A method for generating and transmitting information, recording information, and sending information using a recording information module according to claim 1 or 2, characterized in that: The operation of the sending end of the S1. recording information module includes: S1. Model Generation Results: After receiving the input information, the model begins to generate output text or structured content; S2. Post-processing content generation: The model output is formatted, filtered, and classified at the sending end, and intermediate processing results are recorded; S3. Structured Template Management: The sending end applies preset templates according to task requirements to classify and record the generated results in a structured manner; S4. Task type identifier classification: The sending end adds a corresponding task type identifier to each generated result according to the task content.

5. A method for generating and transmitting information, recording information, and sending information according to the recording information module of claim 4, characterized in that: The operation of the receiving end of the S1. recording information module includes: S1. Receiving Record Information: The receiving end receives and parses the data from the sending end through the interface, and completes unpacking and parsing; S2. Data storage: The receiving end stores the received data according to a structured template, and all types of information are orderly and searchable; S3. Task Identifier Classification and Storage: The receiving end classifies and archives data according to task identifiers, and records explicit and implicit information to support traceability; S4. Data traceability and query: The receiving end uses structured templates and identifiers to trace the source of generated data and query the processing process; S5. Report Generation and Feedback: The receiving end generates feedback reports to display processing progress and quality, and supports evaluation and analysis of input matching degree; The operation of the sending end of the S2. scoring comparison module includes: S1. Data Generation and Standard Sample Preparation S1.1 Model Output Generation: Obtain or generate the output text of the model that needs to be evaluated; S1.2 Reference Sample Preparation: Retrieve standard reference texts corresponding to the model output as benchmark answers for comparison; S1.3 Field Information Annotation: Structured fields are written into the data packet, supporting field-level scoring; S2. Data Normalization and Preprocessing S2.1 Unified Format Processing: Format preprocessing of model output and reference samples to ensure consistency; S2.2 Field Extraction Processing: Key fields are separated in advance to provide independent input sources for field-level scoring; S3. Data Packaging and Transmission S3.1 Data Encapsulation: Model output, reference text, field information, and task ID are encapsulated into a standardized data structure or API request body; S3.2 Request Sending: Send the data packet to the receiving end of the scoring comparison module; S3.3 Record Request Information: Record the task ID, sending time, and request content, and support the association and traceability of scoring results; The receiving end of the S2. scoring comparison module is operated as follows: S1. Data Reception and Parsing S1.1 Request Reception: The receiving end listens to the interface or message queue to receive rating request data packets in real time; S1.2 Data Parsing: The receiving end parses the model output, reference text, and field information to extract the overall text and field content; S2. BLEU Index Calculation and ROUGE Index Calculation S2.1BLEU indicator unit operation S2.11 Data Preprocessing: The receiving end performs word segmentation, standardization, and field division on the text; S2.12 n-gram extraction and matching: The receiving end performs n-gram extraction on the whole and each field respectively, and calculates the word overlap rate; S2.13 BLEU Formula Calculation: The receiver combines precision, weight, and length penalty to calculate the overall and field-level BLEU scores; S2.14 Result Structuring: The receiving end outputs the BLEU scoring results in a structured manner according to fields or overall dimensions; S2.2ROUGE indicator unit operation S2.21 Data ROUGE Preprocessing: The receiving end segments the text into words or characters and divides it by field to prepare for ROUGE calculation; S2.22 n-gram extraction and ROUGE recall statistics: The receiving end calculates the overall and field-level ROUGE-1, ROUGE-2, and ROUGE-L indicators, covering recall rate and longest common subsequence; The method for transmitting the generated text, scoring results, and task template by modifying the sending end of module S3 includes: S1. Receive the generated content S1.1 Receive the generated results of a large language model LLM: Receive the generated text from a large language model LLM and identify its potential structural or semantic problems; S1.2 Similarity Assessment: The generated content is passed to the scoring and comparison module for similarity comparison and assessment; S2. Transmit the scoring comparison results S2.1 Delivering Evaluation Results: Receive generated text from a large language model (LLM) and identify its potential structural or semantic problems; S3. Tag and Label Transmission: The sender transmits the task type identifier and field template; S4. Transfer the generated text to the modification module: Transfer the generated text, scoring results, and related template data to the modification module to start the modification process.

6. A receiving method for automatically modifying the receiving end based on scoring comparison results or structured template rules according to claim 4, characterized in that: The operation of the modified module described in S3 includes: S1. Receive generated content and scoring results: The receiving end obtains the generated text and scoring results, and identifies the deviation area when there is a reference sample; S2. Operation with reference comparison and modification unit: Based on BLEU and ROUGE scores, optimize the generated text; S2.1 Keyword Insertion: Automatically inserts or supplements missing keywords according to preset rules; S2.2 Redundant Information Deletion: Remove redundant content; S2.3 Sentence Structure Adjustment: Optimize grammatical structure; S3. Operation of the no-reference comparison and modification unit: In the scenario of no reference text, the receiving end independently verifies and corrects the generated content according to the predefined structured template and syntax rules; S3.1 Field Integrity Check: Identify and complete missing key field information; S3.2 Illegal character filtering: Removes invalid or erroneous characters; S3.3 Semantic Logic Validation: Detect and correct logical errors and semantic deviations in the text; S4. Automatic modification mechanism: When the text does not conform to the standard, the receiving end triggers the automatic modification mechanism; S4.1 Complete Missing Fields: Automatically add missing key information items; S4.2 Correcting Errors: Correcting linguistic or semantic errors in the text; S4.3 Remove invalid content: Remove meaningless or duplicate information fragments; S5. Output standardized results: The receiving end outputs the modified and optimized text in a structured format and pushes it to the downstream system for integration and application.

7. A method for transmitting generated text and scoring comparison results triggered by a sending end according to claim 4, characterized in that: The operation of the S4 manual verification module includes: S1. Verification Trigger Detection: When the BLEU or ROUGE score is low or the structured output fails, the system automatically triggers the manual verification process. S2. Information Extraction: The sending end extracts and generates text, modification suggestions, and reference samples for verification; S3. Data Organization and Packaging: The sending end will structure and encapsulate the extracted content and generate a difference-highlighted view; S4. Data Push: The sending end stably sends the packaged verification data to the receiving end, supporting anomaly detection and timeout retry; The receiving end of the S4. manual verification and confirmation module performs manual intervention and verification of trajectory records and push of final modification results, including: S1. Information Reception and Parsing: The receiving end receives and parses the data from the sending end, displaying text differences and modification suggestions in a structured comparison view; S2. Manual Intervention and Interaction: The receiving end provides multiple manual operation options, supporting detailed editing and version comparison; S3. Verification result feedback: The receiving end writes the confirmed text and operation record into the system to form the official output version; S4. Process Tracking: The receiving end fully records the verification behavior and operation information, supporting audit trails and quality traceability.

8. A method for transmitting data, formatting data, and adapting the interface for output in a data output module according to claim 4, characterized in that: The S5 operation includes: S1. Receiving the generated content after processing: The sending end receives the final data content after scoring comparison, automatic modification, and manual verification and confirmation; S2. Data Formatting and Encapsulation: The sending end encapsulates the generated content into a standard format; S3. Interface Protocol Adaptation: The sending end adapts the data to the protocol layer according to the target system protocol; S4. Data output preparation: After the data is formatted and adapted, the sending end transmits the data to the receiving end; The method for receiving data through data integration and system interface using the receiving end of the S5 data output module includes: S1. Receiving Data: The receiving end obtains formatted data from the sending end; S2. Data Integration and System Interconnection: Integrate the received data and connect it to the industrial software platform; S3. Data Flow and Application: Facilitating the flow of data within the target system to support business modules; S4. Final Output: The final data is available in real time in the target system.

9. A post-processing system for human-computer interaction language model output that supports field-level scoring comparison and rule engine, characterized in that: The post-processing system includes an information recording module, a scoring comparison module, a modification module, a manual verification and confirmation module, and a data output module; The information recording module is responsible for recording and storing model generation results, structured templates, and task type identifiers to enable data traceability. The scoring and comparison module evaluates the similarity between the generated text and the reference text using the BLEU and ROUGE indices, providing quantitative and visual analysis of the generated results. The modification module performs format correction and semantic optimization based on the scoring comparison results to improve structural consistency and expression accuracy. The manual verification module intervenes manually when the automatic score is below a threshold. The data output module outputs the processed content to the industrial software system in a preset format, supporting multiple data format conversions and interface adaptations.

10. The post-processing system for human-computer interaction language model output supporting field-level scoring comparison and rule engine as described in claim 9, characterized in that: The information recording module includes a result recording unit, a structured template management unit, and a task type identifier classification sheet; the scoring comparison module includes a BLEU indicator unit and a ROUGE indicator unit. The modification module includes a modification unit with reference comparison and a modification unit without reference comparison; The manual verification and confirmation module includes an automatic extraction unit and a modification confirmation unit. The data output module includes a data format conversion unit and an interface adaptation unit; The result recording unit in the information recording module can completely record relevant information from model generation results and post-processing intermediate content; The BLEU index unit in the scoring comparison module can evaluate the similarity between the generated text and the reference text through n-gram matching, and supports the analysis of the overall text and key fields. The ROUGE indicator unit in the scoring and comparison module can evaluate the similarity and recall of generated content by assessing indicators such as n-gram overlap and longest common subsequence. The reference comparison and modification unit can automatically adjust the generated text by comparing the BLEU and ROUGE indices; The automatic extraction unit in the manual verification and confirmation module can automatically extract and generate text, modification suggestions, and reference samples; The interface adaptation unit in the data output module can adapt and output according to the interface protocols of different target systems, and seamlessly connect to different industrial-grade application systems.

11. The post-processing system and optimization method for human-computer interaction language model output supporting field-level scoring comparison and rule engine as described in claims 1-10, characterized in that: The post-processing system and optimization method are applicable to the entire process of standardized output systems in high-end intelligent equipment industrial scenarios. The post-processing system and optimization method described herein are applicable to the construction of intelligent agents, which encapsulate the functional modules of the intelligent agent into independent components with interface definitions.

Citation Information

Patent Citations

  • A digital power control question-answering method and system based on a large language model

    CN119003742B

  • Image data rapid screening method and device based on large model

    CN119377434A

Cited By

  • Methods, devices, electronic equipment, and storage media for extracting configuration parameters from performance appraisal scheme documents

    CN122311141A