Case file quality review method, system, and electronic device

CN122819218APending Publication Date: 2026-09-25YUNLIAN (INNER MONGOLIA) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611012983.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该方式存在以下缺陷:第一,案卷材料数量较多且类型复杂,人工逐页查找和反复比对容易遗漏;第二,不同材料之间存在跨页、跨文书的要素关联,人工在多窗口、多文件之间切换时难以保持较为一致的核对口径;第三,评查结论通常以文字记录或报告形式单独保存,问题与原始材料、具体字段、页面展示区域之间缺少机器可读的绑定关系,后续复核人员难以快速定位;第四,修改目录、补充识别、删除误报、重新校验和报告输出之间缺少统一流程,容易形成“发现问题”和“处理问题”脱节

Benefits of technology

[0009]本申请实施例采用的上述至少一个技术方案能够达到以下有益效果:通过生成统一的批次标识并对案卷材料建立包含材料顺序、名称、类型及影像关联的目录记录,为后续处理提供了可追溯的评查上下文,解决了离散材料难以统一管理的问题;对每条目录记录进行材料类型识别,并基于材料类型抽取结构化字段与目录记录关联,将非结构化的影像或文本转化为可计算、可校验的字段级数据,为自动化评查奠定了数据基础;在同一批次中协同执行全局空值校验、计算机规则校验和大模型自然语言校验,其中全局空值校验快速发现材料或字段缺失问题,计算机规则校验对结构化字段进行确定性规则检查确保可复现性,大模型自然语言校验对需要语义理解的内容进行智能判断,三种校验互补配合,显著提升了评查的全面性和准确性,避免了单一技术方案的局限性;校验结果至少包含批次标识、记录标识、材料分类、字段名称、结果来源类型、问题原因、建议说明和展示状态,使得每条校验结果都能精确定位到具体的批次、材料、字段和展示位置,复核人员可以直接根据结果定位问题,无需反复查找原始材料,大幅提高了复核效率;最终基于结构化的校验结果生成评查报告,保证了报告内容与系统评查状态的一致性,形成了从案卷导入、识别、抽取、校验到报告输出的完整闭环,从而有效降低了人工评查的工作强度和遗漏风险,缩短了评查周期,提升了案卷质量评查的标准化程度和整体工作效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819218A_ABST
    Figure CN122819218A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a case quality evaluation method and system and an electronic device. The method comprises the following steps: receiving case materials to be evaluated, and generating batch identifiers of the case materials to be evaluated; establishing directory records for the case materials to be evaluated; performing content identification on the case materials to be evaluated corresponding to each directory record to obtain material types of the case materials to be evaluated; calling corresponding extraction algorithms according to the material types of the case materials to be evaluated, positioning and extracting key contents of the case materials to be evaluated to obtain extraction results; converting the extraction results into structured fields and associating the extraction results to corresponding directory records; performing global null value checking, computer rule checking and large model natural language checking on the case materials in the same batch to obtain checking results; and generating an evaluation report based on the checking results. The application improves the standardization degree and overall work efficiency of case quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method, system, and electronic device for reviewing case file quality. Background Technology

[0002] Case files typically consist of scanned images, catalog records, various types of documents, tabular materials, natural language descriptions, attachments, and subsequent supplementary materials. These case files are characterized by a wide variety of material types, numerous pages, scattered fields, inconsistent document formats, and the potential for the same fact or information about the same subject to appear repeatedly across different materials. Case file quality assessment requires not only checking the completeness of materials and the absence of empty fields, but also verifying the factual statements, chronological order, personnel or event elements, procedural nodes, and whether there are inconsistencies, incompleteness, or logical anomalies in the content of different materials.

[0003] Current manual review methods typically rely on reviewers opening case file images or electronic documents page by page and checking materials, fields, and content item by item according to the review checklist. This method has the following drawbacks: First, the case file materials are numerous and complex, making it easy to miss items when manually searching and comparing page by page; second, there are cross-page and cross-document element connections between different materials, making it difficult for manual reviewers to maintain a consistent verification standard when switching between multiple windows and files; third, review conclusions are usually saved separately in the form of text records or reports, and there is a lack of machine-readable binding between problems and original materials, specific fields, and page display areas, making it difficult for subsequent reviewers to quickly locate them; fourth, there is a lack of a unified process between modifying the table of contents, supplementing identification, deleting false alarms, re-verifying, and outputting reports, which can easily lead to a disconnect between "identifying problems" and "solving problems." Summary of the Invention

[0004] In view of the above situation, this application provides a case file quality assessment method, system, and electronic device, which aims to solve the above problems or at least partially solve the above problems.

[0005] In a first aspect, embodiments of this application provide a method for reviewing case file quality, the method comprising: Receive case files to be reviewed and generate a batch identifier for the case files to be reviewed; A catalog record shall be established for the case file materials to be reviewed. The catalog record shall include at least one of the following: the order of each material, the name of the material, the material type information, the attachments, and the image association information; Content identification is performed on the case file materials to be reviewed corresponding to each directory record to obtain the material type of the case file materials to be reviewed; Based on the material type of the case file to be reviewed, the corresponding extraction algorithm is invoked to locate and extract the key content of the case file to be reviewed, and the extraction result is obtained; the extraction result is converted into structured fields and associated with the corresponding directory records; Global null value validation, computer rule validation, and large-scale model natural language validation are performed on the case file materials in the same batch to obtain the validation results. The global null value validation is used to check whether key materials or key fields are missing. The computer rule validation is used to perform deterministic rule checks on the extracted structured fields. The large-scale model natural language validation is used to perform semantic validation on the case file materials to be reviewed based on the validation task. The validation results include at least batch identifier, record identifier, material classification or material type, field name or field identifier, result source type, problem reason, suggestion explanation, and display status. An evaluation report is generated based on the verification results.

[0006] Secondly, embodiments of this application also provide a case file quality assessment system, including: The batch management module is used to receive case files to be reviewed and generate batch identifiers for the case files to be reviewed. The catalog and record module is used to create a catalog record for the case file materials to be reviewed. The catalog record includes at least one of the following: the order of each material, the material name, the material type information, the attachments, and the image association information; The content recognition module is used to perform content recognition on the case file materials to be reviewed corresponding to each directory record, and to obtain the material type of the case file materials to be reviewed. The structured extraction module is used to locate and extract the key content of the case materials to be reviewed based on the material type of the case file to be reviewed and call its corresponding extraction algorithm to obtain the extraction results; the extraction results are converted into structured fields and associated with the corresponding directory records; The verification module is used to perform global null value verification, computer rule verification, and large-scale model natural language verification on case file materials in the same batch to obtain verification results. The global null value verification is used to check whether key materials or key fields are missing. The computer rule verification is used to perform deterministic rule checks on the extracted structured fields. The large-scale model natural language verification is used to perform semantic verification on the case file materials to be reviewed based on the verification task. The verification results include at least batch identifier, record identifier, material classification or material type, field name or field identifier, result source type, problem reason, suggestion explanation, and display status. The report generation module is used to generate an evaluation report based on the verification results.

[0007] Thirdly, embodiments of this application also provide an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the steps described in the first aspect.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform the steps described in the first aspect.

[0009] The above-mentioned technical solutions adopted in this application embodiment can achieve the following beneficial effects: by generating a unified batch identifier and establishing a catalog record for case file materials containing material order, name, type, and image association, a traceable review context is provided for subsequent processing, solving the problem of difficult unified management of discrete materials; material type identification is performed on each catalog record, and structured fields are extracted based on the material type and associated with the catalog record, transforming unstructured images or text into calculable and verifiable field-level data, laying a data foundation for automated review; global null value verification, computer rule verification, and large model natural language verification are performed collaboratively in the same batch, where global null value verification quickly discovers material or field missing problems, computer rule verification performs deterministic rule checks on structured fields to ensure reproducibility, and large model natural language verification performs semantic analysis... The system intelligently judges the content of the solutions, and the three types of verification complement each other, significantly improving the comprehensiveness and accuracy of the review and avoiding the limitations of a single technical solution. The verification results include at least batch identifier, record identifier, material classification, field name, result source type, problem reason, suggestion explanation, and display status, so that each verification result can accurately locate the specific batch, material, field, and display position. Reviewers can directly locate the problem based on the results without repeatedly searching for the original materials, which greatly improves the review efficiency. Finally, the review report is generated based on the structured verification results, ensuring the consistency between the report content and the system review status. This forms a complete closed loop from case import, identification, extraction, verification to report output, thereby effectively reducing the workload and omission risk of manual review, shortening the review cycle, and improving the standardization and overall work efficiency of case quality review. Attached Figure Description

[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating the case file quality assessment method provided in an embodiment of this application is shown; Figure 2A flowchart of a case file quality assessment method provided in another embodiment of this application is shown; Figure 3 This paper shows a structural diagram of the case file quality assessment system provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the term "comprising" and its variations should be interpreted as open-ended terms meaning "including but not limited to."

[0013] Figure 1 This document illustrates a flowchart of the case file quality assessment method provided in an embodiment of this application. Figure 1 It can be seen that this application includes at least steps S101-S106: Step S101: Receive the case file materials to be reviewed and generate a batch identifier for the case file materials to be reviewed.

[0014] Among them, the case files to be reviewed include various case files that require quality checks in various systems, including scanned images, electronic documents, tabular materials, natural language descriptions, and subsequent supplementary materials.

[0015] Specifically, the system receives case files to be reviewed, uploaded by users or imported in other ways. All case files imported at once and requiring centralized review are packaged into a single review task unit. For example, a supervisory authority might review 212 pages of case files for a particular case as a batch. The system automatically generates a unique batch identifier for this batch of case files, such as "BATCH_20260606_001," to uniquely identify this review task.

[0016] Step S102: Create a catalog record of the case file materials to be reviewed.

[0017] In some embodiments, catalog records are generated one by one based on the file list of imported materials, an existing electronic catalog, or a user-specified order.

[0018] The catalog record includes at least one of the following: the order of each material, the material name, the material type information, attachments, and image association information. Specifically, the material order indicates the material's sequence number within the entire case file, such as page 1, page 2, ..., page 212. This sequence information determines the order in which the catalog is displayed and is also used to present issues sequentially in the report. The material name is the title or file name of the material, such as "Preliminary Verification Approval Form," "Interview Record," "Search Record," etc. The material name can be directly taken from the file name, or it can be manually entered by the user or automatically filled in by the system based on the recognition results. The material type information indicates the case file category to which the material belongs, such as "Procedure File," "Approval Form," "Interview Record," "Factual Material," etc. The material type information can be automatically determined by the system through the material type recognition module, or it can be manually specified or corrected. This information is crucial for determining the subsequent field extraction templates and validation rules. Attachments refer to the original file storage path or file identifier associated with this catalog record. Attachments can be image files, PDF files, or other electronic documents. The system can load and display original images or document content when needed through attachment association. In addition to the attachment path, image association information can also include metadata such as the image's page number, resolution, storage location, and thumbnail references, making it easy for the front-end page to quickly preview the original image.

[0019] In some embodiments, the system also provides the display and maintenance of catalog records to enable reviewers to perform visual management and manual intervention of case file materials.

[0020] Step S103: Perform content recognition on the case file materials to be reviewed corresponding to each directory record to obtain the material type of the case file materials to be reviewed.

[0021] Specifically, each directory record established in step S102 is used as the basic processing unit, rather than performing batch recognition on the entire batch. This means that each individual case file is sent into the content recognition process separately. This design ensures that the boundaries between materials are preserved, and the type labels and text content of each material can be managed independently, facilitating subsequent field extraction, rule validation, and result location by material.

[0022] Furthermore, the document type identification is performed on the case file materials corresponding to the current directory record. Document type refers to the document category to which the material belongs in the case file, such as "interview record," "preliminary verification approval form," "case filing and investigation approval report," "search record," and "cadre resume." Identification methods can employ one or more combinations of the following: rule matching based on filename, detection based on document title keywords, classification models based on layout features, or semantic classification using a large language model. The document type identification result determines which set of field extraction templates, which validation rules, and how the front-end parsing area should be displayed in subsequent steps.

[0023] Step S104: Based on the material type of the case file to be reviewed, call its corresponding extraction algorithm to locate and extract the key content of the case file to be reviewed, and obtain the extraction result; convert the extraction result into structured fields and associate them with the corresponding directory records.

[0024] Specifically, different types of case files have different information organization and semantic structures. For example, "interview transcripts" require attention to fields such as interview time, interview location, interviewer, interviewee, informed content, question and answer content, and signature; while "preliminary verification approval forms" require extraction of fields such as the handling department, clue code, basic information of the person being verified, summary of the reported issue, approval time, and attachment description. To this end, a corresponding pre-extraction algorithm is pre-configured for each type of material. Different pre-extraction algorithms correspond to different field extraction templates, which define the field names, field types, extraction rules, and default value processing methods for that type of material.

[0025] Furthermore, based on the field extraction template, the case file materials to be reviewed are parsed and extracted. Extraction methods may include, but are not limited to, one or more of the following: rule-based extraction, such as using regular expressions, keyword positioning, template matching, etc., to extract values ​​of specific fields from the case file materials; for tabular materials or documents with a fixed layout, extraction is performed based on the coordinate area of ​​the field on the page or the row and column position of the table; named entity recognition-based extraction, using a pre-trained named entity recognition model to identify entities such as names of people, places, organizations, and dates from the case file materials and map them to corresponding fields; and large language model-assisted extraction, for complex fields that are difficult to extract accurately through rules or small models, a large language model can be invoked to perform semantic extraction based on field definitions and the full-text context.

[0026] After extraction, each field forms a key-value pair: field name and field value. For fields that fail to extract valid values ​​or do not exist, their field values ​​can be marked as empty or with a special placeholder for subsequent null value validation.

[0027] Furthermore, the extracted structured fields are associated with and saved in the current directory records. The association method includes: creating a record for each field value in the structured field table of the database. This record must contain at least the following information: field unique identifier, directory record identifier, batch identifier, field name, field value, material type, and extraction time.

[0028] By using the catalog record identifier, all field values ​​are bound to the catalog record from which they originate, allowing all structured fields under a material to be queried through the catalog record; conversely, any field value can be traced back to the material and batch to which it belongs.

[0029] When users discover errors in the automatically extracted structured fields, such as incomplete, misaligned, or missing field values, they can manually edit the field values ​​through the front-end interface or trigger a "re-identification" command. Upon responding to the re-identification command, the structured field extraction process is re-executed, and the field data in the database is updated. Through this mechanism, the structured field extraction module and manual review form a feedback loop, ensuring the accuracy and reliability of the field data.

[0030] In addition, for some pre-defined special material types, OCR or full-text recognition can be performed on the case files to be reviewed to generate full-text information that is associated with and saved in the catalog records; the full-text information is used for natural language verification of large models and manual full-text viewing.

[0031] Specifically, full-text information recognition is performed on case files containing full-text identification tags to extract the text content into a format suitable for computer processing and human reading. The specific techniques employed vary depending on the material format. For example, for materials identified as requiring full-text access, the system uses an OCR engine to recognize all text content and output the full-text. For non-identified material types, full-text extraction can be performed using structured data acquisition methods.

[0032] After recognition is complete, the obtained full-text content is associated with the record identifier of the current directory and saved to the database. The association method can be as follows: add a new record to the full-text content table, which includes fields such as record identifier, full-text text, recognition method identifier, and creation time. Using the record identifier, the full-text content of the material can be read at any time in subsequent steps for structured field extraction, large-scale model semantic verification, and front-end full-text viewing and display.

[0033] Step S105: Perform global null value verification, computer rule verification, and large model natural language verification on the case file materials in the same batch to obtain the verification results.

[0034] Global null value validation is used to check whether key materials or key fields are missing. Computer rule validation is used to perform logical rule checks on the extracted structured fields, and large model natural language validation is used to perform semantic validation on the case file materials to be reviewed based on the validation task. The validation results should include at least the batch identifier, record identifier, material classification or material type, field name or field identifier, result source type, problem reason, suggestion explanation, and display status.

[0035] Step S106: Generate an evaluation report based on the verification results.

[0036] In some embodiments, all verification results belonging to the same material are summarized under the material name, organized by directory record. For each verification result, the field names involved, specific problem descriptions, root cause analysis, and modification suggestions are listed. The report can also summarize the problem statistics for the entire batch, such as the total number of problems, the number of problems distributed by material type, and the number of problems counted by result source type (global null values / computer rules / large models). PDF format can be used as the output format when generating the report. The generated report can be downloaded and saved.

[0037] Because the report's data comes from structured validation results already bound to batches, records, and fields, rather than temporary page screenshots or manually compiled text, the report content maintains strict consistency with the review status. When reviewers edit, delete, or trigger re-validation of the validation results, the data is automatically updated, ensuring the output report reflects the latest review conclusions.

[0038] from Figure 1As can be seen from the method described, this application generates a unified batch identifier and establishes a catalog record for case file materials that includes material order, name, type, and image association, providing a traceable review context for subsequent processing and solving the problem of difficult unified management of discrete materials. It identifies the material type of each catalog record and extracts structured fields based on the material type, associating them with the catalog record. This transforms unstructured images or text into computable and verifiable field-level data, laying a data foundation for automated review. Within the same batch, it collaboratively performs global null value verification, computer rule verification, and large-scale model natural language verification. Global null value verification quickly identifies missing materials or fields, computer rule verification performs deterministic rule checks on structured fields to ensure reproducibility, and large-scale model natural language verification intelligently judges content requiring semantic understanding. The three verification methods complement each other, significantly improving the comprehensiveness and accuracy of the review and avoiding the limitations of a single technical solution. The verification results include at least batch identifier, record identifier, material classification, field name, result source type, problem cause, suggestion explanation, and display status. This allows each verification result to be accurately located to the specific batch, material, field, and display position. Reviewers can directly locate the problem based on the results without repeatedly searching for the original materials, greatly improving review efficiency. Finally, the review report is generated based on the structured verification results, ensuring the consistency between the report content and the system review status. This forms a complete closed loop from case import, identification, extraction, verification to report output, effectively reducing the workload and omission risk of manual review, shortening the review cycle, and improving the standardization and overall work efficiency of case quality review.

[0039] In some embodiments of this application, after obtaining the material type of the case file materials to be reviewed, the method further includes: determining whether the directory record and material type of the case file materials to be reviewed are correctly identified; if the directory record and material type of the case file materials to be reviewed are incorrectly identified, performing maintenance operations. The maintenance operations include at least one of the following: re-identifying the directory, re-identifying the file type, inserting a category or image before the current directory, replacing the category and image, and deleting the current directory record.

[0040] Specifically, the system's front-end page displays a directory area, which shows the following information: Directory Items: The name or title of each case file material, arranged in order to form a directory list; Material Type: The category of the case file material corresponding to each directory record, such as "Interview Record", "Approval Form", "Preliminary Verification Report", etc.; Record Pagination Information: When the number of directory records exceeds a preset threshold, the system automatically performs pagination and displays the current page number, total number of pages, and number of records per page, supporting page turning.

[0041] Meanwhile, the system receives at least one of the following operation commands from the user regarding the catalog records: Switching the current record: When the user clicks on a catalog item, the system sets that record as the current active record and loads the corresponding original image, full text or HTML parsing content, and associated verification results; Re-identifying the current record: When the user believes that the system's identification of the material type or content of the current record is inaccurate, a re-identification command can be triggered. The system re-executes material type identification, OCR full text recognition, and structured field extraction, and updates the relevant data; Inserting a category or image before the current catalog: The user can insert a new catalog item before the selected catalog record. This catalog item can be a category label or an attached image file; Deleting the current catalog record: When the user confirms that the material corresponding to the current catalog record does not belong to this case file or is an invalid page, the user can delete the catalog record.

[0042] In response to any of the above operation commands, the system performs the corresponding directory maintenance operations, including updating the directory record table in the database, adjusting the record sequence number, reloading the front-end directory list, and refreshing the relevant display content.

[0043] In this embodiment, the aforementioned display and operation functions provide a unified interactive entry point for various types of case file materials, enabling reviewers to quickly locate, switch, correct, or add / delete materials. Simultaneously, it provides a record-level, precise positioning foundation for subsequent material identification, field extraction, rule validation, and review report generation. Furthermore, an effective feedback loop is formed between manual intervention and automated machine processing: when deviations occur in the catalog or identification results, reviewers can directly correct them through catalog maintenance operations. The system then updates the data and re-triggers the validation process, ensuring that the review results always remain consistent with the actual state of the case files.

[0044] In some embodiments of this application, after extracting the structured fields of the case file materials to be reviewed, the method further includes: backfilling the structured fields, identified content, or template content into the HTML display content to form a file parsing page for comparison and verification with the case file materials.

[0045] Specifically, based on the material type recorded in the current directory, the corresponding HTML template is obtained. The HTML template predefines the display structure for that material type, such as the position of field tags, table layout, paragraph styles, and interactive areas. Subsequently, the structured fields extracted in step S104 are filled into the corresponding positions in the HTML template. In addition to the structured fields, the full-text recognition content obtained in step S103 or the fixed content preset in the template can also be backfilled into the HTML page, thereby forming a complete and readable document parsing page.

[0046] Furthermore, the generated HTML content is presented in the user interface through the front-end page. Users can simultaneously view the original image of the case file, such as a scanned image, and the document parsing page on the right, comparing and verifying whether the extracted fields are consistent with the original image. For example, the reviewer checks whether the signature at the end of the interview transcript in the original image is handwritten, and at the same time checks whether the signature field exists in the HTML parsing page on the right.

[0047] In the above embodiments of this application, the step of backfilling structured fields, identified content, or template content into the HTML display content to form the file parsing page specifically uses HTML template filling to generate the display content. However, this application is not limited to the specific HTML backfilling or rendering method described above. In other embodiments, template filling, server-side rendering, front-end rendering, or other page generation methods can be used to display the extraction results. For example, in one alternative, the system can pre-build a complete HTML page on the server, embed the structured fields into the page template, and then send the entire page to the front-end for display. In another alternative, the system can use front-end rendering, where the backend only returns structured JSON data, and the frontend dynamically generates and displays HTML content using a JavaScript framework. Furthermore, a hybrid approach can be used: the page framework and styles are pre-set on the frontend, and the backend only provides API interfaces for field data and identified content, with the frontend responsible for filling and rendering.

[0048] Regardless of the specific implementation method used, as long as the displayed content can be associated with the directory records, fields, and verification results, and allows reviewers to locate and view the original materials, it does not fall outside the scope of protection of this application. These alternative solutions all fall within the concept of this application: "converting structured fields and identified content into page objects that can be manually verified."

[0049] In this embodiment of the application, unstructured case materials are transformed into interactive and comparable structured page objects by using HTML form-based fill-in and display methods, which significantly improves the efficiency and accuracy of manual review.

[0050] In some implementations of this application, the verification methods in step S105 include global null value verification, computer rule verification, and large model natural language verification.

[0051] In some embodiments, global null value validation is performed on the entire batch of case files, and the content checked includes at least one of the following: whether key materials are missing in the batch, whether necessary structured fields are empty, and whether core extraction results have not been generated.

[0052] Specifically, based on the case file type and review rules, it is determined whether a certain type of necessary material should be present within the batch. For example, a complete disciplinary inspection and supervision case file should include key documents such as "Preliminary Verification Approval Form," "Preliminary Verification Report," "Case Filing and Investigation Approval Report," "Interview Records," and "Materials on the Facts of Violations of Discipline and Law." Global null value checking will traverse all catalog records within the batch; if a certain type of necessary material is not identified or does not appear in the catalog, a corresponding missing value check result will be generated.

[0053] For existing materials with extracted structured fields, check whether their key fields have valid values. For example, if the "Conversation Start Time" field in the "Conversation Transcript" is empty, or if no text is extracted from the "Information Content" field, it is considered a field null value issue. This type of check does not depend on whether the materials are complete, but rather on the completeness of the fields in existing materials.

[0054] This check verifies whether the structured extraction process for certain materials, although correctly identified, has been successfully executed and produced the expected field data. For example, if the material type of a certain "approval form" has been correctly identified, but field extraction fails due to low OCR quality or abnormal layout, the system will record "extraction result not generated" as a null value check result, prompting the reviewer to pay attention to the recognition quality of the material.

[0055] Global null value validation is typically performed before computer rule validation and large model semantic validation. Global null value validation can quickly identify fundamental integrity issues, avoiding wasted time performing subsequent complex validations when fields or materials are missing. Validation results are uniformly written into the result data structure and bound to batch identifiers, material types, or field names for front-end display and manual review.

[0056] In some embodiments, the checks performed by computer rule validation include at least one of the following: field existence validation, logical relationship validation between fields, condition judgment validation, cross-material field comparison validation, and multi-field combination rule validation.

[0057] Specifically, field existence checks are used to verify whether a specific field exists in the structured field set, or whether the field's value is not empty. For example, for "conversation transcript" materials, the "conversation start time" field is checked to ensure it exists and is not empty.

[0058] The field logical relationship validation is used to check whether the logical relationship between multiple fields within the same document or across documents is reasonable. For example, it checks whether the "conversation start time" is earlier than the "conversation end time"; and whether the "case filing time" is not earlier than the "preliminary verification time".

[0059] Conditional validation is used to judge field values ​​based on specific conditions. For example, if the value of the "Amount Involved" field exceeds a certain threshold, then the "Leadership Approval" field must exist.

[0060] Cross-material field comparison verification is used to compare the consistency of the same or related fields in different materials. For example, it verifies whether the name, work unit, and position of the person being verified in the "Preliminary Verification Approval Form" are consistent with the corresponding fields in the "Case Filing Review and Investigation Approval Report".

[0061] Multi-field combination rule validation is used to validate cases where multiple fields together satisfy a complex condition. For example, for materials concerning violations of discipline and law, it can validate whether the "description of violations of discipline and law" field simultaneously includes elements such as time, location, and behavior.

[0062] Once the above computer rule validation results are bound to records and fields, they can be read, located, and reported by the front end.

[0063] In some embodiments, large-scale natural language verification includes: constructing a verification task based on batch context, material type of catalog records, full-text information, structured fields, and verification rule descriptions; calling the large language model so that the large language model performs semantic judgment only based on the case file material content and rule requirements in the verification task; and converting the output of the large language model into a unified verification result, which includes a problem description, problem cause, suggestion description, and result source type.

[0064] Specifically, the batch context is used to limit the scope of verification, the material type is used to determine the document category to which verification applies, the full-text information and structured fields provide the original data to be verified, and the verification rule description clarifies the specific semantic issues that need to be judged in this verification. Furthermore, a large language model is invoked, and it is made to make semantic judgments solely based on the content of the case file materials and rule requirements in the verification task, without expanding external knowledge. That is, the large model's judgment is limited to the material information already existing within the current case file batch, without introducing internet knowledge or other irrelevant content from the model's pre-training, to ensure the traceability of the verification results and their strict correspondence with the case file materials. Finally, the output of the large language model is transformed into a unified verification result. The problem description summarizes the specific problems found in the semantic verification, the problem cause explains the basis for the model's judgment, the suggestion provides directions for modification or supplementary material content, and the result source type indicates that the result originates from the large model's natural language verification. This unified verification result uses the same data structure as the results generated by global null value verification and computer rule verification, facilitating unified storage, display, and review.

[0065] Furthermore, it is worth noting that in practice, situations may arise where the contextual information upon which the verification task is based is insufficient, key content in the case file is missing, or the rule description itself is not clear enough, leading to the model's inability to make a definitive judgment. To avoid false alarms or misleading statements caused by the large language model forcibly outputting absolute conclusions under insufficient information, this application adopts the following processing mechanism: Transfer to manual review: Mark this verification task as "Pending manual review" and save the model's output comments as supplementary information. When reviewers view the task on the front-end page, they will see that this issue requires manual judgment based on the original materials and will not be automatically included as a definitive issue in the final report.

[0066] No deterministic verification result: The verification task is abandoned directly, and no verification result is generated. This approach is suitable for scenarios where context is severely lacking and the model cannot provide any meaningful judgment.

[0067] Through the aforementioned processing mechanism, this application ensures the reliability and verifiability of the natural language verification results of the large-scale model, avoiding false conclusions arising from the model's forced responses when information is insufficient, and enabling the large-scale language model to truly become a semantic judgment tool to assist in evaluation. At the same time, this mechanism also leaves ample room for manual review, enabling a reasonable division of labor and collaboration between automatic machine verification and professional human judgment.

[0068] In some embodiments of this application, before generating an evaluation report based on the verification results, the method further includes: displaying the verification results and receiving manual review operations; the manual review operations include at least one of the following: editing or deleting the verification results, and triggering re-identification of the identified content; when re-identification is triggered, global null value verification, computer rule verification, and large model natural language verification are re-executed based on the updated identified content to form new verification results.

[0069] Specifically, the set of verification results generated in step S105 is displayed on the front-end interface. Display methods include, but are not limited to: displaying error markers or error counts next to material items in the directory area; highlighting or inserting error message icons for fields or text fragments associated with the issues on the file parsing page; and centrally displaying all verification results in a separate error list, with each result including an issue description, cause, suggested explanation, and result source type. Reviewers can easily understand at a glance what issues exist in the current batch, and the materials and fields involved in each issue.

[0070] The system provides an interactive entry point in the front-end interface, allowing reviewers to perform manual operations on verification results or identified content. For example, if a reviewer believes the system-generated verification result description is inaccurate, the cause analysis is incomplete, or the suggested explanation needs supplementation, they can directly modify the content of the corresponding field on the interface. The edited result will overwrite the original machine-generated content for use in subsequent report generation. As another example, if a reviewer confirms a verification result is a false alarm, they can delete that result. After deletion, the issue will no longer appear in the result set and will not be included in the final report. Furthermore, if a reviewer finds an error in material type identification when comparing the original image with the system's analyzed content, they can trigger a re-identification command. This command can be issued by clicking the "Re-identify" button next to the directory item, or by selecting the "Re-identify" menu on the analysis page.

[0071] Furthermore, when the system receives a re-identification trigger command, it re-executes steps S103-S105 based on the updated identification content: First, it re-identifies the material type and / or performs OCR / full-text recognition on the current directory records to obtain new structured fields; finally, based on the updated directory records, full-text information, and structured fields, it re-executes global null value verification, computer rule verification, and large-model natural language verification to generate a new set of verification results. The newly generated verification results will overwrite or supplement the original verification results, ensuring that the result set always reflects the latest status of the case file materials.

[0072] In this embodiment, the aforementioned manual review and re-verification mechanism effectively combines automated machine verification with professional human judgment. Reviewers can correct system identification errors and false alarms, while the system can quickly respond to manual corrections, automatically re-execute the verification process, and generate updated results, avoiding the tedious manual modification of each issue description. Furthermore, this mechanism forms a closed-loop workflow of "identification—verification—review—correction—re-identification—re-verification," continuously improving the quality of the review results with human intervention, ultimately generating an accurate and reliable review report.

[0073] Based on the above embodiments, the following is in conjunction with Figure 2 This application provides an explanation of the case file quality assessment methods.

[0074] Case file material import: The system receives case file materials to be reviewed uploaded or imported by users, covering various types of documents such as scanned images, electronic documents, catalog records, tabular materials, natural language statement materials and subsequent supplementary materials in political and legal and discipline inspection and supervision business.

[0075] Create or select a case file batch: The system creates a brand new case file batch for the case file materials imported this time, or allows users to select an existing case file batch and generate a unique batch identifier for the batch. This identifier will serve as a unified association key for all data throughout the entire process.

[0076] Generate catalog records and determine the record order: The system automatically generates catalog records for the batch based on the imported case file materials themselves or the existing catalog information provided by the user. It records the order, name, type, attachment association and image association information of each material, and determines the arrangement order of each catalog record.

[0077] Accessing the Case File Catalog Review Page: The system loads all catalog records and related data for this batch, and you enter the front-end interactive page for the case file catalog review.

[0078] View the catalog, original images, full text, and analysis content: On the review page, users can simultaneously view the catalog list of the case files for this batch, the original images of the currently selected material, the full-text recognition content, and the structured analysis content.

[0079] Determine if the directory or content needs to be re-identified: The user compares the system's identification results with the actual case file materials to determine if there are any deviations or errors in the directory structure, material type identification, or content identification.

[0080] If the judgment result is "yes": the system will perform the operation of re-identifying the catalog, material type or content, and re-identify the specified catalog item or material; after the identification is completed, the system will automatically update the full-text identification content, structured field data and HTML parsing display content, and return to the "view catalog, original image, full text and parsing content" step for the user to check and confirm again.

[0081] If the judgment result is "No": the user triggers a one-click check operation to start the automated review process.

[0082] Perform global null value validation: The system uses the entire case file batch as the execution scope to check whether key materials are missing, whether necessary structured fields are empty, and whether core extraction results have not been generated, thus forming a global null value validation result.

[0083] Perform computer rule verification: Based on the extracted structured fields and preset rule configurations, the system performs deterministic rule checks, including field existence, logical relationships between fields, condition judgments, cross-material field comparisons, and multi-field combination rules, to form computer rule verification results.

[0084] Perform large-scale model semantic verification: For verification items that require natural language understanding, the system constructs a verification task based on material type, full-text recognition content, structured fields, rule descriptions, and questions to be verified, calls the large language model to perform semantic judgment, and forms a large-scale model semantic verification result.

[0085] Unified storage and binding of validation results: The system stores all results generated by global null value validation, computer rule validation, and large model semantic validation as a structured result object, and accurately binds each validation result with the corresponding batch identifier, directory record identifier, field identifier, and page display position.

[0086] Display of directory, parsed content, and error location areas: The system centrally displays all verification results in the directory items, parsed content area, and dedicated error location area on the front-end page, and visually marks the directory items and specific fields with problems.

[0087] Manual review, editing, or deletion of results: Reviewers will review the verification results one by one based on the error location information displayed on the page, comparing them with the original images and analysis content; they will edit and correct inaccurate results and delete false or inapplicable results.

[0088] Determine if re-verification is needed: Based on the verification results, the reviewer determines whether the automated verification needs to be re-executed due to reasons such as material identification errors, manual field corrections, or supplementary case file materials.

[0089] If the judgment result is "yes": return to the "trigger one-click check" step, and the system will re-execute global null value verification, computer rule verification and large model semantic verification based on the updated case file data.

[0090] If the judgment result is "No": execute the operation of generating an evaluation report and merging PDFs.

[0091] Generate review report: Based on all verification results that have been manually reviewed and confirmed, the system automatically generates a standardized case file quality review report and supports outputting the review report as a PDF file, completing the complete closed-loop process of this case file quality review.

[0092] Based on the above embodiments, compared with manual page-by-page review, this application improves the traceability of the review process by incorporating case materials, catalog records, full-text content, structured fields, verification results, and report output into the same review context through batch management. Reviewers can view the catalog, original images, parsed content, and verification results within the same batch, reducing the problem of discrepancies between results and original materials.

[0093] Compared to traditional systems that only store electronic files or manual forms, this application transforms unstructured case file images into displayable, calculable, and verifiable data objects through material classification and recognition, OCR / full-text recognition, and structured field extraction. Different material types can enter different field extraction and display processes, making the processing of various types of case file materials more consistent.

[0094] Compared to a single OCR solution, this application goes beyond simply outputting plain text; it further transforms the recognition results into structured fields and HTML display content. Through HTML form-based backfilling and parsing, reviewers can verify the fields against the original image on a page that closely resembles the material's structure, facilitating the identification or extraction of problems.

[0095] Compared to single-rule validation schemes, this application integrates global null value validation, computer rule validation, and large-scale natural language model validation in stages. For problems that can be formally expressed, deterministic rules are used to enhance reproducibility; for problems requiring semantic understanding, a large language model is used to process contextual relationships and semantic consistency in natural language materials, thereby enhancing the synergistic ability of rule checking and semantic checking.

[0096] Compared to a single large language model approach, this application constrains the large model input through batch, record, full-text, field, and rule contexts, and uniformly writes the large model results into a queryable, editable, deleteable, and reportable result structure. Thus, the output of the large model is transformed into structured verification results bound to specific materials and fields, which can be directly used for review and report generation, forming traceable evaluation conclusions.

[0097] By binding the verification results with batches, catalog records, fields, material categories, and page display areas, this application makes error location more precise. Reviewers can access specific materials from the catalog error messages, view the issues by combining the original images, full text, or HTML parsing, and then edit or delete the results. This mechanism helps reduce the workload of repeatedly searching for material locations during review.

[0098] Through report generation, this application can transform the structured results of the review process into deliverable outputs. The report originates from the bound and reviewed result data, which helps to maintain consistency between the report content and the system review status, forming a closed loop of import, identification, verification, review, and reporting.

[0099] The following uses 639 batches of simulated case files that have been imported into the system and completed the review as a specific example to illustrate the specific usage method, operation steps, and work process of this application in the intelligent case file quality review scenario. The case file content in this example is simulated data and does not need to be confidentialized as in real cases. In order to make the example reflect the actual operating status of the system, the following description directly combines batch, time, case file images, field extraction results, verification information, and report generation status, but does not advocate any processing time, recognition accuracy, or efficiency indicators not supported by system records.

[0100] I. Implementation Examples, Batch Time, and Data Range: This example corresponds to the case review task with system batch ID 639, batch number DJ202600003, batch tag DJ, year 2026, serial number 3, set tag party_and_government, and personnel level tag junior_officer. The case file subject of this batch is A. The batch creation time recorded by the system is January 17, 2026, 17:27:26, and the last update time is May 9, 2026, 15:05:02. The case filing time field is not currently recorded for this batch.

[0101] In terms of data scale, this batch contains 212 case records, 421 structured field data entries, including 32 master records with structured field data, and 4 full-text recognition data entries. The verification result table records 41 review results, including 17 computer rule verification results and 24 large-model semantic verification results. This fact only reflects the current database status and does not affect the system's functional chain of generating PDF review reports based on existing verification results.

[0102] From the perspective of the composition of the case file materials, batch 639 primarily processes image-based case file records, with attachments named in a sequential numbering format such as 0001.jpg, 0002.jpg, 0003.jpg, etc. The catalog recognition results include both main category materials and text and table subpages. There are 97 text subpages and 11 table subpages; the remaining main category materials include procedural files, approval forms, letters assigning problem clues, minutes of special meetings on clue handling, preliminary verification approval forms, preliminary verification plans, preliminary verification reports, meeting minutes, case filing and investigation approval reports, investigation plans, detention measure filing forms, interview records, interrogation records, notices of rights and obligations of the investigated person, search records, self-criticisms, materials on disciplinary and legal violations, letters of repentance, and cadre resumes, etc. Therefore, this embodiment is not a single document recognition scenario, but a comprehensive case file review scenario encompassing procedural approval, factual materials, interviews and interrogations, notices of rights, detention filing, factual determination, and cadre identity materials.

[0103] The simulated case in this embodiment involves: Personnel A, a section chief in a certain unit, is suspected of serious violations of discipline and law in the verification of a certain subsidy. The case file is compiled around the handling of leads, preliminary verification, case filing and investigation, detention measures, interviews and interrogations, factual materials, and relevant procedural materials. The goal of the system review is not to determine the substantive disciplinary action, but to conduct a quality check on the completeness of the elements, procedural standardization, consistency of expression, field completeness, and cross-material relevance of the case file materials.

[0104] II. Case file loading, batch catalog generation, and image preview: Operators access the case file catalog review page on the system front end, specifying the batch parameter as batch=639. After reading the batch parameter, the front end calls the batch record interface to load the case file catalog for that batch. A catalog list appears on the left side of the page, the middle area displays the original image of the case file corresponding to the current record, and the right area displays the file parsing results, full-text content entry, editing entry, entry to copy or download HTML content, and error location information for that record.

[0105] In this embodiment, the system first organizes 212 records using batch 639 as the unified context. The identification results at the beginning of the directory include: `0001.jpg` is identified as "Program Volume"; `0002.jpg` is identified as "[Approval Form]"; `0003.jpg` is identified as "Letter of Assignment of Problem Clues", etc.

[0106] In the middle and end of the directory, the system also identified `0070.jpg` as "Record of Detention Measures", `0072.jpg` as "Interview Record", and `0076.jpg` as "Interrogation Record of a Supervisory Commission", etc. The above directory results provide a basis for material types in subsequent structured extraction and rule validation.

[0107] When a reviewer clicks on a directory item, the system sets that record as the current record. For example, clicking on the interview transcript `0072.jpg` displays the image of the interview transcript in the center of the page, with the system-parsed interview transcript fields and verification results on the right; clicking on the preliminary verification report `0019.jpg` loads the structured fields of the report and its associated verification results; clicking on the review and investigation plan `0029.jpg` displays the full-text recognition fragment of the plan, field extraction results, and expression issues found in the large-scale model semantic verification. In this way, the system organizes multi-page, multi-type case files into a unified data structure of "batch—directory record—image attachment—parsed content—verification result".

[0108] III. Material Recognition, Full-Text Recognition, and Structured Feature Extraction: After material identification is completed, the system performs full-text recognition and structured element extraction according to different material types. For text subpages and table subpages without configured structured extraction templates or used as auxiliary pages, the system retains their table of contents and image records; for main category materials with configured extraction rules, the system generates structured field data and HTML parsing content. Batch 639 contains 421 structured field data entries, with field types including text fields, long text fields, and JSON fields. Field names cover titles, names, genders, ages, political affiliations, work units and positions, signing units, signing dates, legal citations, notification content, Q&A content, signatures or fingerprint information, receipt time, and number of pages for materials on violations of discipline and laws.

[0109] Taking the "Clues Handling Special Meeting Minutes" file `0006.jpg` as an example, the system extracts the "Title" as "Clues Handling Special Meeting Minutes" and the "Meeting Time" as "2025.5.23", while leaving empty values ​​for fields such as "Organizer" and "Document Number". These types of fields can be used to subsequently determine whether the meeting minutes are complete and whether the time can be corroborated by other procedural materials.

[0110] Taking the "Approval Form - Preliminary Verification Approval Form of a Certain Discipline Inspection and Supervision Commission" in `0009.jpg` as an example, the system extracted 34 fields, mainly including: the handling department is "Fifth Discipline Inspection and Supervision Office", the date is May 23, 2025, the clue code is "20250005", the source of the clue is "assigned by the Discipline Inspection Commission of a certain city", the gender of the person being investigated, Ning, is "female", age is "43", ethnicity is "XX", political affiliation is "XX", work unit and position is "head of a certain management section", the person in charge is "Wei", and the person handling the matter is "Zhao". The system also extracted a summary of the reported problem, which can be summarized as follows: From October to November 2023, Ning failed to perform her verification duties diligently, resulting in relevant dealers fraudulently obtaining certain subsidies and causing losses to the state; in May 2024, she also helped others apply for agricultural machinery under false pretenses and pass the verification, suspected of accepting bribes. This summary is entered into the structured data set as a long text field, which can be used for semantic verification of large models and cross-material verification.

[0111] As can be seen from the above examples, this application does not simply convert images into text, but rather establishes a connection between case file images, full text, fields, HTML parsing content, and subsequent verification objects based on the material type; reviewers can view the original images and parsing content on the page simultaneously, thereby verifying the field extraction results and locating problems.

[0112] IV. Joint Verification Operation and Actual Verification Results: After the operator clicks the one-click check, the system performs joint verification in stages: first, it performs global null value or necessary field type checks; second, it performs computer rule verification; and finally, for items requiring natural language understanding, cross-material understanding, or complex expression judgment, it performs large-scale model semantic verification. Batch 639 currently has 41 verification results, including 17 computer rule verification results and 24 large-scale model semantic verification results.

[0113] The 41 results generated in this embodiment cover multiple dimensions, including missing values, document names, program descriptions, time and location, signatures, title consistency, legal citations, cross-material existence, factual material characterization, duplicate descriptions, and material quantity verification. Computer rules are suitable for handling ruleable issues such as empty fields, fixed descriptions, time, and titles; large-scale model semantic validation is suitable for handling issues that are difficult to fully rule-form, such as natural language descriptions, cross-material associations, legal citation order, long text duplication, and factual characterization. Both types of results are uniformly written into the same result set and can be displayed on the page by record and field.

[0114] V. Feedback on Evaluation Results, Manual Review and Correction Procedures: After verification, the system displays the results on the front-end directory review page. The left-hand directory indicates which materials have verification results; when reviewers click on a material in the directory, the corresponding case file image is displayed in the middle area, and the document parsing content and error messages are displayed on the right. Taking the `0072.jpg` interview transcript as an example, after reviewers click on this directory item, they can view the original interview transcript in the image area, view the interview start time, end time, interview location, interviewee information, and handwritten signature field in the parsing area, and see prompts in the error area such as interview location, interview duration, handwritten signature specifications, confirmation of informal interview location Q&A, and evidence materials from NPC deputies. Reviewers use this information to determine whether the problem actually exists. If they believe the verification result is accurate, they retain the problem; if they find that the extracted fields are inconsistent with the original image text, they can first correct the parsing content or re-identify, and then re-verify.

[0115] After reviewing the data, the reviewers can edit or delete the verification results. For materials with incorrect classification, a re-identification of the file type can be triggered; for materials with incorrect field extraction, re-content recognition or manual field editing can be triggered; for batches with supplementary materials, a one-click check can be re-executed. When the system re-checks, it uses the updated table of contents, full text, structured fields, and HTML parsing results as input to regenerate computer rules and large-scale model semantic verification results for reviewers to confirm again. This process ensures that the review results are not just a one-time output, but form a closed loop around the actual case file modification process: "identifying problems—verifying originals—correcting materials or fields—re-verifying—re-re-reviewing".

[0116] VI. Evaluation Report Generation Process and Current Report Status: After the reviewers confirm the results, the system can generate a PDF review report based on the saved verification results of batch 639. The report generation module reads the set of verification results, corresponding records, material types, field names, error pages, problem descriptions, error content, and suggestions for this batch, and organizes them into a downloadable PDF report.

[0117] In this embodiment, the database query shows that the number of `EXPORT_CHECK_REPORT_LOGS` records for batch 639 is 0, meaning that there are no saved report export attachment logs in the current query state. Therefore, the following text will not describe it as "report downloaded" or "report attachment formed", but rather as: the system has a functional path to generate a PDF review report based on 41 verification results, and the current batch already has result data that can be read by the report generation module; if the operator clicks "generate report" on the page, the system can summarize the verification results from materials such as the preliminary verification approval form, preliminary verification situation report, case filing and investigation approval report, investigation plan, interview transcript, interrogation transcript, notice of rights and obligations, search transcript, letter of self-criticism, materials on facts of violations of discipline and law, letter of repentance, and cadre resume into the report by record and field.

[0118] In terms of report organization, the system can group issues according to the material records. For example, the report may specify that the preliminary verification approval form `0009.jpg` has a blank date for the deputy secretary's approval and missing attachments; the preliminary verification report `0019.jpg` has issues with the order of the signing unit, cited clauses, place of origin, and basic information; the case filing and investigation approval report `0024.jpg` has issues with the title and the order of legal citations; the investigation plan `0029.jpg` has issues with the description of case filing; and the interview transcript `0072.jpg` has issues with the interview location, interview duration, handwritten signature, and confirmation of the location of informal interviews. The following issues were found in the following documents: `0083.jpg` The Notice of Rights and Obligations has issues with the document title and date of signature; `0086.jpg` The search record has issues with its correlation with the search plan; `0115.jpg`, `0117.jpg`, `0145.jpg`, `0147.jpg`, and `0148.jpg` Factual materials have issues with the characterization of violations of discipline and law, the citation of clauses, and the duplication of basic information; `0151.jpg` The Confession Letter has a prompt to verify the quantity of handwritten and copied materials; `0176.jpg` The Cadre Resume has a missing date of receipt or retrieval.

[0119] This embodiment started at 17:27 on January 17, 2026. Uploading 280 pages (Internet environment) took 44.22 seconds, providing data took 11 minutes and 48 seconds, and verification took 2 minutes and 22 seconds.

[0120] In summary, the complete operation process of this embodiment can be summarized as follows: The operator enters the `case / catalog?batch=639` page, and the system loads 212 image-based case records from batch `DJ202600003`; the system generates a catalog according to material type and displays the images; structured fields are extracted from materials such as clue handling, preliminary investigation, case filing, review and investigation, detention, interviews, interrogations, notification of rights, searches, reviews, factual materials, confessions, and cadre resumes, forming 421 field data entries and 4 full-text recognition results; after the operator triggers a one-click check, the system generates 17 computer rule verification results and 24 large-model semantic verification results; the reviewer checks each item on the page in conjunction with the images and analysis content, editing, deleting, re-recognizing, or re-verifying as necessary; finally, the system can generate a PDF review report based on the confirmed results. This specific embodiment illustrates that this application can transform unstructured image materials in comprehensive discipline inspection and supervision case files into evaluation objects that can be displayed, extracted, verified, located, reviewed, and reported, thereby forming an intelligent evaluation process that is feasible and reproducible.

[0121] In summary, the case file quality assessment method provided in this application can achieve the following beneficial effects: 1. Significantly improves review efficiency and reduces labor costs: This application automates the entire process of case file uploading, identification, element extraction, verification, and report generation, eliminating the need for manual page-by-page review, recording, and verification. Combined with lightweight compression and parsing of large case files and dynamic computing power scheduling algorithms, it completely solves the problems of time-consuming and inefficient manual reviews. Actual testing shows that a single 200-page case file takes only 17 minutes (compared to just 3 minutes for manual operation), representing an efficiency improvement of over 95% compared to traditional manual reviews (which take two staff members more than two days).

[0122] 2. Reduce omissions in case review and improve the quality of case file review: This application adopts a two-layer accurate recognition mechanism of OCR + multimodal large model, combined with full-dimensional automated verification (program, element, time sequence, cross-document, semantics, etc.), covering all core points of discipline inspection and supervision case file review, avoiding omissions caused by human review fatigue and subjective judgment bias. Among them, the accuracy rate of element extraction is over 99%, the accuracy rate of handwritten content recognition is over 97%, and the rule verification coverage is 100%. It can accurately identify key issues that are easily overlooked by humans, such as missing fingerprints, reversed time sequence, and contradictory information, effectively avoiding potential case file quality problems and ensuring the accuracy and comprehensiveness of the review results.

[0123] 3. Shorten the review cycle and solve the problem of long-delayed cases: This application significantly reduces the process of multiple rounds of repeated reviews through visual error feedback, one-click re-verification, and convenient manual correction, avoiding the time-consuming back-and-forth of "feedback-modification-re-review" in manual reviews. Traditional manual reviews of a single case file require multiple rounds of re-review, with a cycle of 3-7 days. The review cycle of a single case file (including modification and re-review) can be shortened to 1-2 hours, significantly reducing the case file review cycle, ensuring timely case completion, and improving the efficiency of discipline inspection and supervision work.

[0124] 4. Achieve standardized and regulated review and reduce human error: This application constructs a standardized metadata foundation and rule verification library, abstracts the disciplinary inspection and supervision case file review standards into reusable and configurable rules, and realizes unified configuration and dynamic iteration of rules through a visual editor, adapting to the unified case file standards at the municipal level. This avoids the problems of inconsistent review standards and large human error caused by differences in individual professional abilities and work habits in manual review, and realizes standardized and regulated management of case file review.

[0125] 5. Easy to operate, highly practical, and adaptable to multiple application scenarios: This application adopts a three-dimensional visualization display interface with left, center, and right sides, multi-level error marking, and one-click jump positioning. The operation is simple and easy to understand, and no professional intelligent technology foundation is required for the staff. At the same time, it supports multi-format case file adaptation, dynamic configuration of local large models, and regional adaptation. It can be widely used in various case file review scenarios of discipline inspection and supervision agencies at all levels, and can be extended to any case file review scenario such as public security, judiciary, and market supervision. It has extremely strong practicality and scalability.

[0126] 6. Reduce workload and improve staff performance: This application replaces tedious and repetitive tasks such as manual page-by-page review, element extraction, manual recording, and report writing, freeing staff from heavy manual labor and enabling them to devote more energy to core tasks such as problem rectification and case analysis. This effectively reduces workload and improves the performance and work enthusiasm of discipline inspection and supervision staff.

[0127] In some embodiments of this application, a case file quality assessment system is provided, which corresponds one-to-one with the case file quality assessment methods described in the above embodiments. For example... Figure 3 As shown, the case file quality assessment system includes a batch management module 101, a catalog and record module 102, a content recognition module 103, a structured extraction module 104, a verification module 105, and a report generation module 106. Batch management module 101 is used to receive case files to be reviewed and generate batch identifiers for the case files to be reviewed. The catalog and record module 102 is used to establish a catalog record for the case file materials to be reviewed. The catalog record includes at least one of the following: the order of each material, the material name, the material type information, the attachments, and the image association information; Content recognition module 103 is used to perform content recognition on the case file materials to be reviewed corresponding to each directory record, and obtain the material type of the case file materials to be reviewed; The structured extraction module 104 is used to locate and extract the key content of the case materials to be reviewed according to the material type of the case file to be reviewed and call its corresponding extraction algorithm to obtain the extraction result; convert the extraction result into structured fields and associate them with the corresponding directory records; The verification module 105 is used to perform global null value verification, computer rule verification, and large-scale model natural language verification on case file materials in the same batch to obtain verification results. The global null value verification is used to check whether key materials or key fields are missing. The computer rule verification is used to perform deterministic rule checks on the extracted structured fields. The large-scale model natural language verification is used to perform semantic verification on the case file materials to be reviewed based on the verification task. The verification results include at least batch identifier, record identifier, material classification or material type, field name or field identifier, result source type, problem reason, suggestion explanation, and display status. The report generation module 106 is used to generate an evaluation report based on the verification results.

[0128] In some embodiments of this application, the content recognition module 103 is specifically used to perform OCR or full-text recognition on the case file materials to be reviewed for a preset special material type, forming full-text information associated with and saved in the catalog record; the full-text information is used for large model natural language verification and manual full-text viewing.

[0129] In some embodiments of this application, it is determined whether the catalog record and material type of the case file to be reviewed are correctly identified; if the catalog record and material type of the case file to be reviewed are incorrectly identified, a maintenance operation is performed, the maintenance operation including at least one of the following: re-identifying the catalog, re-identifying the file type, inserting a category or image before the current catalog, replacing the category and image, and deleting the current catalog record.

[0130] In some embodiments of this application, the image and parsing display module is used to backfill the structured fields, identified content or template content into the HTML display content to form a document parsing page for comparison and verification with case file materials.

[0131] In some embodiments of this application, the global null value verification module is used to check at least one of the following: whether key materials are missing in the batch, whether necessary structured fields are empty, and whether core extraction results have not been generated, with the entire case file batch as the execution scope.

[0132] In some embodiments of this application, the computer rule verification module is used to check at least one of the following: field existence verification, logical relationship verification between fields, condition judgment verification, cross-material field comparison verification, and multi-field combination rule verification.

[0133] In some embodiments of this application, the large-scale natural language verification module is used to construct a verification task based on the batch, the material type of the catalog record, the full-text information, the structured fields, and the verification rule description; call the large-scale language model so that the large-scale language model performs semantic judgment only based on the case file material content and rule requirements in the verification task; and convert the output of the large-scale language model into a unified verification result, the unified verification result including problem description, problem cause, suggestion description, and result source type.

[0134] In some embodiments of this application, a re-identification and re-verification closed-loop module is used to display the verification result and receive manual review operations; the manual review operations include at least one of the following: editing or deleting the verification result, and triggering re-identification of the identified content; When re-identification is triggered, global null value verification, computer rule verification, and large model natural language verification are re-executed based on the updated identification content to form a new verification result.

[0135] It should be noted that any of the above-mentioned case file quality assessment systems can implement the aforementioned case file quality assessment methods one by one, which will not be elaborated here.

[0136] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Figure 4 As shown, at the hardware level, this electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or it may include non-volatile memory, such as at least one disk drive. Of course, this electronic device may also include other hardware required for other business operations.

[0137] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0138] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0139] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a case file quality assessment system at the logical level. The processor executes the program stored in memory and specifically performs the aforementioned methods.

[0140] The processor may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0141] This electronic device can execute the case file quality assessment methods provided in several embodiments of this application, and realize a case file quality assessment system in... Figure 3 The functions of the embodiments shown are not described in detail here.

[0142] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform the case file quality assessment method provided in several embodiments of this application.

[0143] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0145] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0146] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0147] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0148] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0149] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0150] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0151] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0152] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for evaluating the quality of case files, characterized in that, The method includes: Receive case files to be reviewed and generate a batch identifier for the case files to be reviewed; A catalog record shall be established for the case file materials to be reviewed. The catalog record shall include at least one of the following: the order of each material, the name of the material, the material type information, the attachments, and the image association information; For each directory record, the material type information feature of the case file materials to be reviewed is matched to obtain the material type of the case file materials to be reviewed. Based on the material type of the case file to be reviewed, the corresponding extraction algorithm is invoked to locate and extract the key content of the case file to be reviewed, and the extraction result is obtained; the extraction result is converted into structured fields and associated with the corresponding directory records; Global null value validation, computer rule validation, and large-scale model natural language validation are performed on the case file materials in the same batch to obtain the validation results. The global null value validation is used to check whether key materials or key fields are missing. The computer rule validation is used to perform logical rule checks on the extracted structured fields. The large-scale model natural language validation is used to perform semantic validation on the case file materials to be reviewed based on the validation task. The validation results include at least batch identifier, record identifier, material classification or material type, field name or field identifier, result source type, problem reason, suggestion explanation, and display status. An evaluation report is generated based on the verification results.

2. The method according to claim 1, characterized in that, The method further includes: For preset special material types, OCR or full-text recognition is performed on the case file materials to be reviewed to form full-text information associated with the catalog records; the full-text information is used for large model natural language verification and manual full-text viewing.

3. The method according to claim 1 or 2, characterized in that, After obtaining the material type of the case file materials to be reviewed, the method further includes: Determine whether the catalog records and material types of the case files to be reviewed are correctly identified; If the catalog record and material type identification of the case file materials to be reviewed are incorrect, a maintenance operation is performed. The maintenance operation includes at least one of the following: re-identifying the catalog, re-identifying the material type, inserting a category or image before the current catalog, replacing the category and image, and deleting the current catalog record.

4. The method according to claim 1, characterized in that, After extracting the structured fields of the case file materials to be reviewed, the following is also included: The structured fields, identified content, or template content are then backfilled into the HTML display content to form a file parsing page, which is used for comparison and verification with case file materials.

5. The method according to claim 1, characterized in that, The global null value check is performed on the entire batch of case files, and the check includes at least one of the following: whether key materials are missing in the batch, whether necessary structured fields are empty, and whether core extraction results have not been generated.

6. The method according to claim 1, characterized in that, The computer rule verification includes at least one of the following: field existence verification, logical relationship verification between fields, condition judgment verification, cross-material field comparison verification, and multi-field combination rule verification.

7. The method according to claim 2, characterized in that, The large-scale model natural language validation specifically includes: A verification task is constructed based on the batch context, the material type of the catalog record, the full-text information, the structured fields, and the verification rule description. The large language model is invoked, and the large language model performs semantic judgments solely based on the content of the case file materials and rule requirements in the verification task; The output of the large language model is transformed into a unified verification result, which includes a problem description, the cause of the problem, a suggestion, and the source type of the result.

8. The method according to claim 1, characterized in that, Before generating an evaluation report based on the verification results, the method further includes: Display the verification results and receive manual review operations; the manual review operations include at least one of the following: editing or deleting the verification results, and triggering the re-identification of the identified content; When re-identification is triggered, global null value verification, computer rule verification, and large model natural language verification are re-executed based on the updated identification content to form a new verification result.

9. A case file quality assessment system, characterized in that, The system includes: The batch management module is used to receive case files to be reviewed and generate batch identifiers for the case files to be reviewed. The catalog and record module is used to create a catalog record for the case file materials to be reviewed. The catalog record includes at least one of the following: the order of each material, the material name, the material type information, the attachments, and the image association information; The content recognition module is used to perform content recognition on the case file materials to be reviewed corresponding to each directory record, and to obtain the material type of the case file materials to be reviewed. The structured extraction module is used to locate and extract the key content of the case materials to be reviewed based on the material type of the case file to be reviewed and call its corresponding extraction algorithm to obtain the extraction results; the extraction results are converted into structured fields and associated with the corresponding directory records; The verification module is used to perform global null value verification, computer rule verification, and large-scale model natural language verification on case file materials in the same batch to obtain verification results. The global null value verification is used to check whether key materials or key fields are missing. The computer rule verification is used to perform deterministic rule checks on the extracted structured fields. The large-scale model natural language verification is used to perform semantic verification on the case file materials to be reviewed based on the verification task. The verification results include at least batch identifier, record identifier, material classification or material type, field name or field identifier, result source type, problem reason, suggestion explanation, and display status. The report generation module is used to generate an evaluation report based on the verification results.

10. An electronic device, comprising: processor; And a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the steps of the case file quality assessment method as described in any one of claims 1-8.