AI-based multi-modal evidence package-based manuscript pre-check and reply method and system

CN122819166APending Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611308474.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]第一,文档解析结果与原稿位置之间缺少稳定绑定

Benefits of technology

[0052](1)通过来源锚点机制,确保模型输出能够精确定位至原文节点,使生成内容具备可追溯、可验证的源头,增强结论的可靠性和透明度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819166A_ABST
    Figure CN122819166A_ABST
Patent Text Reader

Abstract

The application discloses a manuscript pre-checking and rewriting method and system based on an AI multi-modal evidence package, and belongs to the field of artificial intelligence assisted pre-checking. The method comprises the following steps: obtaining a manuscript and generating a manuscript version identifier, parsing the manuscript into document nodes and text blocks, and obtaining source anchor points and confidence degrees; secondly, generating a picture reference object to represent a picture-text reference relationship, constructing an internal multi-modal evidence package in combination with the manuscript version identifier, the source anchor points, the text blocks and the picture reference object, inputting the internal multi-modal evidence package into an AI checking model to generate a plurality of checking results, each of which comprises a state field and an output field, writing the results that pass the checking to an editing and proofreading workpiece, recording the problem-free results that do not hit the artificial review boundary as editing and proofreading passing, and generating visual items and transferring the problem-free results that hit the artificial review boundary and the candidate problems or evidence gaps to artificial review. The method makes the model output results locatable, traceable and reviewable, and effectively avoids storing conclusions without basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of digital publishing and artificial intelligence-assisted pre-screening, specifically relating to a method and system for pre-screening and writing back text and image manuscripts based on AI multimodal evidence packages. Background Technology

[0002] A graphic manuscript refers to a published manuscript that includes at least text content and may contain images, captions, text citations, tables, or formulas. This includes textbook manuscripts, supplementary teaching materials, popular science manuscripts, picture book manuscripts, or other manuscripts with mixed text and graphics. It does not include question banks, courseware, videos, audio, or answer analysis files that exist independently as external supporting resources, nor does it include H5 pages that need to be analyzed separately by page node, media segment, and interactive state.

[0003] When receiving manuscripts with images, publishing houses typically need to process different formats such as DOCX, PDF, and Markdown. These manuscripts contain both textual content such as chapter titles, main text, formulas, and tables, and visual content such as illustrations, figure captions, and text citations. Existing intelligent proofreading solutions can recognize documents or images and provide proofreading suggestions, but they still have the following shortcomings in the pre-screening scenario for manuscripts with images for publication:

[0004] First, there is a lack of stable binding between the document parsing results and the location of the original manuscript. After the model outputs candidate questions, editors often only see a question description and find it difficult to directly locate the chapter, paragraph, figure title, or illustration in the original manuscript.

[0005] Second, there is a lack of a unified evidentiary object among image titles, text citations, and image assets. Existing solutions tend to handle image recognition, text review, and rule verification in a fragmented manner, resulting in a lack of verifiable evidence chains for image-text consistency issues.

[0006] Third, the model output lacks structured verification before being stored in the database. If the output lacks source anchors, verification basis, or manual review markings, the system may still generate review conclusions that cannot be located or reviewed.

[0007] Therefore, a technical solution is needed that can organize document nodes, text blocks, image asset identifiers and image access paths, image titles, text citations, rule versions, and manual review boundaries of graphic manuscripts into verifiable evidence packages, and stably write candidate issues or evidence gaps back to the original text anchor points. Summary of the Invention

[0008] To address existing technical issues such as how to establish cross-format source anchors in the pre-screening of text and image manuscripts; how to organize image asset identifiers, image access paths, image titles, text citations, and rule constraints into a verifiable internal multimodal evidence package; how to block model outputs lacking source anchors, image citation objects, or verification basis through output constraints and structured verification; and how to write back candidate issues or evidence gaps to the corresponding positions in the review workpiece, this invention provides a text and image manuscript pre-screening write-back method and system based on AI multimodal evidence packages.

[0009] In a first aspect, this invention proposes a method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages, including:

[0010] S1, Obtain the text and image manuscript to be reviewed and generate a manuscript version identifier for the text and image manuscript;

[0011] S2, the text and image manuscript is parsed into various document nodes, further resulting in text blocks, and source anchors are generated for each document node and text block; the document nodes include title nodes, paragraph nodes, image nodes, image caption nodes, and body text reference nodes; the text blocks are obtained by merging title nodes and paragraph nodes;

[0012] S3, based on document nodes and text blocks, generates an image reference object to represent the reference relationship between the image and the manuscript text;

[0013] S4. Based on the manuscript version identifier, source anchor, text segmentation, and image reference objects, an internal multimodal evidence package is generated; the internal multimodal evidence package is used to limit the processing scope of the AI ​​verification model.

[0014] S5, input the internal multimodal evidence package into the AI ​​verification model to generate multiple verification results containing state fields and output fields, and perform basic field verification on each verification result; the state fields are no problem, candidate problem, or evidence gap;

[0015] For verification results that pass the basic field verification and whose status field is "no problem" or "candidate problem", proceed directly to S6; for verification results that fail the basic field verification, or verification results that pass the basic field verification but whose status field is "evidence gap", further evidence gap-specific verification will be performed, and if it passes, proceed to S6.

[0016] S6. Write the verification results back to the review workbook. For verification results with a status field of "no problem" and that do not hit the boundary of manual review, no manual review is required. For verification results with a status field of "no problem" and that hit the boundary of manual review, as well as verification results with a status field of "candidate problem" or "evidence gap", generate visual entries in the review workbook and perform manual review.

[0017] Furthermore, in S2, the process of generating text blocks specifically involves:

[0018] Based on the position of each document node in the text and image document, the chapter path of that document node is formed and normalized to obtain the normalized value of the chapter path;

[0019] Within a preset length threshold, adjacent title nodes or adjacent paragraph nodes located in the same chapter path are merged into a text block; at the same time, paragraph nodes exceeding the preset length threshold are split to obtain multiple text blocks.

[0020] Furthermore, in S2, the source anchor also includes the source anchor confidence score, which is obtained by calculating the completeness of the chapter path, the effective length of the original text fragment, the matching degree of the preceding and following text, and the uniqueness of the node of the document node or text block corresponding to the source anchor, and then summing them by weight.

[0021] Chapter path completeness is a binary criterion for determining whether the document node or text block corresponding to the source anchor point can be traced back to at least one chapter title in the graphic manuscript, and the value is assigned based on the determination result.

[0022] The effective length of the original text fragment is determined by comparing the fragment length with the preset effective range, minimum positioning length, and maximum positioning length of the source anchor point of the corresponding text block or text reference node; for the source anchor points of other document types, a fixed value is directly used.

[0023] The context matching score is calculated by comparing the prefix text of the source anchor record with the previous document node or text block, and comparing the suffix text of the source anchor record with the next document node or text block; the score is determined based on the consistency comparison results of the prefix and suffix text.

[0024] Node uniqueness is determined by counting the total number of times the text block corresponding to the source anchor appears in the text and image manuscript, and then taking the reciprocal of the total number of occurrences.

[0025] Furthermore, S3 specifically refers to:

[0026] S301, Generate an image asset identifier for each image; the image asset identifier is obtained by concatenating the manuscript version identifier, image sequence number and image binary verification value in a fixed field order, and then performing a hash operation;

[0027] S302, based on the text and image manuscript, the image title node and the text reference node are respectively associated and bound to the image asset identifier of the corresponding image, and an image reference object is further generated; the image reference object includes the image reference object identifier, image asset identifier, image alternative text, image access path, identifier of the text block adjacent to the image, source anchor point of the corresponding image node, associated and bound image title node identifier, associated and bound text reference node identifier, and matching status; the image alternative text is a text description of the image content, and the matching status is used to describe whether the image title node or the text reference node matches the image.

[0028] Furthermore, in S4, the internal multimodal evidence package also includes model invocation context, rule version snapshot, subject plugin identifier, manuscript stage, manual review boundary, and output schema version;

[0029] The model invocation context is used to store the attribute information, running information, and output information of the AI ​​verification model;

[0030] The rule version snapshot is used to define the verification rules that the AI ​​verification model follows when performing verification.

[0031] The subject plugin identifier is used to indicate the subject to which the text and image manuscript belongs, and loads the corresponding subject-specific verification logic, terminology and image recognition dimensions;

[0032] The term "manuscript stage" indicates the current publication stage of a text and image manuscript.

[0033] The manual review boundary is used to mark content in graphic manuscripts related to knowledge accuracy, copyright, national standards, industry standard clauses, safe operating procedures, and political orientation.

[0034] The output schema version is used to define the state fields output by the AI ​​validation model and the output fields that each state field must include.

[0035] Furthermore, in S5, the AI ​​verification model is a large language model, a multimodal large model, or a visual language model.

[0036] Furthermore, in S5, the basic field validation includes output status validity validation, source anchor index retrospective validation, reference object consistency validation, validation basis integrity validation, and manual review mark integrity validation.

[0037] The output status validity check requires that the status field of the check result must be either "No Problem", "Candidate Problem", or "Evidence Gap".

[0038] The source anchor index can be used to look up the source anchor referenced in the output field of the validation result, and the corresponding record must be retrieved from all source anchors;

[0039] The consistency check of referenced objects requires that when the output field of the check result involves references to images, captions, or text, the corresponding image reference object must be referenced. At the same time, the image reference object must belong to the current text and image manuscript, and the source anchor point in the image reference object must be consistent with the source anchor point in the output field or within the allowed range of adjacent nodes.

[0040] The verification basis integrity check requires that the verification basis be provided in the output fields of the verification result.

[0041] The manual review mark integrity check output field must include the manual review mark field, and the value of this field must be a valid boolean value.

[0042] Furthermore, the source anchor index retrospective verification also includes downgraded location processing, the specific process of which is as follows:

[0043] The source anchors referenced in the output fields of the verification results are used as candidate source anchors. The matching values ​​of each source anchor in the text and image manuscript are calculated and weighted by fuzzy matching from four dimensions: chapter path, original text fragment, prefix text and suffix text. The result is obtained as the restored location reliability of each source anchor in the text and image manuscript.

[0044] If there is a source anchor with a recovery location confidence score higher than the threshold, the candidate source anchor is replaced with the source anchor with the highest recovery location confidence score, and the basic field validation is performed again.

[0045] If the recovery of location reliability is below the threshold, but one or more features in the four dimensions of chapter path, original text fragment, prefix text or suffix text have a matching value higher than the corresponding threshold, then a location evidence gap is generated.

[0046] If the reliability of the restored location is lower than the threshold, and the matching values ​​of all source anchors' chapter paths, original text fragments, prefix texts, and suffix texts are all less than or equal to the corresponding threshold, then the verification result will be prevented from being written back to the review workpiece.

[0047] Furthermore, in S5, the process of performing basic field validation on the output result is as follows:

[0048] For verification results where the status field is "no problem" or "candidate problem", if the output result passes the basic field verification, it will directly enter S6; if any verification in the basic field verification fails, it will be converted into an evidence gap and the field will be filled. If the filling is successful, it will enter the evidence gap special verification; if it cannot be filled, it will directly block the entry into the database.

[0049] For verification results where the status field is an evidence gap, if the output result passes the basic field verification, it will proceed to the evidence gap-specific verification; if any verification fails, the field will be completed. If the completion is successful, it will proceed to the evidence gap-specific verification; if the completion cannot be successful, the data entry will be blocked directly.

[0050] Secondly, this invention proposes a text and image manuscript pre-inspection and rewrite system based on AI multimodal evidence packages, which is used to implement the above-mentioned text and image manuscript pre-inspection and rewrite method based on AI multimodal evidence packages.

[0051] The beneficial effects of this invention are:

[0052] (1) By using the source anchoring mechanism, the model output can be accurately located to the original text node, so that the generated content has a traceable and verifiable source, and the reliability and transparency of the conclusions are enhanced.

[0053] (2) Introducing the image reference object and internal multimodal evidence package, so that all candidate issues related to the image and text are accompanied by verifiable evidence, which makes it easier for reviewers to check the conclusions related to the image and text item by item.

[0054] (3) By using evidence package hash and model call context, the model input, output, structured verification record and write-back status are uniformly mapped to the same evidence scope to achieve full-link data closure and avoid cross-link information fragmentation.

[0055] (4) By using output constraints and structured verification, outputs lacking valid source anchors are downgraded or blocked from entering the database, thus preventing unfounded conclusions from entering the system and ensuring the quality of the final output.

[0056] (5) Based on partial matching features, location evidence gaps are formed, and outputs that do not meet the requirements of mandatory fields and consistency are transferred to evidence gap processing to ensure that the system has a clear path for handling incomplete information, rather than forcibly generating conclusions. Attached Figure Description

[0057] Figure 1 This is the overall processing flowchart.

[0058] Figure 2 Flowchart for input parsing and structured modeling.

[0059] Figure 3 This is a flowchart for constraint processing and state verification.

[0060] Figure 4 This is a flowchart for anomaly location and feedback output. Detailed Implementation

[0061] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0062] This invention proposes an AI-based multimodal evidence package-based pre-detection and write-back method for text and image manuscripts. This method addresses the text, images, figure titles, text citations, tables, and formulas within the same version of a text and image manuscript. It parses the manuscript into a unified semantic document structure, generating source anchors and image citation objects. These objects, along with rule versions, manuscript stages, manual review boundaries, and output constraints, are encapsulated into an internal multimodal evidence package. A large language model, multimodal large model, or visual language model is invoked to output "no problems," "candidate problems," or "evidence gaps." The model output is then validated against anchors, image citation objects, verification criteria, and review markers. The validated results are then written back to the review workpiece. This method performs downgraded localization during manuscript version updates, recording the validation and write-back status, enabling the localization, review, blocking, and tracing of candidate problems and evidence gaps.

[0063] like Figure 1-4 As shown, the specific steps are as follows:

[0064] Step 1: Receive the text and image manuscripts, generate a manuscript version identifier, and establish a manuscript version asset corresponding to the manuscript version identifier.

[0065] In this step, the system receives text and image documents in any format, such as DOCX, PDF, or Markdown, automatically generates a globally unique document version identifier (e.g., V20260619-01), and simultaneously creates a corresponding document version asset library, storing the document's filename, source format, checksum, uploader, document version identifier, and original asset access address.

[0066] Among them, the filename represents the original name of the file when the user uploads it; the source format represents the file format type, including DOCX, PDF, or Markdown; the checksum represents the document's hash value; the uploader is used to indicate the operator who uploaded the text and image manuscript; the manuscript version identifier is generated by the system and is used to identify the current version of the manuscript; and the original asset address represents the actual storage location path of the file on the system server.

[0067] For external supporting resource files such as question banks, courseware, videos, audio, or answer explanations, generate external citation records or independent resource pre-inspection task identifiers, and save them in association with the manuscript version assets, but do not incorporate them into the multimodal evidence package within this text and image manuscript.

[0068] At the same time, the system records the subject classification and publication stage (such as preliminary review, secondary review, and final review) of the manuscript when it is received, for use in the subsequent step five when constructing the internal multimodal evidence package.

[0069] Step 2: Parse the text and image manuscript into a unified semantic document structure according to the source format.

[0070] The overall process is as follows Figure 2 As shown. Specifically:

[0071] Step 2.1: Parse DOCX or MD

[0072] In this step, for DOCX format text and image documents, the DOCX file is first unpacked, each paragraph in the document is traversed, and the titles, paragraphs, tables, formulas, figure captions, and embedded images in the document are identified; for Markdown documents, the titles in the Markdown source code, as well as the corresponding title levels, paragraphs, image syntax, and body references are directly parsed.

[0073] Step 2.2: Convert PDF to MD

[0074] For PDF document documents, first save the document version asset and generate the access address, then convert the PDF to Markdown with image positioning marks, and then parse the title, paragraph, image, figure title and text reference from the Markdown with image positioning marks;

[0075] Step 2.3: Identify the title image and text citations

[0076] For each identified object in steps 2.1 and 2.2, generate a document node identifier and node sequence number, and identify reference expressions such as Figure X, Table X, and Formula X in the main text of the manuscript. Extract the figure number, table number, and formula number from them, and bind the corresponding figure title, image, table, and formula nodes. At the same time, mark the paragraph node where the reference expression is located as a reference node in the main text.

[0077] Ultimately, the resulting document nodes include title nodes, paragraph nodes, table nodes, formula nodes, figure caption nodes, image nodes, and text reference nodes. Each node records a document node identifier, node type, node sequence number, chapter path, original location (i.e., the location information of the document node in the original manuscript), and node content or asset reference. Specifically, the node content records the text content of plain text nodes such as titles, paragraphs, and figure captions; the asset reference records the reference identifier or storage path pointing to the corresponding asset (image file, formula image, etc.) for non-text nodes such as images and formulas.

[0078] Step 2.4: Generate chapter paths and text blocks

[0079] First, a chapter path is formed based on the title stack where each identified object is located. The title stack refers to the complete title hierarchy chain from the root title (such as Chapter 3) to the section where the current node is located (such as Section 3.2.1), for example, Chapter 3 / Section 3.2 / Section 3.2.1.

[0080] Secondly, based on the generated chapter paths, the main text document nodes (including title nodes and paragraph nodes) are merged into text blocks according to the chapter path, node order, document node type, and preset length threshold. Alternatively, paragraph nodes that exceed the preset length threshold are split into multiple text blocks.

[0081] When dividing the text into blocks, the system first normalizes the chapter paths, removing consecutive whitespace and format control characters while preserving the heading hierarchy order. The resulting normalized string of chapter paths is used as the normalized value of the chapter paths.

[0082] Each text block record includes the set of document node identifiers to be merged, the start and end node sequence numbers, the block number, the original text fragment, the prefix text (i.e., the original text fragment of the block preceding it), and the suffix text (the original text fragment of the block following it). Furthermore, the identifier for each text block is obtained by concatenating the manuscript version identifier, the chapter path normalization value, the starting document node identifier, the block number, and the block content hash in a fixed order. The block content hash is obtained by performing a hash operation on the original text fragment of the text block and taking a predetermined number of bits from the hash value.

[0083] Step 2.5: Output structured data packets.

[0084] At this point, a complete unified semantic document structure is obtained, including a complete set of document nodes, a set of text chunks, and a set of original image assets, thus completing the construction of the structured input package.

[0085] Step 3: Source Anchor Generation and Index Building

[0086] This step generates source anchors for all document nodes and text chunks, and establishes a source anchor index for the current manuscript version. Each generated source anchor includes at least the manuscript version identifier, a semantic document identifier for identifying the unified semantic document, a document node identifier, a chunk identifier, a chapter path, a chunk type, a chunk number, the original text fragment, the prefix text, the suffix text, the anchor hash, and the anchor confidence.

[0087] The specific operation is divided into the following three steps.

[0088] Step 3.1: Perform anchor hash calculation

[0089] First, the chapter path of each node or block is normalized by removing consecutive whitespace and format control characters, while preserving the heading hierarchy order in the chapter titles.

[0090] Secondly, the original text fragments from the text blocks are truncated and normalized to form standard fragments for computation. Document nodes that did not participate in text block segmentation are left unprocessed.

[0091] Subsequently, the manuscript version identifier, semantic document identifier, document node identifier, chunk identifier, chapter path, chunk type, chunk number, standard fragment, prefix text, and suffix text are concatenated into the anchor text string in a fixed order; for document nodes that did not participate in the text chunking operation in step two, the chunking fields in their source anchors are directly left empty.

[0092] Finally, a hash operation is performed on the anchor string, and the predetermined number of bits in the hash value is used as the anchor hash.

[0093] Step 3.2: Anchor point confidence weighted calculation

[0094] Anchor point confidence is used to determine whether a source anchor point can directly participate in candidate question generation. It is calculated by weighting the following indicators: chapter path completeness, effective length of the original text segment, contextual matching degree, and node uniqueness. The calculation can be expressed as: 0.30 × chapter path completeness + 0.25 × effective length of the original text segment + 0.25 × contextual matching degree + 0.20 × node uniqueness.

[0095] If the confidence level of the anchor point is lower than the confidence threshold, then the source anchor point is prohibited from directly supporting the generation of a deterministic candidate problem in step six, and automatically enters the evidence gap special verification process in step seven, generating or converting it into a system standard evidence gap. After passing the evidence gap special verification, it enters the supplementary evidence flow. In a specific embodiment of the present invention, the source anchor point confidence threshold is 0.70.

[0096] The chapter path completeness is used to indicate whether the current document node or text block can be traced back to at least one chapter title. If it can be traced back to at least one chapter title, the value is 1; otherwise, the value is 0.

[0097] The effective length of the original text fragment is a length index value between 0 and 1. When the standard fragment length is within the preset effective range, the value is 1. When the standard fragment is empty (e.g., the image node does not have a standard fragment), less than the minimum positioning length, or greater than the maximum positioning length, the value is 0. When the standard fragment length is in the transition range between the effective range boundary and the minimum / maximum positioning length, it is linearly converted to a value between 0 and 1 based on the distance from the effective range boundary. In one example configuration, the effective length of the original text fragment is 1 when the standard fragment length is 30 to 300 characters, 0 when it is less than 15 characters or more than 500 characters, and converted proportionally based on the distance from the effective range boundary for 15 to 29 characters and 301 to 500 characters.

[0098] The context matching score, ranging from 0 to 1, indicates the degree of matching between the prefix and suffix text within the unified semantic document structure. A context matching score of 1.0 is achieved if the prefix text of the current anchor record matches the original text fragment of the preceding node or text block, and the suffix text matches the original text fragment of the following node or text block. A score of 0.5 is achieved if only one matches, and a score of 0 is achieved if neither matches or if the preceding or following node is missing.

[0099] The node uniqueness index, ranging from 0 to 1, indicates the degree of repetition of the same standard fragment in the current manuscript version. Its value is 1 divided by the number of times the standard fragment appears in the entire manuscript. For non-textual document nodes such as image nodes, node uniqueness is not calculated.

[0100] Step 3.3: Create source anchor index

[0101] Finally, all source anchors are stored in the current manuscript version anchor index. This index uses fields such as manuscript version identifier, block identifier, document node identifier, and anchor hash as search keys, and supports quick location of the corresponding complete anchor record by anchor identifier, which serves as the sole location basis for subsequent AI output location, verification, and write-back.

[0102] In subsequent steps, any candidate questions or evidence gaps output by the model must be traced back to their source anchors through this index to confirm that the referenced locations actually exist and can be located in the current manuscript version.

[0103] Step 4: Standardized Construction of Multimodal Reference Objects

[0104] This step generates three types of structured reference objects—images, tables, and formulas—based on image nodes, table nodes, and formula nodes, respectively. It associates and binds non-text document nodes with corresponding source anchors, figure titles, table titles, text references, chapter paths, etc., establishing structured evidentiary relationships between images and text, tables and text, and formulas and text.

[0105] (1) Generate image reference object

[0106] The construction of an image reference object involves three steps: generating image asset identifiers, associating and binding image titles and text references, and recording the field records of the image reference object.

[0107] First, the system generates a globally unique image asset identifier for each image. The image asset identifier is constructed from the manuscript version identifier, the image sequence number, and the image binary check value in a fixed field order.

[0108] Specifically, after performing a hash operation on the binary data of the image file to obtain the image binary verification value, the manuscript version identifier of the current manuscript, the image sequence number of the image in the manuscript, and the image binary verification value are respectively cleaned of blanks, unified in case, and invisible characters are deleted. Then, the corresponding processed contents are concatenated into the image asset original string in the order of manuscript version identifier, image sequence number, and image binary verification value. A hash operation is performed on the original string, and the first 8 to 16 bits of the hash value are taken and combined with the image sequence number to form the image asset identifier.

[0109] If the image binary verification value is temporarily unavailable, the system generates a temporary image asset identifier using the normalized image storage address or image access path as a substitute field, and records the identifier's status as pending verification. The official image asset identifier is recalculated after the image binary verification value is obtained. After the image asset identifier is generated, the system also records the image access path, the image substitute text describing the image content, and the storage address.

[0110] Secondly, the system associates and binds the figure title nodes and text citation nodes identified in step two with the image asset identifier. For figure title nodes, the system identifies the figure number and title in the adjacent text, and records the document node and chapter path where the figure title is located. For text citation nodes, the system identifies the citation expressions in the manuscript text and extracts the figure number, the chapter path where the citation is located, and the block identifier where the citation is located.

[0111] After binding is completed, the system calculates the matching status of the figure title node, the text reference node, and the image based on the consistency of the figure number, the distance of the chapter path, the distance of the node order, and the confidence of the source anchor. If the figure number is consistent and the distance of the chapter path, the distance of the node order, and the confidence of the source anchor meet the preset conditions, the matching is successful; if any condition is not met, the figure title node and the text reference node are marked as pending confirmation and will be converted into evidence gaps in subsequent steps.

[0112] Finally, the image reference object identifier, image asset identifier, image alternative text, image access path, text block identifier adjacent to the image, associated source anchor, image title node identifier, text reference node identifier, and matching status are organized into an image reference object, and the image title node identifier and text reference node identifier are associated with this image reference object. The image reference object identifier is obtained by concatenating and hashing the manuscript version identifier, image asset identifier, image title node identifier, text reference node identifier, and source anchor in a fixed order; the associated source anchor is the source anchor corresponding to the image node of this image.

[0113] (2) Generate table reference objects

[0114] The system extracts table nodes from the unified semantic document structure and records the table node identifier, table title, and chapter path of each table node. Simultaneously, from the text reference nodes identified in step two, the system filters out nodes that reference tables and extracts their table number, chapter path, and block identifier.

[0115] The system determines the matching status between table nodes and text reference nodes based on table number consistency, chapter path distance, and node order distance. When the table number matches and the chapter path and node order meet preset conditions, a definite association is established; otherwise, it is marked as pending confirmation and converted into an evidence gap in subsequent steps.

[0116] Ultimately, the record content of the table reference object includes the table node identifier, table title, path of the chapter where the table is located, text reference node identifier, and associated source anchor point, where the associated source anchor point points to the source anchor point generated in step three for the table node.

[0117] (3) Generate formula reference objects

[0118] The system extracts formula nodes from the unified semantic document structure and records their formula node identifier, formula number, chapter path, and formula carrier type. The formula carrier type indicates whether the formula is stored as an image or as machine-readable text.

[0119] Meanwhile, the system filters out the nodes that reference formulas from the text reference nodes identified in step two, and extracts their formula number, the document node identifier where the reference is located, and the block identifier where the reference is located.

[0120] The system determines the matching status between formula nodes and text reference nodes based on formula number consistency, chapter path distance, and node order distance. When formula numbers match and both chapter path distance and node order distance meet preset conditions, a definite association is established, and the matching status is marked as matched; otherwise, it is marked as pending confirmation, and the process proceeds to the evidence gap process.

[0121] Finally, the formula reference object records the formula node identifier, formula number, chapter path where the formula is located, text reference node identifier, formula carrier type, associated source anchor point, and matching status, where the associated source anchor point points to the source anchor point generated in step three for this formula node.

[0122] Step 5: Construction of Internal Multimodal Evidence Package

[0123] This step, based on all the structured data obtained in steps two through four, assembles rules and constraints, encapsulates the model calling context, and generates the internal multimodal evidence package for the manuscript.

[0124] The system first retrieves rule version information (including rule identifiers and version numbers) from the preset rule base based on the subject classification and publication stage recorded in step one, matching the manuscript type, subject, and publication stage. This information is used to create a rule version snapshot. Next, it loads the subject verification plugin (such as a plugin for mathematics, physics, mechanics of materials, or engineering drawing) corresponding to the text and image manuscript. This plugin contains subject-specific verification logic, terminology, and image recognition dimensions, and forms a subject plugin identifier. Simultaneously, it presets boundaries for manual review. Even if the model output passes the structured verification in step six, it marks the manuscript as requiring manual review and retains the manual review status in the review workpiece if the correctness of the manuscript's knowledge conclusions, the use of copyrighted images or text, citations of national or industry standard clauses, safety operating procedures, or politically oriented statements are involved.

[0125] Secondly, constructing the evidence package specifically includes the following steps:

[0126] Step 5.1: Loading basic data

[0127] The system writes all the text blocks and source anchor point sets generated in step two, as well as the image reference objects, table reference objects, and formula reference objects generated in step four, into the evidence package, which serves as the input data range for model calls.

[0128] Step 5.2: Assembling Rules and Constraints

[0129] The matching rule version snapshot, subject plugin identifier, manuscript stage, and manual review boundary are entered into the evidence package, and the rule identifier, rule version, and matching conditions are recorded in the evidence package. For standard clauses cited in the manuscript, the system records the standard clause name, clause number, clause version, document node identifier where the clause is located, and review status. Before proceeding to step eight (candidate issue write-back), it is necessary to check whether the clause version is within the allowed range of the current rule version. If the clause version is missing or a matching record cannot be found in the rule base, the model output is converted into an evidence gap.

[0130] Simultaneously, configure the output schema version. This schema defines that the validation model can only output one of three states: no problem, candidate problem, or evidence gap. Outputs outside this range will be directly intercepted in subsequent validation stages. The required fields for each state are as follows:

[0131] (1) When the output status is no problem, it should at least include the source anchor reference, inspection scope, verification basis, output status and manual review mark, and carry the reference of the corresponding reference object when it involves the reference of pictures, figure titles, tables or formulas;

[0132] (2) When the output status is a candidate problem, it should at least include the problem type, problem description, source anchor reference, verification basis, rule basis field used to indicate which rule was used to determine the existence of the candidate problem, output status and manual review mark, and carry the reference of the corresponding reference object when it involves references to pictures, figure titles, tables or formulas;

[0133] (3) When the model should check a certain content, but cannot give a definite conclusion due to the lack of necessary information, it outputs an evidence gap. When the output status is evidence gap, it should at least include the gap type indicating the missing evidence, the missing evidence, the source anchor reference indicating the location of the gap, the suggested supplementary evidence, the output status, and the manual review mark; when images are involved, it should include image references, when tables are involved, it should include table references, and when formulas are involved, it should include formula references.

[0134] Step 5.3: Encapsulate the model invocation context

[0135] The system reserves fields in the evidence package to store the supplier, model name, model version, prompt word version, evidence package hash, input hash, output hash, prompt word hash, response hash, time consumption, and error code of the verification model. These fields are filled in sequentially before and after the model call to ensure a one-to-one correspondence between the backend model call records and the internal multimodal evidence package, model output, and structured verification records, enabling end-to-end audit traceability.

[0136] Among them, the input hash, output hash, prompt word hash, and response hash correspond to the hash values ​​of the input content, output content, prompt word, and complete response context of the validation model call, respectively.

[0137] Step 5.4: Evidence Packet Hash Generation

[0138] The system standardizes and concatenates the following data in a fixed order: manuscript version identifier, text block identifier set, source anchor point identifier set, image reference object identifier set, table reference object identifier set, formula reference object identifier set, rule version identifier, and output schema version. A hash operation is then performed on the concatenated string to obtain the evidence package hash, which serves as the unique identifier for that evidence package. If core fields such as source anchor points, image reference objects, or rule versions are missing, the system marks the evidence package as incomplete, preventing it from proceeding to subsequent model calls and directly triggering the evidence gap process or manual verification process.

[0139] Step Six: Batch Pre-detection Inference of Multimodal AI Models

[0140] The complete internal multimodal evidence package generated in step five is input into the validation model. This model can be a large language model, a multimodal large model, or a visual language model. The model's output is strictly constrained by the output schema, allowing only three states: no problem, candidate problem, and evidence gap. The model validates each text block, image reference object, table reference object, and formula reference object in the structured output package according to the rule version and subject plugin.

[0141] Before calling, record the supplier, model name, model version, prompt word version, evidence package hash, prompt word hash, and input hash of the verification model. After returning the result, record the output hash, response hash, time consumption, and error code for subsequent verification failure tracing.

[0142] Step 7: Structured Validation and Anomaly Handling of Model Output

[0143] This step is the core verification process of this invention, used to perform two layers of verification on the model output: basic field verification and evidence gap verification.

[0144] Step 7.1: Basic Field Validation

[0145] After receiving the model output, the system first performs basic field validation. The validation dimensions include the following five aspects:

[0146] (1) Output status validity check

[0147] Check if the state field output by the model is one of three: no problem, candidate problem, or evidence gap. If the state does not belong to any of the above three or the state field is missing, the basic field validation is directly deemed to have failed.

[0148] (2) The source anchor index can be checked back.

[0149] The system checks whether the source anchor references carried in the model output can be retrieved in the source anchor index of the current manuscript version. Based on the manuscript version identifier, block identifier, and anchor hash provided in the model output, the system performs an exact match in the index. If the match is successful, the process proceeds; if the match fails, the process proceeds to the downgraded location processing in step 7.2.

[0150] (3) Consistency check of image / table / formula reference objects

[0151] The check verifies whether the model output includes a reference to the image reference object when it involves images, figure captions, or text citations, and whether the reference belongs to the current manuscript version and whether its associated source anchor is consistent with the source anchor in the model output or within the allowed range of adjacent nodes. If an image is involved but no image reference object is included, or the referenced object is inconsistent with the source anchor, the validation fails.

[0152] The validation logic for referenced objects in tables and formulas is the same.

[0153] (4) Verification of the integrity of the basis

[0154] Check whether the model output provides sufficient validation evidence. A "no problem" status must include the scope of the check and validation evidence; a "candidate problem" status must include a problem description and validation evidence; and an "evidence gap" status must include the gap type, missing evidence, and suggested supplementary evidence. If any required field is missing, the validation will fail.

[0155] (5) Manual verification of the integrity of the marking

[0156] Check if the model output includes a manual review marker field, and whether the value of this field is a valid boolean value (True / False).

[0157] Finally, for output statuses of no problem or candidate problem, if all verifications pass, proceed directly to step eight (write-back); if any verification fails, generate an evidence gap and attempt to fill in the field. If it can be filled in, convert it into a system standard evidence gap and proceed to step 7.3; if it cannot be filled in, directly block the data entry.

[0158] If the output status is "evidence gap", proceed to step 7.3 if all verifications pass, and attempt to complete the field if any verification fails. If the field can be completed, it will be converted to a system standard evidence gap and proceed to step 7.3. If the field cannot be completed, the data entry will be blocked directly.

[0159] Step 7.2: Anchor point anomaly downgrade positioning handling

[0160] Based on the chapter path, original text fragment, prefix text, and suffix text of the source anchor point output by the model, the system performs fuzzy matching in the current manuscript, and weights and sums the matching results of each dimension to obtain the restored location reliability of each source anchor point in the index (the threshold is preferably 0.75, which can be adjusted by the rule version or subject plugin according to the manuscript format and publication stage).

[0161] When there is a source anchor point whose recovery location confidence is higher than the threshold, the system will take the source anchor point with the highest matching degree as the current version's retrievable anchor point, write it into the verification record, mark it as needing manual review, and return to step 7.1 for re-verification based on the source anchor point.

[0162] When the location confidence is lower than the threshold, but the matching value calculated by at least one feature in the chapter path, original text fragment, prefix text, or suffix text is higher than the threshold, a location evidence gap is formed. The location evidence gap that passes the consistency check of the cited object, the integrity check of the verification basis, and the integrity check of the manual review mark will proceed to step 7.3. The location evidence gap that fails the check will be filled in with fields.

[0163] When the recovery location reliability is below the threshold and there are no partial matching features, or when the document parsing result of the current manuscript version lacks key structures such as heading hierarchy and block order and cannot provide a context for manual location, the system determines that the associated source anchor point cannot be located at any position in the current manuscript, that is, the model outputs illusory content with no verifiable source, and then directly blocks the entry into the database. The system records the failure reason field, model version, input hash, output hash and error code.

[0164] Step 7.3: System Standard Evidence Gap Generation and Special Verification

[0165] This step is used to perform specific verification on all evidence gaps. Evidence gaps fall into three categories: evidence gaps verified through basic fields, location evidence gaps caused by the inability to recover the source anchor point due to downgraded positioning, and evidence gaps generated by the system.

[0166] The evidence gaps generated by the system are divided into three categories: those triggered by abnormal prior data, those with no issues or candidate issues that fail basic field validation, and those where the model output shows an evidence gap status but lacks the required fields for the evidence gap.

[0167] (1) Generation triggered by abnormal preceding data

[0168] When the system detects situations such as inaccessible images, unrecognizable image text, uncertain image title attribution, missing text citations, abnormal image asset verification values, inaccessible formula images, unrecognizable formula text, text citations that cannot uniquely correspond to formula nodes, or table nodes that cannot uniquely correspond to text citations, it automatically generates gap types, missing evidence, and suggested supplementary evidence based on the missing field type or anomaly type, forming the system's standard evidence gap.

[0169] (2) No problem or candidate problem if the basic field validation fails.

[0170] The system fills in the required fields to form a standard evidence gap; if it cannot fill in the gap, it directly blocks the entry into the database.

[0171] (3) The model output shows an evidence gap state but lacks required fields.

[0172] The system generates or converts these gaps into standard system evidence gaps. The gap type is specifically set according to the missing evidence. For example, if there is a lack of image references, the gap type is set to "Image Evidence Missing"; if there is a lack of verification evidence, the gap type is set to "Verification Evidence Missing".

[0173] Before proceeding to step eight (writing back), all evidence gaps must undergo a special verification process, including: checking whether the gap contains valid source anchors, old source anchors, or anchors that can be retrieved in the current version; if images are involved, checking whether they contain image asset identifiers or image reference objects; if tables or formulas are involved, checking whether they contain corresponding reference objects; whether the gap type is an allowed gap type enumeration value; and whether missing evidence and suggested supplementary evidence have been filled in the output results.

[0174] If the special verification of evidence gaps fails, the output result of the verification model will be blocked from entering the database and the write-back will be prevented. Only evidence gaps that pass the special verification of evidence gaps will enter the write-back process in step eight.

[0175] Step 8: Write back categorized results and complete lifecycle state transition

[0176] This step is used to write back the results of the structured verification in step seven to the review workpiece and maintain the full lifecycle flow status of the results. The review workpiece includes, but is not limited to, structured documents such as HTML, XML, and JSON, or cloud storage objects of online collaborative editors. As long as it can save document node identifiers, source anchors, write-back status, and manual review status, it can be used as the review workpiece in this step.

[0177] The write-back operation essentially stores the issues identified by the validation model in the work interface, allowing reviewers to see these issues in the corresponding locations in the original manuscript and to confirm or reject them. Based on the type of validation result, there are three paths: write-back for no issues, write-back for candidate issues, and write-back for evidence gaps. Each write-back record retains both the validation status and the write-back status (including pending review, confirmed, rejected, and transferred to expert review), ensuring full traceability from model output to manual processing.

[0178] (1) No problem results are written back

[0179] For results that pass the basic field validation and have an output status of "no problem," the system further determines whether they meet the human review boundary. For results that do not meet the human review boundary, no manual processing is required; they are directly archived as approved records. The system records them as approved and writes back the source anchor, inspection scope, and verification basis to the corresponding document node in the review workpiece. For results that meet the human review boundary, a visual entry is generated in the corresponding position of the review workpiece, and a record awaiting human review is generated.

[0180] (2) Candidate question writing

[0181] For results that pass the basic field validation and are output as candidate questions, the system associates the unique identifier of the candidate question with the source anchor identifier or the current version's traceable anchor, which points to the specific location of the question in the original manuscript (which chapter, which section, which paragraph, which document node).

[0182] When candidate questions involve images, figure captions, or text references, the system associates the unique identifier of the candidate question with the image reference object identifier. The image reference object identifier points to the complete information package of the image (image file, figure caption, and text reference relationship). When tables are involved, the system associates the candidate question with the table reference object identifier. When formulas are involved, the system associates the candidate question with the formula reference object identifier.

[0183] Finally, write all the following information about the candidate issues into the review worksheet: candidate issue identifier, issue type, issue description, source anchor identifier (or anchor identifier that can be retrieved in the current version), image reference object identifier (or table reference object identifier, formula reference object identifier), verification basis, rule basis field, manual review mark, verification status (fill in "passed") and write-back status (initially fill in "pending review").

[0184] After the writing is completed, the system generates a visual entry at the corresponding document node position of the text and image manuscript in the review workpiece, showing the problem type, problem description and verification basis of the candidate problem, and marks the candidate problem as pending review status, with three operation buttons: confirm, reject and close, and then enters the manual review process.

[0185] When an editor views candidate issues in a proofreading worksheet, they can perform the following operations:

[0186] 1) Confirmation Operation: After the editor confirms that the candidate issue is valid, the system will update the status from pending review to confirmed, and record the issue as a valid review conclusion.

[0187] 2) Rejection Operation: If the editor believes that the model identification is incorrect or the problem is not valid, the system will update the status back to "rejected" and record the reason for rejection. The original evidence package and model output hash are retained for subsequent auditing.

[0188] 3) Closing Operation: If the editor believes that the problem does not need to be handled or has been resolved in other aspects, the system will update the status back to closed.

[0189] 4) Transfer to expert review: When the editor is unable to make an independent judgment, the system will update the status back to expert review and record the reason for transfer and the person in charge.

[0190] Each time the status changes, the system records the person who changed the status, the time of the status change, the source anchor point, the object referenced by the image, the hash of the evidence package, the hash of the model output, and the manual processing opinion.

[0191] For high-risk items marked in the manual review boundary (including professional accuracy, copyright, standard clauses, security matters, or political orientation), manual review is mandatory. Editors must confirm or reject the corresponding no-issue results or candidate issues before proceeding to the next stage of publication.

[0192] Through this status transition mechanism, the system can distinguish between model candidate issues, manually confirmed issues, issues with insufficient evidence, and closed issues, thus preventing candidate issues from losing their source location or being repeatedly written back after the manuscript version is updated.

[0193] (3) Writing back the evidence gap

[0194] 1) Writing back the gaps in ordinary evidence

[0195] The system determines the specific location (chapter, section, paragraph, or document node) of the gap in the original text based on the source anchor identifier carried in the evidence gap record or the anchor identifier that can be retrieved in the current version. When images are involved, the system simultaneously determines the corresponding image file based on the image reference object identifier or image asset identifier; when tables are involved, it determines the corresponding table based on the table reference object identifier; and when formulas are involved, it determines the corresponding formula based on the formula reference object identifier.

[0196] The system will write all the following information about the evidence gap into the review workpiece: evidence gap identifier, gap type, missing evidence (such as missing high-resolution original image, missing machine-readable text description in the image), related source anchor identifier or current version back-searchable anchor identifier, related image asset identifier or image reference object identifier, related table reference object identifier, related formula reference object identifier, suggested supplementary evidence, output status, manual review mark, verification status, and write-back status (the initial write-back status is pending supplementation or supplemented evidence).

[0197] Once the data is written, the system generates a visual entry near the corresponding document node, image, table, or formula in the proofreading workpiece. This entry displays the gap type, missing evidence, and suggested supplementary evidence, along with a supplementary evidence button for editors to view and add materials.

[0198] The system records the write-back time, verification status (passed), write-back status (pending supplementary evidence), status changer ("system" when the system operates automatically), evidence package hash, and model output hash for the evidence gap.

[0199] After reviewing the missing evidence entry in the proofreading worksheet, the editor can supplement the evidence as suggested. Upon confirmation and submission, the write-back status will be updated from "Pending Supplementation" to "Supplemented," and the system will record the operator, operation time, description of supplementary materials, and storage location of supplementary materials. If supplementation is not possible or deemed unnecessary, the editor can click the "Close" button to update the write-back status to "Closed," and the system will record the reason for closing.

[0200] Simultaneously, based on the supplementary evidence, the corresponding image reference objects, table reference objects, or formula reference objects are regenerated or updated. Taking images as an example, if the supplemented image is a high-resolution original image, the system re-executes image asset registration, image text recognition, and image title / text reference association to generate a new image reference object; if the supplemented image is text description, the system injects the text into the verification basis field of the corresponding image reference object; if the supplemented image title or text reference is a corrected image title or text reference, the system updates the association relationship of the image title node identifier or text reference node identifier.

[0201] Replace the original corresponding object in the internal multimodal evidence package with the updated reference object, and return to step six using the replaced internal multimodal evidence package.

[0202] 2) Writing back the gaps in location evidence

[0203] For location evidence gaps caused by the inability to recover the source anchor point due to downgraded positioning, the system will write the old source anchor point identifier, current manuscript version identifier, reason for recovery failure, and manual processing suggestions carried in the location evidence gap record into the manual confirmation node or the dedicated node for supplementary evidence transfer in the review workpiece. This node is an independent processing area and is not associated with any document node or non-text object.

[0204] Based on the old source anchors, current manuscript version identifiers, reasons for recovery failures, and candidate matching fragments recorded in the location evidence gaps, editors manually determine whether the problem should exist in the current manuscript version and to which specific location it should be attributed.

[0205] The method proposed in this invention will be described in detail below with reference to specific embodiments.

[0206] Example 1: Candidate Issue Generation and Writeback for Normal Path of DOCX Text and Image Manuscripts

[0207] This embodiment corresponds to the complete normal path from step one to step eight.

[0208] The system first receives a DOCX format graphic manuscript, generates a manuscript version identifier V20260619-01, and establishes the corresponding manuscript version asset.

[0209] Subsequently, the system identifies the source format as DOCX. After unpacking the DOCX file, it identifies chapter titles, paragraphs, figure titles, embedded images, and text citations according to the document reading order. For each identified object, a document node identifier and node sequence number are generated, and a chapter path is formed based on the current title stack, constructing a unified semantic document structure. Text-type document nodes are segmented according to the chapter path, node order, block type, and preset length threshold, generating a set of text blocks. Taking the block with text block identifier B0007 as an example, the corresponding chapter path is Chapter 3 / Section 3.2, the block type is text, and the original text fragment length is 126 characters.

[0210] The system generates source anchors for document nodes and text blocks. It concatenates the manuscript version identifier V20260619-01, text block identifier B0007, chapter path "Chapter 3 / Section 3.2", block type "body text", standard fragment, prefix text and suffix text in a fixed order to form the anchor text string and performs a hash operation. The hash value is taken as the anchor hash prefix H8F21C. The source anchor identifier combination is generated as A-V20260619-01-B0007-H8F21C.

[0211] The anchor confidence score was calculated by weighting the completeness of the chapter path, the effective length of the original text segment, the matching degree between the preceding and following text, and the uniqueness of the node. The result was 0.93, which is higher than the rule threshold of 0.70. This source anchor can be directly used to generate candidate questions.

[0212] Embedded images are extracted from the DOCX file. An image asset identifier (IMG-V20260619-01-003-C9A12F, abbreviated as IMG-003) is generated based on the manuscript version identifier V20260619-01, image sequence number 003, and the image binary checksum prefix C9A12F. The image title is identified as 3-2 with the title "Load Direction Diagram," and the text reference to "As shown in Figure 3-2" is also identified. The node containing the image title, the node containing the text reference, the chapter path, and the associated source anchor point are recorded. The matching status of the image reference object is calculated based on the consistency of the image number, the distance of the chapter path, and the distance of the node sequence, indicating a match.

[0213] The text chunks, source anchor A-V20260619-01-B0007-H8F21C, image references, rule version PUB-RULE-2.1, manuscript stage "initial review", output schema version, and manual review boundaries are encapsulated into an internal multimodal evidence package, and the evidence package hash is calculated as the evidence package identifier. The internal multimodal evidence package is input into the model, and the model outputs candidate question Q-0005 within the evidence package scope. The output status is candidate question, and the verification criteria include figure titles, text references, and image references. Manual review is marked as requiring manual review.

[0214] Upon receiving Q-0005, first verify if the output status is a candidate issue. Then, based on the source anchor point A-V20260619-01-B0007-H8F21C, check the source anchor point index to confirm that the manuscript version identifier, block identifier, and anchor point hash are consistent. Next, verify whether the image reference object belongs to the same manuscript version and whether its associated source anchor point is consistent with the source anchor point in the model output. After the above basic field verifications pass, follow step six to bind Q-0005 to the source anchor point A-V20260619-01-B0007-H8F21C and the image asset IMG-003, write it back to the text node corresponding to B0007 in the review workpiece, and record the verification status as passed, the write-back status as pending review, the manual review status, and the write-back time in the review workpiece.

[0215] Example 2: Handling Boundary Paths with Uncertain Title Attribution

[0216] In this embodiment, the generation of the image reference object in step four and the abnormal triggering of the pre-data in step seven are corresponding to the two steps.

[0217] The system receives an image link (IMG-011) from a converted PDF. Following step two, it identifies an empty title and the text reference is "see image below." The source anchor confidence score is 0.58, below the rule threshold of 0.70. At this point, the image reference is marked as pending confirmation. The system generates evidence gap G-0011, with the gap type being uncertain title attribution, and sets the manuscript version to a manuscript-level manual confirmation status. This evidence gap is a pre-existing data anomaly caused by uncertain title attribution before constructing the internal multimodal evidence package; the system treats it as a standard system evidence gap.

[0218] This embodiment illustrates that when a definite association cannot be established between an image referenced object, the system does not generate a definitive candidate question but instead converts it into an evidence gap or manual confirmation.

[0219] Example 3: Degraded Location and Blocking When Source Anchor Points are Missing or Cannot Be Retrieved

[0220] This embodiment corresponds to the downgraded location processing in step seven.

[0221] The problem description for model return result E-0009 is that the figure title and the text reference are inconsistent, but the returned fields lack source anchors and verification basis. First, we try to perform downgraded positioning based on chapter path, original text fragment, prefix text, or suffix text. Since E-0009 does not carry positioning features that can be used for partial matching, it is impossible to recover the anchors that can be retrieved in the current version. At the same time, it lacks verification basis. The system sets the verification status of E-0009 to fail. The reason for failure is missing source anchors and missing verification basis. The system blocks the result from being entered into the database. The system only records the failure reason field, input hash, output hash, model version, and error code. It does not write candidate issues back to the review workpiece.

[0222] Example 4: Candidate Question Generation and Write-back in Scenarios Where Images and Text Do Not Match

[0223] This embodiment corresponds to the complete path from step one to step eight, demonstrating the normal handling process for text-image consistency conflicts.

[0224] The text block B0012 in the image and text manuscript contains the text "As shown in Figure 1-1, the curve shows an upward trend," with the corresponding chapter path being "Chapter 1 / Section 1.1." The system generates a source anchor point A-V20260619-02-B0012-H31D9A, identifies Figure 1-1 referenced in the text, and establishes an image reference object FR-011 with the image asset IMG-011, the figure title "Temperature Change Curve" (Figure 1-1), and the image block B0011. The matching status is "Matched," and the associated source anchor point and the source anchor point of the text block are within the same chapter and adjacent node range.

[0225] The main text of B0012, the source anchor A-V20260619-02-B0012-H31D9A, the image reference object FR-011, the image address, the rule version PUB-RULE-2.1, the subject plugin version PLUG-CURVE-1.0, the output schema version, and the manual review boundary are encapsulated into an internal multimodal evidence package.

[0226] The verification model identified that the main trend of the curve in the image was downward, and output candidate problem Q-0012. The problem type is suspected conflict between text and image consistency. The problem description is that the text states "the curve is on an upward trend" while the main direction of the curve in the image is downward. The verification basis includes the text reference, the image reference object FR-011, and the image recognition result. The manual review is marked as requiring manual review.

[0227] Upon receiving Q-0012, the system first verifies whether the output status is a candidate issue. Then, based on the source anchor point A-V20260619-02-B0012-H31D9A, it checks the source anchor point index to confirm that the manuscript version identifier, block identifier, and anchor point hash are consistent. Next, it verifies whether the image reference object FR-011 belongs to the same manuscript version and whether its figure number, figure title, and text reference correspond to each other. After the above basic field verifications pass, Q-0012 is written back to the text node corresponding to B0012 in the review workpiece, simultaneously displaying the associated image IMG-011, figure title, text reference, and verification basis. The responsible editor can perform confirmation, rejection, or referral to expert review operations in the review workpiece. If the responsible editor confirms the issue is valid, the system changes its write-back status from pending review to confirmed. If the editor believes the model identification is incorrect, the system records the rejection reason and retains the original evidence package and output hash for subsequent auditing.

[0228] Example 5: Pre-check for table and formula references

[0229] This embodiment corresponds to the generation of table node references and formula node references in step two, and the processing of table node references and formula node references contained in the evidence package in step three.

[0230] The text block B0031 in the document contains the text "Calculation results are shown in Table 4-2". The system generates a source anchor point A-V20260619-03-B0031-H42AC and extracts the table node TBL-042, table title, the path of the chapter containing the table, and the text reference node from the semantic document structure. The table reference object includes the table node identifier TBL-042, table title, the path of the chapter containing the table, the text reference node identifier, and the associated source anchor point A-V20260619-03-B0031-H42AC.

[0231] For formula-related scenarios, the system extracts the formula node EQ-017, formula number, formula image or machine-readable formula text, and text references. The formula reference object includes the formula node identifier EQ-017, formula number, the chapter path where the formula is located, the text reference node identifier, the formula carrier type, and the associated source anchor. In this scenario, the formula reference is used to point the formula number and reference expression in the text to the corresponding formula node, and is written into the evidence package along with the document node identifier and the block identifier where the reference is located.

[0232] Based on the same inventive concept, this invention proposes a text and image manuscript pre-detection and write-back system based on AI multimodal evidence packages, comprising:

[0233] The document parsing module is used to obtain the text and image manuscript to be reviewed and generate a manuscript version identifier for the manuscript; it parses the text and image manuscript into various document nodes and further obtains text blocks;

[0234] Anchor point generation module, which is used to generate source anchor points for each document node and text block;

[0235] The image reference construction module is used to generate image reference objects that represent the reference relationship between images and document text, based on document nodes and text blocks.

[0236] The evidence package construction module is used to generate an internal multimodal evidence package based on the manuscript version identifier, source anchor, text blocks, and image reference objects;

[0237] The model invocation module is used to input the internal multimodal evidence package into the AI ​​verification model and generate multiple verification results containing state fields and output fields.

[0238] The output verification module performs basic field verification on each verification result. For verification results that pass the basic field verification and whose status field is "no problem" or "candidate problem", they directly enter the write-back module. For verification results that fail the basic field verification, or verification results that pass the basic field verification and whose status field is "evidence gap", a special verification for evidence gap is performed. If the verification passes, the result enters the write-back module.

[0239] The write-back module is used to write back the verification results to the review workpiece. For verification results with a status field of "no problem" and that do not hit the boundary of manual review, no manual review is required. For verification results with a status field of "no problem" and that hit the boundary of manual review, as well as verification results with a status field of "candidate problem" or "evidence gap", a visual entry is generated in the review workpiece and manual review is performed.

[0240] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages, characterized in that, include: S1, Obtain the text and image manuscript to be reviewed and generate a manuscript version identifier for the text and image manuscript; S2, the text and image manuscript is parsed into multiple document nodes, further resulting in text blocks, and source anchors are generated for each document node and text block; The document nodes include title nodes, paragraph nodes, image nodes, image caption nodes, and body text reference nodes; the text blocks are obtained by merging title nodes and paragraph nodes; S3, based on document nodes and text blocks, generates an image reference object to represent the reference relationship between the image and the manuscript text; S4. Based on the manuscript version identifier, source anchor, text segmentation, and image reference objects, an internal multimodal evidence package is generated; the internal multimodal evidence package is used to limit the processing scope of the AI ​​verification model. S5, input the internal multimodal evidence package into the AI ​​verification model to generate multiple verification results containing state fields and output fields, and perform basic field verification on each verification result; the state fields are no problem, candidate problem, or evidence gap; For verification results that pass the basic field verification and whose status field is "no problem" or "candidate problem", proceed directly to S6; for verification results that fail the basic field verification, or verification results that pass the basic field verification but whose status field is "evidence gap", further evidence gap-specific verification will be performed, and if it passes, proceed to S6. S6. Write the verification results back to the review workbook. For verification results with a status field of "no problem" and that do not hit the boundary of manual review, no manual review is required. For verification results with a status field of "no problem" and that hit the boundary of manual review, as well as verification results with a status field of "candidate problem" or "evidence gap", generate visual entries in the review workbook and perform manual review.

2. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 1, characterized in that, In step S2, the process of generating text blocks is specifically as follows: Based on the position of each document node in the text and image document, the chapter path of that document node is formed and normalized to obtain the normalized value of the chapter path; Within a preset length threshold, adjacent title nodes or adjacent paragraph nodes located in the same chapter path are merged into a text block; at the same time, paragraph nodes exceeding the preset length threshold are split to obtain multiple text blocks.

3. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 1, characterized in that, In S2, the source anchor also includes the source anchor confidence score, which is obtained by calculating the chapter path completeness, effective length of the original text fragment, context matching degree, and node uniqueness of the document node or text block corresponding to the source anchor, and then summing them by weight. Chapter path completeness is a binary criterion for determining whether the document node or text block corresponding to the source anchor point can be traced back to at least one chapter title in the graphic manuscript, and the value is assigned based on the determination result. The effective length of the original text fragment is determined by comparing the fragment length with the preset effective range, minimum positioning length, and maximum positioning length of the source anchor point of the corresponding text block or text reference node; for the source anchor points of other document types, a fixed value is directly used. The context matching score is calculated by comparing the prefix text of the source anchor record with the previous document node or text block, and comparing the suffix text of the source anchor record with the next document node or text block; the score is determined based on the consistency comparison results of the prefix and suffix text. Node uniqueness is determined by counting the total number of times the text block corresponding to the source anchor appears in the text and image manuscript, and then taking the reciprocal of the total number of occurrences.

4. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 1, characterized in that, Specifically, S3 is: S301, Generate an image asset identifier for each image; the image asset identifier is obtained by concatenating the manuscript version identifier, image sequence number and image binary verification value in a fixed field order, and then performing a hash operation; S302, based on the text and image manuscript, the image title node and the text reference node are respectively associated and bound to the image asset identifier of the corresponding image, and an image reference object is further generated; the image reference object includes the image reference object identifier, image asset identifier, image alternative text, image access path, identifier of the text block adjacent to the image, source anchor point of the corresponding image node, associated and bound image title node identifier, associated and bound text reference node identifier, and matching status; the image alternative text is a text description of the image content, and the matching status is used to describe whether the image title node or the text reference node matches the image.

5. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 1, characterized in that, In S4, the internal multimodal evidence package also includes model call context, rule version snapshot, subject plugin identifier, manuscript stage, manual review boundary and output schema version; The model invocation context is used to store the attribute information, running information, and output information of the AI ​​verification model; The rule version snapshot is used to define the verification rules that the AI ​​verification model follows when performing verification. The subject plugin identifier is used to indicate the subject to which the text and image manuscript belongs, and loads the corresponding subject-specific verification logic, terminology and image recognition dimensions; The term "manuscript stage" indicates the current publication stage of a text and image manuscript. The manual review boundary is used to mark content in graphic manuscripts related to knowledge accuracy, copyright, national standards, industry standard clauses, safe operating procedures, and political orientation. The output schema version is used to define the state fields output by the AI ​​validation model and the output fields that each state field must include.

6. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 1, characterized in that, In S5, the AI ​​verification model is a large language model, a multimodal large model, or a visual language model.

7. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 1, characterized in that, In S5, the basic field validation includes output status validity validation, source anchor index retrospective validation, reference object consistency validation, validation basis integrity validation, and manual review mark integrity validation. The output status validity check requires that the status field of the check result must be either "No Problem", "Candidate Problem", or "Evidence Gap". The source anchor index can be used to look up the source anchor referenced in the output field of the validation result, and the corresponding record must be retrieved from all source anchors; The consistency check of referenced objects requires that when the output field of the check result involves references to images, captions, or text, the corresponding image reference object must be referenced. At the same time, the image reference object must belong to the current text and image manuscript, and the source anchor point in the image reference object must be consistent with the source anchor point in the output field or within the allowed range of adjacent nodes. The validation basis integrity check requires that validation basis be provided in the output fields of the validation result; The manual review mark integrity check output field must include the manual review mark field, and the value of this field must be a valid boolean value.

8. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 7, characterized in that, The source anchor index can be back-checked and verified, which also includes a downgraded location process. The specific process of the downgraded location process is as follows: The source anchors referenced in the output fields of the verification results are used as candidate source anchors. The matching values ​​of each source anchor in the text and image manuscript are calculated and weighted by fuzzy matching from four dimensions: chapter path, original text fragment, prefix text and suffix text. The result is obtained as the restored location reliability of each source anchor in the text and image manuscript. If there is a source anchor with a recovery location confidence score higher than the threshold, the candidate source anchor is replaced with the source anchor with the highest recovery location confidence score, and the basic field validation is performed again. If the recovery of location reliability is below the threshold, but one or more features in the four dimensions of chapter path, original text fragment, prefix text or suffix text have a matching value higher than the corresponding threshold, then a location evidence gap is generated. If the reliability of the restored location is lower than the threshold, and the matching values ​​of all source anchors' chapter paths, original text fragments, prefix texts, and suffix texts are all less than or equal to the corresponding threshold, then the verification result will be prevented from being written back to the review workpiece.

9. The method for pre-detection and write-back of text and image manuscripts based on AI multimodal evidence packages according to claim 7, characterized in that, In S5, the process of performing basic field validation on the output result is as follows: For verification results where the status field is "no problem" or "candidate problem", if the output result passes the basic field verification, it will directly enter S6; if any verification in the basic field verification fails, it will be converted into an evidence gap and the field will be filled. If the filling is successful, it will enter the evidence gap special verification; if it cannot be filled, it will directly block the entry into the database. For verification results where the status field is an evidence gap, if the output result passes the basic field verification, it will proceed to the evidence gap-specific verification; if any verification fails, the field will be completed. If the completion is successful, it will proceed to the evidence gap-specific verification; if the completion cannot be successful, the data entry will be blocked directly.

10. A text and image manuscript pre-inspection and write-back system based on AI multimodal evidence packages, used to implement the text and image manuscript pre-inspection and write-back method based on AI multimodal evidence packages as described in claim 1, characterized in that, include: The document parsing module is used to obtain the text and image manuscript to be reviewed and generate a manuscript version identifier for the manuscript; it parses the text and image manuscript into various document nodes and further obtains text blocks; Anchor point generation module, which is used to generate source anchor points for each document node and text block; The image reference construction module is used to generate image reference objects that represent the reference relationship between images and document text, based on document nodes and text blocks. The evidence package construction module is used to generate an internal multimodal evidence package based on the manuscript version identifier, source anchor, text blocks, and image reference objects; The model invocation module is used to input the internal multimodal evidence package into the AI ​​verification model and generate multiple verification results containing state fields and output fields. The output verification module performs basic field verification on each verification result; for verification results that pass the basic field verification and whose status field is "no problem" or "candidate problem", they directly enter the write-back module. For verification results that pass the basic field verification and whose status field is "no problem" or "candidate problem", proceed directly to S6; for verification results that fail the basic field verification or pass the basic field verification but whose status field is "evidence gap", further evidence gap-specific verification is performed, and if it passes, proceed to the write-back module. The write-back module is used to write back the verification results to the review workpiece. For verification results with a status field of "no problem" and that do not hit the boundary of manual review, no manual review is required. For verification results with a status field of "no problem" and that hit the boundary of manual review, as well as verification results with a status field of "candidate problem" or "evidence gap", a visual entry is generated in the review workpiece and manual review is performed.