An intelligent identification method for graphic-text evidence relationship for project acceptance
Patent Information
- Application Number
- CN202611016179.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-09
AI Technical Summary
[0002]现有视觉-语言预训练模型主要学习图文语义对应关系,缺少对事实类型、证据层级与证据充分性的显式建模,当图像与事实语义相关却缺少结论、签章、复查意见等关键证据要素时,容易将相关但证据不足的图像误判为有效证据;文档智能类方法以信息抽取为目标,并不判断所抽取信息是否足以支撑特定事实;基于多模态大模型提示的外部推理方式则可控性与可解释性不足
(1)本发明依据事实类型检索证据需求模板并逐证据要素计算证据缺口度,将证据判断从"图文是否相关"推进到"图像是否足以证明特定事实",区分语义相关与证据支撑,降低相关但缺少关键证据的图像被误判为有效证据的比例。
Smart Images

Figure CN122594936B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and multimodal information processing technology, and in particular to an intelligent method for identifying relationships between textual and graphical evidence for project acceptance. Background Technology
[0002] Existing visual-language pre-trained models primarily learn the semantic correspondence between images and text, lacking explicit modeling of fact types, evidence levels, and evidence sufficiency. When an image is semantically related to a fact but lacks key evidentiary elements such as conclusions, signatures, or review opinions, it is easy to misjudge a relevant but insufficiently supported image as valid evidence. Document intelligence methods aim at information extraction but do not determine whether the extracted information is sufficient to support a specific fact. External reasoning methods based on multimodal large model prompts lack controllability and interpretability.
[0003] Therefore, existing technologies have the following technical problems when automatically reviewing project acceptance image materials: it is difficult to distinguish between "image-text relevance" and "evidence support", and it is easy to misjudge images that are relevant but lack key evidentiary elements as valid evidence; it lacks element-by-element evidence gap measurement for the evidentiary elements required for the fact type; it has a weak ability to identify high-risk samples such as insufficient evidence and cross-modal conflicts, and it uses the same static reasoning for all samples, making it difficult to focus on key evidence and output interpretable sources of risk, resulting in unreliable automatic review results and being unfavorable for assisting manual review.
[0004] Therefore, there is a need to provide an intelligent method for recognizing the relationship between textual and graphical evidence for project acceptance, in order to solve the above problems. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent identification method for textual and graphical evidence relationships for project acceptance, which can distinguish between semantic relevance and evidence support, thereby improving the reliability and interpretability of automatic project acceptance review.
[0006] To achieve the above objectives, this invention provides an intelligent method for identifying relationships between textual and graphical evidence for project acceptance, comprising the following steps: S1: Obtain the project's acceptance description text and candidate image materials; perform atomic fact decomposition on the acceptance description text to obtain several atomic acceptance facts, each atomic acceptance fact including fact text and fact type; perform optical character recognition and image type recognition on the candidate image materials to obtain the OCR text and image evidence type corresponding to the candidate image materials; S2: Retrieve the corresponding evidence requirement template from the preset evidence requirement template library according to the fact type. The evidence requirement template contains several evidence elements corresponding to the fact type. Generate evidence requirement prompts from the evidence requirement templates, input the evidence requirement prompts into the evidence requirement submodule to obtain the fact requirement representation, and input the fact text into the text encoder. S3: Multiple cross-modal processing submodules are used to process candidate image materials and OCR text to obtain image evidence representation; S4: Fuse the representation of factual needs with the representation of image evidence to obtain the initial evidence relationship prediction; S5: Calculate the intensity of reflection; S6: Determine whether the reflection intensity is greater than the preset reflection threshold; When the intensity of reflection exceeds the preset reflection threshold, active evidence detection is triggered and the trigger flag is set to 1. The action strategy selects at least one detection action related to evidence elements, OCR areas or cross-modal conflicts from the preset detection action set and executes it to obtain supplementary evidence representation. When the intensity of reflection is not greater than the preset reflection threshold, active evidence detection is not triggered and the trigger flag is set to 0; S7: When the trigger flag is 1, the correction amount obtained from the supplementary evidence representation is superimposed on the initial evidence relationship prediction according to the reflection intensity, using the reflection intensity as the gating coefficient, to obtain the final evidence relationship prediction; When the trigger flag is 0, the initial evidence relationship prediction is used as the final evidence relationship prediction; and the evidence relationship between the atomic acceptance facts and the candidate image materials is output as supporting, insufficient, or conflicting.
[0007] Preferably, the specific method in S1 includes: Image type recognition extracts image evidence types as one of the following: on-site photos, equipment photos, system screenshots, acceptance reports, signed and stamped pages, rectification photos, or contract documents; optical character recognition outputs the coordinates of the OCR field in the image's layout area, which is used for image area positioning and OCR field positioning in active evidence detection; candidate image materials are organized into candidate image material groups according to acceptance items, rectification items, or time information, and each candidate image material group includes one or more candidate image materials.
[0008] Preferably, the fact types in S2 include at least construction completion, system launch, organization acceptance, acceptance conclusion, signed voucher, or rectification closure; The sources of evidentiary elements should at least be the acceptance conclusion, signatures and seals, acceptance date, acceptance entity, status before rectification, status after rectification, or review opinions; Factual demand representation satisfy: ; In the formula, For text encoders, For the evidence requirements submodule, Indicates the feature fusion operator, This is a prompt for evidence requirements generated from the evidence requirements template.
[0009] Preferably, in S3, the cross-modal processing submodule includes at least a semantic observation submodule, a local detection submodule, and a conflict assistance submodule; the image output by the semantic observation submodule is fused with the OCR joint semantic observation representation, OCR text representation, image evidence type representation, and local evidence representation output by the local detection submodule, and a cross-modal conflict representation is obtained through the conflict assistance submodule; The evidence requirement submodule is a requirement encoding submodule constructed based on text branches of a text encoder or a cross-modal pre-trained backbone; The semantic observation submodule, the local detection submodule, and the conflict assistance submodule share the same cross-modal pre-training backbone. Each role-based cross-modal submodule utilizes a pre-trained cross-modal backbone. Apply the first Individual adapters for each role With character hint vector And thus we obtain: ; In the formula, For the first Each character's independent cue vector; This is a semantic observation submodule. For local detection submodule, For conflict assistance submodules; Joint semantic observation representation of output image and OCR , Output local evidence representation , Output cross-modal conflict representation Image and OCR Joint Semantic Observation Representation Local evidence indicates Image evidence representation Representation of cross-modal conflict The following conditions must be met: ; ; ; ; In the formula, This represents the sequence number of the candidate image material. This refers to the sequence number of the atomic acceptance fact; For the first Image input or visual features obtained after preprocessing candidate image materials through size normalization, image enhancement and visual coding; For the first The OCR text, OCR field sequence, and layout area coordinates of each candidate image material; This represents a factual demand; For OCR text encoder, For image evidence type Type embedding function, image evidence type embedding function For trainable embedding functions or pre-defined embedding maps, For joint semantic observation representation of images and OCR, This is to represent partial evidence. For the representation of image evidence, This represents cross-modal conflict.
[0010] Preferably, in S4, the initial evidence relationship prediction is a scoring vector, which is the initial score on the three types of evidence relationships: supporting, insufficient evidence, and conflict. The initial evidence relationship prediction is normalized to obtain the initial probability. Scoring Vector Represented by factual needs With visual evidence After fusion classification head get: ; ; in For the initial evidence relationship prediction in the 1st Probability of similar evidentiary relationships The terms are, in order, supporting evidence, insufficient evidence, and conflict.
[0011] Preferably, the intensity of reflection in S5 is obtained by the prediction uncertainty of the initial evidence relationship prediction, the modal conflict degree characterized by the difference between the prediction distribution of atomic acceptance facts and candidate image materials in the text and type joint branches, the image branch and the OCR branch, and the evidence gap degree calculated according to the evidence requirement template for each evidence element. Evidence gap The calculation includes: for each evidentiary element in the evidentiary requirement template. From the evidence requirement submodule Output and Image Evidence Representation Satisfaction discrimination head Calculate satisfaction The weighted sum of the unmet elements of each piece of evidence is taken as the evidentiary gap. : ; ; in, In order to be with the first The number of evidence elements in the evidence requirement template that matches the fact type of each atomic acceptance fact. It is a positive integer; For the Sigmoid function, elements of evidence Factor weights, factor weights Evidence required by a pre-set template or obtained through training; intensity of reflection satisfy: ; In the formula, , , Non-negative weights For reflection intensity bias parameters; To predict uncertainty, by The normalized entropy is determined as follows: ; In the formula, modal conflict degree Determined by the conflict degree mapping function: ; in, For cross-modal conflict representation; For the unimodal prediction distribution of the joint branch of text and type, For the single-modal prediction distribution of the image branch, This represents the single-modal prediction distribution for the OCR branch. Represented by factual needs Representation of image evidence types The input text and type are combined to form an auxiliary classification head, which is then used to obtain the result. Represented by factual needs With visual encoding function Candidate image materials Output visual representation The result is obtained after inputting the image branch-aided classification head. Represented by factual needs With OCR text encoder The output OCR representation is obtained after inputting the OCR branch auxiliary classification head; visual encoding function. Cross-modal pre-training backbone Visual branches constitute; To normalize to the [0,1] symmetric Jensen-Shannon divergence operator, For And three pairwise Jensen-Shannon divergences are collision degree mapping functions with input and output normalized to [0,1].
[0012] Preferably, S6 specifically includes: The preset reflection threshold is The preset detection action set is as follows ;when At that time, the action strategy calculates the action score for the detected action. Score the action Sort in descending order, and take the top results. The sequence numbers of each detection action constitute the selected action set. ,in For the preset number of actions and ,exist Weights are obtained by normalization. And by aggregating the evidence, we can obtain supplementary evidence. : ; ; ; in, For the first One detection action, To express the needs of the facts Image evidence representation Cross-modal conflict representation Intensity of reflection With fact type A scoring network for input actions. For the first The evidence extraction function corresponding to each detection action is expressed in terms of factual requirements. Candidate image materials OCR text Representation of cross-modal conflict As input, output supplementary evidence representation corresponding to the detection action. To provide supplementary evidence; Preset detection action set The detection methods include visual region recoding detection, OCR field re-identification detection, signature conclusion region detection, rectification closed-loop element comparison detection, and cross-modal consistency calculation detection. Visual region recoding detection includes target region detection or region cropping of candidate image materials, and visual feature recoding of cropped regions; OCR field re-identification detection includes extracting OCR fields related to evidence elements based on page area coordinates, and re-identifying or encoding OCR fields at the field level; Signature conclusion area detection detection includes detecting signature areas, acceptance conclusion fields, and date fields, and determining their matching relationship with factual requirement representations; Rectification closed-loop element comparison detection includes comparing the pre-rectification image, post-rectification image, rectification time field, and review opinion field in candidate image material groups associated with the same rectification item; Cross-modal consistency calculation detection includes calculating the consistency score between factual text, OCR fields, image area features, and cross-modal conflict representations.
[0013] Preferred final evidence relationship prediction The following relationship must be satisfied: ; In the formula, This serves as a trigger for proactive evidence detection. For the correction vector; when hour , ; when hour , To and Zero vectors of the same dimension; Output evidence relationship The following relationship must be satisfied: ; in, To provide supplementary evidence For the input correction header, In order, they represent supporting evidence, insufficient evidence, and conflict; when hour, ,when At that time, the correction amount according to The gating coefficient acts on .
[0014] Preferably, the encoder, evidence requirement submodule, role-based cross-modal submodule, OCR encoder, image evidence type embedding function, fusion classification head, single-modal auxiliary classification head, satisfaction discrimination head, conflict degree mapping function, action scoring network, and correction head are obtained through multi-task joint training, with a total loss. satisfy: ; ; ; In the formula, To supervise and apply the evidence relationship based on factual level, and to the fusion classification head The result after residual calibration The weighted three-class cross-entropy loss; For element satisfaction loss; For single-mode auxiliary loss; The action loss is to perform weak supervision on the selected detection action by using preset risk type labels or weak supervision labels generated by fact type and modal conflict situation; Loss due to evidence requirement ranking It is a binary cross-entropy function. ;in, The first output of the satisfaction judgment head The predictive satisfaction of each evidence element For the first Satisfaction of each evidentiary element: supervisory label or weak supervisory label; Predicted by the final evidence relationship The results are based on the facts With images The probability of the supporting class, To The supporting images To and Images that are semantically related but lack sufficient evidence. For interval hyperparameters; , , , The non-negative weights of each loss term.
[0015] Therefore, the present invention employs the above-mentioned intelligent identification method for graphic and textual evidence relationships for project acceptance, and the technical effects are as follows: (1) The present invention retrieves evidence requirement templates based on fact types and calculates the evidence gap degree for each evidence element, advancing the evidence judgment from "whether the text and images are related" to "whether the image is sufficient to prove a specific fact", distinguishing between semantic relevance and evidence support, and reducing the proportion of images that are related but lack key evidence being misjudged as valid evidence.
[0016] (2) In this invention, multiple sub-modules share the same cross-modal pre-training backbone and learn different sub-tasks only with independent adapters and role prompts. While retaining pre-training knowledge, the parameters are efficiently adapted to the domain materials, reducing the risk of small sample overfitting.
[0017] (3) The present invention uses the reflection intensity composed of prediction uncertainty, modal conflict degree and evidence gap degree as a gating, triggers active evidence detection only on difficult samples, and performs residual controlled calibration on the initial prediction, taking into account the stability of simple samples and the error correction of difficult samples. At the same time, it can output missing evidence elements or conflict fields as risk interpretation, thereby improving reliability and interpretability. Attached Figure Description
[0018] Figure 1 This is a flowchart of an intelligent identification method for textual and graphical evidence relationships for project acceptance, according to the present invention. Figure 2 This is an organizational framework diagram of the modules in an intelligent identification method for graphic and textual evidence relationships for project acceptance, as described in an embodiment of the present invention. Detailed Implementation
[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0021] Example 1 like Figures 1-2 As shown, this invention provides an intelligent identification method for textual and graphical evidence relationships for project acceptance, comprising the following steps: S1: Obtain the project's acceptance description text and candidate image materials; perform atomic fact decomposition on the acceptance description text to obtain several atomic acceptance facts, each atomic acceptance fact including fact text and fact type; perform optical character recognition and image type recognition on the candidate image materials to obtain the OCR text and image evidence type corresponding to the candidate image materials; The specific methods in S1 include: Record No. The factual text of the acceptance facts of each atom is as follows: Its factual type is , No. The candidate image materials are Its OCR text is The type of its image evidence is ,in , The sequence number; image type recognition identifies the type of image evidence. Extract one of the following: on-site photos, equipment photos, system screenshots, acceptance reports, signed and stamped pages, rectification photos, or contract documents; Optical Character Recognition (OCR) field coordinates are output in the image's layout area coordinates, which are used for image area positioning and OCR field positioning in active evidence detection; Candidate image materials are organized into candidate image material groups according to acceptance items, rectification items, or time information, and each candidate image material group includes one or more candidate image materials.
[0022] S2: Retrieve the corresponding evidence requirement template from the preset evidence requirement template library according to the fact type. The evidence requirement template contains several evidence elements corresponding to the fact type. Generate evidence requirement prompts from the evidence requirement templates, input the evidence requirement prompts into the evidence requirement submodule to obtain the fact requirement representation, and input the fact text into the text encoder. The fact types in S2 include at least construction completion, system launch, organization of acceptance, acceptance conclusion, signed and stamped vouchers, or rectification closure; The corresponding evidence requirement template consists of a set of evidence elements. It means that among them For the first Each element of evidence The total number of evidentiary elements for this template; evidentiary elements The source must be at least the acceptance conclusion, signature and seal, acceptance date, acceptance entity, status before rectification, status after rectification, or review opinion; Factual demand representation satisfy: ; In the formula, For text encoders, For the evidence requirements submodule, Indicates the feature fusion operator, This is a prompt for evidence requirements generated from the evidence requirements template.
[0023] S3: Multiple cross-modal processing submodules are used to process candidate image materials and OCR text to obtain image evidence representation; In S3, the cross-modal processing submodule includes at least a semantic observation submodule, a local detection submodule, and a conflict assistance submodule. The image output by the semantic observation submodule is fused with the joint semantic observation representation of OCR, the OCR text representation, the image evidence type representation, and the local evidence representation output by the local detection submodule. The cross-modal conflict representation is obtained through the conflict assistance submodule. The evidence requirement submodule is a requirement encoding submodule constructed based on text branches of a text encoder or a cross-modal pre-trained backbone; The semantic observation submodule, the local detection submodule, and the conflict assistance submodule share the same cross-modal pre-training backbone. Each role-based cross-modal submodule utilizes a pre-trained cross-modal backbone. Apply the first Individual adapters for each role With character hint vector And thus we obtain: ; In the formula, For the first Each character's independent cue vector; This is a semantic observation submodule. For local detection submodule, For conflict assistance submodules; Joint semantic observation representation of output image and OCR , Output local evidence representation , Output cross-modal conflict representation Image and OCR Joint Semantic Observation Representation Local evidence indicates Image evidence representation Representation of cross-modal conflict The following conditions must be met: ; ; ; ; In the formula, This represents the sequence number of the candidate image material. This refers to the sequence number of the atomic acceptance fact; For the first Image input or visual features obtained after preprocessing candidate image materials through size normalization, image enhancement and visual coding; For the first The OCR text, OCR field sequence, and layout area coordinates of each candidate image material; This represents a factual demand; For OCR text encoder, For image evidence type Type embedding function, image evidence type embedding function For trainable embedding functions or pre-defined embedding maps, For joint semantic observation representation of images and OCR, This is to represent partial evidence. For the representation of image evidence, This represents cross-modal conflict.
[0024] S4: Fuse the representation of factual needs with the representation of image evidence to obtain the initial evidence relationship prediction; In S4, the initial evidence relationship prediction is a scoring vector, which is the initial score on the three types of evidence relationships: supporting, insufficient, and conflicting. The initial probability is obtained after the initial evidence relationship prediction is normalized. Scoring Vector Represented by factual needs With visual evidence After fusion classification head get: ; ; in For the initial evidence relationship prediction in the 1st Probability of similar evidentiary relationships The terms are, in order, supporting evidence, insufficient evidence, and conflict.
[0025] S5: Calculate the intensity of reflection; the intensity of reflection is obtained by the prediction uncertainty of the initial evidence relationship prediction, the modal conflict degree characterized by the difference between the predicted distributions of the atomic acceptance facts and candidate image materials in the text and type joint branches, the image branch and the OCR branch, and the evidence gap degree calculated for each evidence element according to the evidence requirement template. Evidence gap The calculation includes: for each evidentiary element in the evidentiary requirement template. From the evidence requirement submodule Output and Image Evidence Representation Satisfaction discrimination head Calculate satisfaction The weighted sum of the unmet elements of each piece of evidence is taken as the evidentiary gap. : ; ; in, In order to be with the first The number of evidence elements in the evidence requirement template that matches the fact type of each atomic acceptance fact. It is a positive integer; For the Sigmoid function, elements of evidence Factor weights, factor weights Evidence required by a pre-set template or obtained through training; intensity of reflection satisfy: ; In the formula, , , Non-negative weights , , and The method for obtaining it is: during the model training phase, [the following is done] , , and The parameters are used as training parameters for the reflexive strength gating function and are trained together with the satisfaction discriminant head, conflict degree mapping function, action scoring network, and correction head; to ensure , , Set unconstrained intermediate parameters for non-negative weights. , , and order , ; To reflect on the strength bias parameters; during training, the total loss is used. To optimize the target, updates are performed through backpropagation. , , and The result after training is , , and Solidified into model parameters for calculating reflection intensity during the inference phase. ; To predict uncertainty, by The normalized entropy is determined as follows: ; In the formula, modal conflict degree Determined by the conflict degree mapping function: ; in, For the unimodal prediction distribution of the joint branch of text and type, For the single-modal prediction distribution of the image branch, This represents the single-modal prediction distribution for the OCR branch. Represented by factual needs Representation of image evidence types The input text and type are combined to form an auxiliary classification head, which is then used to obtain the result. Represented by factual needs With visual encoding function Candidate image materials Output visual representation The result is obtained after inputting the image branch-aided classification head. Represented by factual needs With OCR text encoder The output OCR representation is obtained after inputting the OCR branch auxiliary classification head; visual encoding function. Cross-modal pre-training backbone Visual branches constitute; To normalize to the [0,1] symmetric Jensen-Shannon divergence operator, For And three pairwise Jensen-Shannon divergences are collision degree mapping functions with input and output normalized to [0,1].
[0026] S6: Determine whether the reflection intensity is greater than the preset reflection threshold; When the intensity of reflection exceeds the preset reflection threshold, active evidence detection is triggered and the trigger flag is set to 1. The action strategy selects at least one detection action related to evidence elements, OCR areas or cross-modal conflicts from the preset detection action set and executes it to obtain supplementary evidence representation. When the intensity of reflection is not greater than the preset reflection threshold, active evidence detection is not triggered and the trigger flag is set to 0; S6 specifically includes: The preset reflection threshold is The preset detection action set is as follows Reflection Threshold The method for obtaining the validation sample set is as follows: a validation sample set is set up outside the training set, and the reflection strength is calculated on the validation samples. Based on the initial evidence relationship of the verification samples, it predicts whether there are errors, missing evidence elements, or cross-modal conflicts, and indicates whether active evidence detection needs to be triggered; within a preset candidate threshold set... Candidate thresholds are selected one by one. ,according to The triggering result is obtained, and the criterion for determining the trigger is the maximum F1 score between the triggering result and the triggering label of the validation sample or the minimum validation error cost. After confirmation This is written into the model configuration as a fixed threshold and used during the inference phase to determine whether to trigger active evidence detection.
[0027] when At that time, the action strategy calculates the action score for the detected action. Score the action Sort in descending order, and take the top results. The sequence numbers of each detection action constitute the selected action set. ,in For the preset number of actions and ,exist Weights are obtained by normalization. And by aggregating the evidence, we can obtain supplementary evidence. : ; ; ; in, For the first One detection action, To express the needs of the facts Image evidence representation Cross-modal conflict representation Intensity of reflection With fact type A scoring network for input actions. For the first The evidence extraction function corresponding to each detection action is expressed in terms of factual requirements. Candidate image materials OCR text Representation of cross-modal conflict As input, output supplementary evidence representation corresponding to the detection action. To provide supplementary evidence; Preset detection action set The detection methods include visual region recoding detection, OCR field re-identification detection, signature conclusion region detection, rectification closed-loop element comparison detection, and cross-modal consistency calculation detection. Visual region recoding detection includes target region detection or region cropping of candidate image materials, and visual feature recoding of cropped regions; OCR field re-identification detection includes extracting OCR fields related to evidence elements based on page area coordinates, and re-identifying or encoding OCR fields at the field level; Signature conclusion area detection detection includes detecting signature areas, acceptance conclusion fields, and date fields, and determining their matching relationship with factual requirement representations; Rectification closed-loop element comparison detection includes comparing the pre-rectification image, post-rectification image, rectification time field, and review opinion field in candidate image material groups associated with the same rectification item; Cross-modal consistency calculation detection includes calculating the consistency score between factual text, OCR fields, image area features, and cross-modal conflict representations.
[0028] Action scoring network When the corresponding fact type or conflict situation meets the preset conditions, the action score for the corresponding detection action is determined. Add a non-negative bias term, with preset conditions including: when To add a non-negative bias term to the detection and probe of the signature / signature area when accepting conclusions or signing documents, when... To add a non-negative bias term to the comparison and detection of rectification loop elements during rectification loop closure, when Middle field and fact text Inconsistent or cross-modal conflict representation When characterizing conflict risk, a non-negative bias term is added to the cross-modal consistency calculation probe.
[0029] S7: When the trigger flag is 1, the correction amount obtained from the supplementary evidence representation is superimposed on the initial evidence relationship prediction according to the reflection intensity, using the reflection intensity as the gating coefficient, to obtain the final evidence relationship prediction; Final Evidence Relationship Prediction The following relationship must be satisfied: ; In the formula, This serves as a trigger for proactive evidence detection. For the correction vector; when hour , ; when hour , To and Zero vectors of the same dimension; Output evidence relationship The following relationship must be satisfied: ; in, To provide supplementary evidence For the input correction header, In order, they represent supporting evidence, insufficient evidence, and conflict; when hour, ,when At that time, the correction amount according to The gating coefficient acts on .
[0030] When the trigger flag is 0, the initial evidence relationship prediction is used as the final evidence relationship prediction; and the evidence relationship between the atomic acceptance facts and the candidate image materials is output as supporting, insufficient, or conflicting.
[0031] This encoder, evidence requirement submodule, role-based cross-modal submodule, OCR encoder, image evidence type embedding function, fusion classification head, single-modal auxiliary classification head, satisfaction discrimination head, conflict degree mapping function, action scoring network, and correction head are obtained through multi-task joint training. The total loss is... satisfy: ; ; ; In the formula, To supervise and apply the evidence relationship based on factual level, and to the fusion classification head The result after residual calibration The weighted three-class cross-entropy loss; For element satisfaction loss; For single-mode auxiliary loss; The action loss is to perform weak supervision on the selected detection action by using preset risk type labels or weak supervision labels generated by fact type and modal conflict situation; Loss due to evidence requirement ranking It is a binary cross-entropy function. ;in, The first output of the satisfaction judgment head The predictive satisfaction of each evidence element For the first Satisfaction of each evidentiary element: supervisory label or weak supervisory label; Predicted by the final evidence relationship The results are based on the facts With images The probability of the supporting class, To The supporting images To and Images that are semantically related but lack sufficient evidence. For interval hyperparameters; The acquisition method is as follows: construct a model from positive sample images. and difficult sample images The validation sample pairs consist of, where To verify the facts of atomic acceptance The image that forms the support To and Images that are semantically relevant but lack sufficient evidence; candidate images are selected from a predefined set of candidate intervals. Calculate the ranking loss and evaluate the support class probability of positive samples. Compared to the probability of supporting classes for hard-to-bear samples The ability to differentiate is assessed; candidate pairs are selected that minimize the validation set ranking loss and achieve the best evidence relationship classification performance. As .
[0032] , , , The non-negative weights of each loss term, , , , The method of obtaining it is: As the primary task loss and with its coefficient fixed at 1, for , , , After performing batch mean normalization or moving average normalization respectively, search the preset non-negative candidate weight set on the validation set. , , , The combination is determined based on the optimal comprehensive index of the three-category classification of evidence relationships, the accuracy rate of evidence element satisfaction identification, and the accuracy rate of active evidence detection triggering. The determined combination... , , , It is used consistently during final training and model training before inference.
[0033] Freeze cross-modal pre-training backbone during training The parameters are optimized only for the character cue vector. ,adapter OCR encoder , Integration and classification head Single-modal auxiliary classification head, satisfaction judgment head Conflict degree mapping function Action scoring network With correction head and embedding functions in image evidence types When for trainable embedding functions Optimize, in Incorrect when using preset embedding mapping Optimize, and visual encoding function With cross-modal pre-training backbone Freeze; or further, apply a learning rate lower than the other parameters to the cross-modal pre-trained backbone. Fine-tuning is performed on some parameters or low-rank adaptation parameters of the visual encoding function. With cross-modal pre-training backbone The visual branches were partially fine-tuned.
[0034] Therefore, the present invention adopts the above-mentioned intelligent identification method for graphic and textual evidence relationships for project acceptance. It uses the reflection intensity composed of prediction uncertainty, modal conflict degree and evidence gap degree as a gating to trigger active evidence detection on difficult samples and perform residual controlled calibration on the initial prediction. It takes into account the stability of simple samples and the error correction of difficult samples. At the same time, it can output missing evidence elements or conflict fields as risk interpretation, thereby improving reliability and interpretability.
[0035] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A project acceptance-oriented intelligent identification method for graphic-text evidence relationship, characterized in that, Includes the following steps: S1: Obtain the project's acceptance description text and candidate image materials; perform atomic fact decomposition on the acceptance description text to obtain several atomic acceptance facts, each atomic acceptance fact including fact text and fact type; Perform optical character recognition and image type recognition on candidate image materials to obtain OCR text and image evidence type corresponding to the candidate image materials; S2: Retrieve the corresponding evidence requirement template from the preset evidence requirement template library according to the fact type. The evidence requirement template contains several evidence elements corresponding to the fact type. Generate evidence requirement prompts from the evidence requirement templates, input the evidence requirement prompts into the evidence requirement submodule to obtain the fact requirement representation, and input the fact text into the text encoder. S3: Multiple cross-modal processing submodules are used to process candidate image materials and OCR text to obtain image evidence representation; S4: Fuse the representation of factual needs with the representation of image evidence to obtain the initial evidence relationship prediction; S5: Calculate the intensity of reflection; the intensity of reflection is obtained by the prediction uncertainty of the initial evidence relationship prediction, the modal conflict degree characterized by the difference between the predicted distributions of the atomic acceptance facts and candidate image materials in the text and type joint branches, the image branch and the OCR branch, and the evidence gap degree calculated for each evidence element according to the evidence requirement template. S6: Determine whether the reflection intensity is greater than the preset reflection threshold; When the intensity of reflection exceeds the preset reflection threshold, active evidence detection is triggered and the trigger flag is set to 1. The action strategy selects at least one detection action related to evidence elements, OCR areas or cross-modal conflicts from the preset detection action set and executes it to obtain supplementary evidence representation. When the intensity of reflection is not greater than the preset reflection threshold, active evidence detection is not triggered and the trigger flag is set to 0; S7: When the trigger flag is 1, the correction amount obtained from the supplementary evidence representation is superimposed on the initial evidence relationship prediction according to the reflection intensity, using the reflection intensity as the gating coefficient, to obtain the final evidence relationship prediction; When the trigger flag is 0, the initial evidence relationship prediction is used as the final evidence relationship prediction; It outputs the evidentiary relationship between the atomic acceptance facts and the candidate image materials, indicating whether the evidence is supporting, insufficient, or conflicting.
2. The intelligent recognition method for graphic and textual evidence relationships for project acceptance as described in claim 1, characterized in that, The specific methods in S1 include: Image type recognition extracts image evidence types as one of the following: on-site photos, equipment photos, system screenshots, acceptance reports, signed and stamped pages, rectification photos, or contract documents; optical character recognition outputs the coordinates of the OCR field in the image's layout area, which is used for image area positioning and OCR field positioning in active evidence detection; candidate image materials are organized into candidate image material groups according to acceptance items, rectification items, or time information, and each candidate image material group includes one or more candidate image materials.
3. The intelligent recognition method for textual and graphical evidence relationships for project acceptance as described in claim 1, characterized in that, The fact types in S2 include at least construction completion, system launch, organization of acceptance, acceptance conclusion, signed and stamped vouchers, or rectification closure; The sources of evidentiary elements should at least be the acceptance conclusion, signatures and seals, acceptance date, acceptance entity, status before rectification, status after rectification, or review opinions; Factual demand representation satisfy: ; In the formula, For text encoders, State the facts, For the evidence requirements submodule, Indicates the feature fusion operator, This is a prompt for evidence requirements generated from the evidence requirements template.
4. The intelligent recognition method for graphic and textual evidence relationships for project acceptance as described in claim 1, characterized in that, In S3, the cross-modal processing submodule includes at least a semantic observation submodule, a local detection submodule, and a conflict assistance submodule. The image output by the semantic observation submodule is fused with the joint semantic observation representation of OCR, the OCR text representation, the image evidence type representation, and the local evidence representation output by the local detection submodule. The cross-modal conflict representation is obtained through the conflict assistance submodule. The evidence requirement submodule is a requirement encoding submodule constructed based on text branches of a text encoder or a cross-modal pre-trained backbone; The semantic observation submodule, the local detection submodule, and the conflict assistance submodule share the same cross-modal pre-training backbone. Each role-based cross-modal submodule utilizes a pre-trained cross-modal backbone. Apply the first Individual adapters for each role With character hint vector And thus we obtain: ; In the formula, For the first Each character's independent cue vector; This is a semantic observation submodule. For local detection submodule, For conflict assistance submodules; Joint semantic observation representation of output image and OCR , Output local evidence representation , Output cross-modal conflict representation Image and OCR Joint Semantic Observation Representation Local evidence indicates Image evidence representation Representation of cross-modal conflict The following conditions must be met: ; ; ; ; In the formula, This represents the sequence number of the candidate image material. This refers to the sequence number of the atomic acceptance fact; For the first Image input or visual features obtained after preprocessing candidate image materials through size normalization, image enhancement and visual coding; For the first The OCR text, OCR field sequence, and layout area coordinates of each candidate image material; This represents a factual demand; For OCR text encoder, For image evidence type Type embedding function, image evidence type embedding function For trainable embedding functions or pre-defined embedding maps, For joint semantic observation representation of images and OCR, This is to represent partial evidence. For the representation of image evidence, This represents cross-modal conflict.
5. The intelligent recognition method for graphic and textual evidence relationships for project acceptance as described in claim 1, characterized in that, In S4, the initial evidence relationship prediction is a scoring vector, which is the initial score on the three types of evidence relationships: supporting, insufficient, and conflicting. The initial probability is obtained after the initial evidence relationship prediction is normalized. Scoring Vector Represented by factual needs With visual evidence After fusion classification head get: ; ; in For the initial evidence relationship prediction in the 1st Probability of similar evidentiary relationships The terms are, in order, supporting evidence, insufficient evidence, and conflict.
6. The intelligent recognition method for textual and graphical evidence relationships for project acceptance as described in claim 5, characterized in that, Evidence gap in S5 The calculation includes: for each evidentiary element in the evidentiary requirement template. The evidence requirement submodule Output and Image Evidence Representation Satisfaction discrimination head Calculate satisfaction The weighted sum of the unmet elements of each piece of evidence is taken as the evidentiary gap. ; ; ; in, In order to be with the first The number of evidence elements in the evidence requirement template that matches the fact type of each atomic acceptance fact. It is a positive integer; For the Sigmoid function, elements of evidence Factor weights, factor weights Evidence required by a pre-set template or obtained through training; intensity of reflection satisfy: ; In the formula, , , Non-negative weights For reflection intensity bias parameters; To predict uncertainty, by The normalized entropy is determined as follows: ; In the formula, modal conflict degree Determined by the conflict degree mapping function: ; in, For cross-modal conflict representation; For the unimodal prediction distribution of the joint branch of text and type, For the single-modal prediction distribution of the image branch, This represents the single-modal prediction distribution for the OCR branch. Represented by factual needs Representation of image evidence types The input text and type are combined to form an auxiliary classification head, which is then used to obtain the result. Represented by factual needs With visual encoding function Candidate image materials Output visual representation The result is obtained after inputting the image branch-aided classification head. Represented by factual needs With OCR text encoder The output OCR representation is obtained after inputting the OCR branch auxiliary classification head; visual encoding function. Cross-modal pre-training backbone Visual branches constitute; To normalize to the [0,1] symmetric Jensen-Shannon divergence operator, For And three pairwise Jensen-Shannon divergences are collision degree mapping functions with input and output normalized to [0,1].
7. The intelligent recognition method for graphic and textual evidence relationships for project acceptance as described in claim 6, characterized in that, S6 specifically includes: The preset reflection threshold is The preset detection action set is as follows ;when At that time, the action strategy calculates the action score for the detected action. Score the action Sort in descending order, and take the top results. The sequence numbers of each detection action constitute the selected action set. ,in For the preset number of actions and ,exist Weights are obtained by normalization. And by aggregating the evidence, we can obtain supplementary evidence. : ; ; ; in, For the first One detection action, To express the needs of the facts Image evidence representation Cross-modal conflict representation Intensity of reflection With fact type A scoring network for input actions. For the first The evidence extraction function corresponding to each detection action is expressed in terms of factual requirements. Candidate image materials OCR text Representation of cross-modal conflict As input, output supplementary evidence representation corresponding to the detection action. To provide supplementary evidence; Preset detection action set The detection methods include visual region recoding detection, OCR field re-identification detection, signature conclusion region detection, rectification closed-loop element comparison detection, and cross-modal consistency calculation detection. Visual region recoding detection includes target region detection or region cropping of candidate image materials, and visual feature recoding of cropped regions; OCR field re-identification detection includes extracting OCR fields related to evidence elements based on page area coordinates, and re-identifying or encoding OCR fields at the field level; Signature conclusion area detection detection includes detecting signature areas, acceptance conclusion fields, and date fields, and determining their matching relationship with factual requirement representations; Rectification closed-loop element comparison detection includes comparing the pre-rectification image, post-rectification image, rectification time field, and review opinion field in candidate image material groups associated with the same rectification item; Cross-modal consistency calculation detection includes calculating the consistency score between factual text, OCR fields, image area features, and cross-modal conflict representations.
8. The intelligent recognition method for graphic and textual evidence relationships for project acceptance as described in claim 7, characterized in that, Final Evidence Relationship Prediction The following relationship must be satisfied: ; In the formula, This serves as a trigger for proactive evidence detection. For the correction vector; when hour , ; when hour , To and Zero vectors of the same dimension; Output evidence relationship The following relationship must be satisfied: ; in, To supplement evidence For the input correction header, In order, they represent supporting evidence, insufficient evidence, and conflict; when hour, ,when At that time, the correction amount according to The gating coefficient acts on .
9. The intelligent recognition method for graphic and textual evidence relationships for project acceptance as described in claim 8, characterized in that, This encoder, evidence requirement submodule, role-based cross-modal submodule, OCR encoder, image evidence type embedding function, fusion classification head, single-modal auxiliary classification head, satisfaction discrimination head, conflict degree mapping function, action scoring network, and correction head are obtained through multi-task joint training. The total loss is... satisfy: ; ; ; In the formula, To supervise and apply the evidence relationship based on factual level, and to the fusion classification head The result after residual calibration The weighted three-class cross-entropy loss; For element satisfaction loss; For single-mode auxiliary loss; The action loss is to perform weak supervision on the selected detection action by using preset risk type labels or weak supervision labels generated by fact type and modal conflict situation; Loss due to evidence requirement ranking It is a binary cross-entropy function. ;in, The first output of the satisfaction judgment head The predictive satisfaction of each evidence element For the first Satisfaction of each evidentiary element is monitored by a supervisory label or a weakly monitored label. Predicted by the final evidence relationship The results are based on the facts With images The probability of the supporting class, To The supporting images To and Images that are semantically related but lack sufficient evidence. For interval hyperparameters; , , , The non-negative weights of each loss term.
Citation Information
Patent Citations
Diagnostic ICD automatic coding method and system based on medical record semantic understanding
CN115859914A
False news identification method and system for multi-modal evidence collection
CN122200309A