Method for analyzing and extracting key information of letters and visits information based on machine learning

CN122594484APending Publication Date: 2026-08-18WUHAN CHUYU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610674803.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]但在同一信访事项存在多源材料、版本折叠、字段覆盖不完整以及跨材料内容冲突时,现有技术多停留在字段层的静态置信融合或简单投票,缺乏面向材料级别的证据承诺表达与可追溯冲突清单,难以形成按字段与材料差异化的动态折扣抑制机制;同时缺少以最小修复一致推断为核心的全局一致推断与规则修复、补证修复二分输出,导致冲突解释与处置建议难以闭环回流更新,影响关键信息提取的一致性与可控性

Benefits of technology

本发明通过将同一信访事项的多源信访材料归并为信访事项工作空间并形成材料包集合,在统一的文本与结构规范化处理基础上生成证据片段集合、字段候选集合及候选证据片段,实现对信访关键信息从材料到字段的连续可追溯建模;进一步为每份材料生成包含哈希聚合承诺、字段覆盖摘要、结构一致性可检核摘要、载体质量摘要与来源分布摘要的证据承诺向量集合,并输出承诺冲突清单与覆盖缺口结果,使跨材料冲突与缺口由隐性问题转化为可检核、可回溯的结构化输入,从源头降低多材料环境下字段抽取的不确定传播。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594484A_ABST
    Figure CN122594484A_ABST
Patent Text Reader

Abstract

The application discloses a method for analyzing and extracting key information of letters and visits based on machine learning, comprising the following steps: obtaining multi-source materials of the same letter and visit matter and forming a material package; generating evidence fragments and field candidates and binding candidate evidence for the material package; generating evidence commitment vectors, commitment conflict lists and coverage gaps for each material; synthesizing commitment layer consistency to obtain reliable source posterior and dynamic discount, completing field layer evidence synthesis and conflict contribution disassembly; constructing a minimum repair consistent inference factor graph for global consistent inference; generating an explanation-repair integrated package and backflow updating to form a closed loop. The application adopts the dynamic discount of evidence commitment vectors and commitment layer consistency synthesis, the closed loop method of the minimum repair consistent inference factor graph, realizes traceable extraction of key information of letters and visits and minimum repair of conflicts, and has the advantages of strong consistency, strong explainability and high self-correction ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of petition information processing technology, and in particular to a method for petition information analysis and key information extraction based on machine learning. Background Technology

[0002] Existing petition information processing technologies typically revolve around the collection and archiving of petition materials, text standardization, element extraction, and classification. By uniformly formatting and structurally parsing carriers such as letters and visits, online forms, scanned documents, and receipts, fields such as petitioner, time, location, involved unit, request, and nature of problem are extracted. Some solutions introduce sequence labeling or classification models to output candidate fields and combine rule matching or similarity calculation to achieve classification and processing path recommendation, thereby improving the efficiency of initial classification and information entry of petitions.

[0003] However, when there are multiple sources of materials, version folding, incomplete field coverage, and cross-material content conflicts in the same petition, existing technologies mostly remain at the static confidence fusion or simple voting at the field level. They lack evidence commitment expression and traceable conflict list at the material level, making it difficult to form a dynamic discount suppression mechanism based on field and material differences. At the same time, they lack global consistent inference and rule repair and supplementary evidence repair binary output with minimum repair consistent inference as the core, which makes it difficult for conflict interpretation and handling suggestions to be updated in a closed loop, affecting the consistency and controllability of key information extraction.

[0004] Therefore, how to provide machine learning-based methods for analyzing petition information and extracting key information is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a machine learning-based method for analyzing petition information and extracting key information. This invention employs a closed-loop method of dynamic discounting and minimum repair consistency inference factor graph, which combines evidence commitment vectors and commitment layer consistency synthesis, to achieve traceable extraction of key petition information and minimum conflict repair. It has the advantages of strong consistency, strong interpretability, and high self-correction capability.

[0006] The method for analyzing petition information and extracting key information based on machine learning according to embodiments of the present invention includes the following steps: Collect multiple sources of petition materials for the same petition matter and merge them into a petition matter workspace to generate a material package collection; The material package set is subjected to text and structure normalization processing to generate a set of evidence fragments, and a field candidate set is further generated and bound to the candidate evidence fragments; Based on the evidence fragment set and the material package set, an evidence commitment vector set is generated for each material, and a commitment conflict list and coverage gap result are generated; Perform commitment-layer consistency synthesis on the evidence commitment vector set to generate reliable posterior results and dynamic discount results. Based on the dynamic discount results, perform field-layer evidence synthesis on the field candidate set to output the field synthesis result set and the field conflict contribution decomposition result. Based on the set of field synthesis results, the list of commitment conflicts, the coverage gap results and candidate evidence fragments, a minimum repair consistency inference factor graph is constructed and global consistency inference is performed to obtain the global field combination results, field consistency conclusions and minimum repair set; Based on the field consistency conclusions, field conflict contribution decomposition results, and minimum repair set, an explanation-repair package is generated to export key information results. The source reliable posterior results, dynamic discount results, and minimum repair consistency inference factor graph are then fed back to form a closed loop.

[0007] Optionally, the generation of the material package set specifically includes: Obtain multiple sources of petition materials corresponding to the same petition matter, generate material source markers, collection time markers and carrier type markers, and obtain the original set of multiple sources of petition materials; The original set of multi-source petition materials is processed to be uniformly formatted and parsable, resulting in a unified formatted set of materials. For each document in the unified format document set, generate a document fingerprint set and a content summary set; Deduplication and version folding are performed based on the material fingerprint set to form an effective material set and version cluster set; The system merges the valid material set and version cluster set of petitions. Materials belonging to the same petition are grouped into the petition workspace according to the petition identifier. The materials are organized into material packages using the material source tag and carrier type tag as grouping keys, and the merge integrity is checked.

[0008] Optionally, the generation of the candidate field set and candidate evidence fragments specifically includes: Read the material text and carrier type tag of each material package in the material package set to obtain a standardized material text set; Based on the standardized material text set and the original image carrier, structural normalization processing is performed to obtain a set of structural elements; The standardized material text set is subjected to text normalization processing to form a standardized element text set, and source location information is generated; Based on the set of structural elements, the standardized material text set is segmented into evidence fragments to generate candidate fragments. The scores of the evidence fragments are calculated, and the candidate fragments that meet the threshold conditions are selected and packaged into a set of evidence fragments. Generate a candidate set of fields and a candidate probability score based on a standardized set of element texts and a set of evidence fragments; Based on the candidate probability score, candidate normalization is performed on the field candidate set and bound to the candidate evidence fragments. The candidate confidence is calculated and encapsulated into a candidate evidence fragment with the corresponding evidence fragment set.

[0009] Optionally, the evidence segmentation is performed in the order of priority for title area, priority for table row unit, and priority for body paragraph. The evidence segment score is obtained by calculating the fragment completeness score, structural reliability score, and carrier quality score of the candidate fragments and then summing them according to preset non-negative weights.

[0010] Optionally, the generation of the commitment conflict list and coverage gap results specifically includes: Read the evidence fragment set, the field candidate set, the candidate evidence fragment and the material package set, and construct the evidence commitment vector field template for each material in the material package set; Generate a hash aggregation commitment for each material in the material package set; Generate a field overlay summary for each material in the material package set; Generate a structurally consistent checkable summary for each material in the material package set; For each material in the material package set, generate a carrier quality summary and a source distribution summary; Based on the evidence commitment vector field template, each material is assembled with hash aggregation commitment, field coverage summary, structural consistency checkable summary, carrier quality summary, source distribution summary and candidate evidence fragment citation mark to form an evidence commitment vector. The evidence commitment vector set is then generated by aggregating the material package sets. A list of commitment conflicts is obtained by performing pairwise consistency comparisons on the evidence commitment vectors of different materials within the same petition workspace. Based on the field coverage summary, the coverage field marks of all materials in the workspace of the same petition matter are summarized field by field in the preset field system to generate the coverage gap result.

[0011] Optionally, the generation of the field synthesis result set and the field conflict contribution decomposition result specifically includes: Read the evidence commitment vector set, commitment conflict list, field candidate set and candidate evidence fragment, and generate a commitment layer aligned input set for each evidence commitment vector in the evidence commitment vector set; Calculate the commitment conflict metric set based on the commitment layer aligned input set and the commitment conflict list; The reliable posterior result of the source is calculated by performing a reliable posterior result of the source based on the set of commitment conflict metrics and the source distribution summary; Dynamic discounting results are generated based on reliable posterior results from sources, a set of commitment conflict metrics, and coverage gap results. Based on the dynamic discount results, perform field-level evidence synthesis on the candidate field set and output the synthesized field result set; The results of decomposing field conflict contributions are generated based on the dynamic discount results and the field synthesis result set.

[0012] Optionally, the generation of the global field combination result, field consistency conclusion, and minimum repair set specifically includes: Read the set of field synthesis results, the list of commitment conflicts, the coverage gap results and candidate evidence fragments, and construct a minimum repair consistency inference factor graph; Generate a candidate aligned input set based on the field synthesis result set and candidate evidence fragments; A set of consistency constraint inputs is generated based on the mapping between the commitment conflict list and the coverage gap results; Perform global consistency inference on the minimum repair consistency inference factor graph, calculate the consistency target quantity based on the candidate alignment input set and the consistency constraint input set, search for the field assignment combination with the minimum consistency target quantity, output the global field combination result and generate the field consistency conclusion; When the field consistency conclusion is inconsistent or requires supplementary verification, the constraint entries that lead to an increase in the consistency target quantity are extracted from the consistency constraint input set as a candidate repair entry set, and divided into a rule repair candidate set and a supplementary verification repair candidate set. The minimum set of rule repair candidates and the minimum set of supplementary evidence repair candidates are each subjected to a minimum set of repairs.

[0013] Optionally, the minimumity decision uses the minimum number of repair items as the first minimumity criterion and the maximum decrease in the consistency target quantity as the second minimumity criterion. First, the minimum rule repair subset that satisfies the first minimumity criterion is selected from the rule repair candidate set and encapsulated into a rule repair set. Then, the minimum supplementary evidence repair subset that satisfies the first minimumity criterion is selected from the supplementary evidence repair candidate set and encapsulated into a supplementary evidence repair set. The rule repair set and the supplementary evidence repair set are encapsulated into a minimum repair set in a binary output form.

[0014] Optionally, the generation of the key information results and closed loop specifically includes: Read the field consistency conclusions, field conflict contribution decomposition results, minimum repair set, global field combination results and candidate evidence fragments, perform field-based encapsulation on the global field combination results according to the preset field system, and generate key information results; The generation of the explanation-repair package is triggered based on the field consistency conclusion. When the field consistency conclusion is consistent, an explanation-repair package containing a consistency confirmation mark and a set of field source location information is generated. When the field consistency conclusion is pending verification or inconsistent, the subsequent conflict explanation link and repair suggestion link are entered and an explanation-repair package is generated. A set of conflict explanation entries is generated based on the field conflict contribution decomposition results; Generate a set of repair suggestion items based on the minimum repair set; Based on the conflict interpretation item set and the repair suggestion item set, an interpretation-repair package is assembled and bound with key information results to generate an exportable result package. The explanation-repair package triggers a backflow update, forming a closed loop.

[0015] The beneficial effects of this invention are: This invention consolidates multi-source petition materials for the same petition matter into a petition matter workspace and forms a material package set. Based on unified text and structure standardization processing, it generates a set of evidence fragments, a set of candidate fields, and candidate evidence fragments, achieving continuous and traceable modeling of key petition information from materials to fields. Furthermore, it generates a set of evidence commitment vectors for each material, including hash aggregation commitments, field coverage summaries, structural consistency verifiable summaries, carrier quality summaries, and source distribution summaries, and outputs a commitment conflict list and coverage gap results. This transforms cross-material conflicts and gaps from implicit issues into verifiable and traceable structured inputs, reducing the uncertainty propagation of field extraction in a multi-material environment from the source.

[0016] Based on this, the present invention generates reliable posterior results and dynamic discount results through commitment-layer consistency synthesis, and performs field-layer evidence synthesis on field candidates using field-to-material discount mapping to obtain a set of field synthesis results and a decomposition result of field conflict contributions. This allows for refined suppression and interpretation of reliability differences between different sources and different fields. Simultaneously, a minimum repair consistency inference factor graph is constructed to perform global consistency inference, outputting global field combination results, field consistency conclusions, and a minimum repair set. The rule repair and supplementary evidence repair are encapsulated into an interpretation-repair package to export key information results. Then, a closed-loop update action is used to backflow and update the reliable posterior results, dynamic discount results, and minimum repair consistency inference factor graph, ensuring that the extraction of key information in petitions maintains consistency, controllability, and continuous self-correction capability even in conflict scenarios. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the machine learning-based petition information analysis and key information extraction method proposed in this invention; Figure 2 This is a flowchart illustrating the evidence commitment vector and conflict gap generation process of the machine learning-based petition information analysis and key information extraction method proposed in this invention. Figure 3 This is a flowchart of the dynamic discount and minimum repair closed-loop inference process for the machine learning-based petition information analysis and key information extraction method proposed in this invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figures 1-3 The method for analyzing and extracting key information from petition information based on machine learning includes the following steps: Collect multiple sources of petition materials for the same petition matter and merge them into a petition matter workspace to generate a material package collection; The material package set is subjected to text and structure normalization processing to generate a set of evidence fragments, and a field candidate set is further generated and bound to the candidate evidence fragments; Based on the evidence fragment set and the material package set, an evidence commitment vector set is generated for each material, and a commitment conflict list and coverage gap result are generated; Perform commitment-layer consistency synthesis on the evidence commitment vector set to generate reliable posterior results and dynamic discount results. Based on the dynamic discount results, perform field-layer evidence synthesis on the field candidate set to output the field synthesis result set and the field conflict contribution decomposition result. Based on the set of field synthesis results, the list of commitment conflicts, the coverage gap results and candidate evidence fragments, a minimum repair consistency inference factor graph is constructed and global consistency inference is performed to obtain the global field combination results, field consistency conclusions and minimum repair set; Based on the field consistency conclusions, field conflict contribution decomposition results, and minimum repair set, an explanation-repair package is generated to export key information results. The source reliable posterior results, dynamic discount results, and minimum repair consistency inference factor graph are then fed back to form a closed loop.

[0020] In this embodiment, the generation of the material package set specifically includes: Obtain multiple sources of petition materials corresponding to the same petition matter, generate material source markers, collection time markers and carrier type markers, and obtain the original set of multiple sources of petition materials; The multi-source petition materials include petition registration texts, online petition form texts, scanned images, photographs, audio and video transcription texts, processing receipt texts, and historical related materials texts; The original set of multi-source petition materials is processed to be uniformly formatted and parsable, resulting in a unified formatted set of materials. The unified formatting and parsing process includes converting scanned images and photo images into searchable text while retaining the original image carrier; unifying audio and video transcription text and processing receipt text into paragraph-based text; and converting online form text into field-based text while retaining the field name sequence. For each document in the unified format document set, generate a document fingerprint set and a content summary set; The material fingerprint is obtained by irreversibly mapping the byte sequence of the material, and the content summary is obtained by segmented sampling and normalized splicing of the paragraph sequence. The material fingerprint set and the content summary set are used for deduplication and merging determination. Deduplication and version folding are performed based on the material fingerprint set to form an effective material set and version cluster set; The deduplication and version folding include identifying materials with the same material fingerprint as duplicate materials and retaining only one copy as valid material, and identifying materials with different material fingerprints but whose content summary similarity meets the threshold condition as homologous version materials and folding them into the same version cluster. The content summary similarity is obtained by calculating the cosine similarity after vectorizing the content summary. Merge the petition items for the valid material set and version cluster set. Put the materials belonging to the same petition item into the petition item workspace according to the petition item identifier. Organize the materials into material package sets by the material source mark and carrier type mark as the grouping key and perform merge integrity verification. The material package includes the material text, material source marker, acquisition time marker, carrier type marker, material fingerprint, and version cluster marker; The merge integrity check includes verifying the consistency between the material text and the material fingerprint of each material package, the non-decreasing consistency of the acquisition time stamp within the same material package, and the traceability consistency of the material source mark within the material package. For material packages that do not meet the consistency requirements, the package is split or reorganized and the version cluster mark is updated.

[0021] In this embodiment, the generation of the candidate field set and candidate evidence fragments specifically includes: Read the material text and carrier type tag of each material package in the material package set to obtain a standardized material text set; The standardized material text set performs unified encoding, full-width and half-width character normalization, whitespace folding and punctuation regularization on paragraph text and field text, and retains the page order association with the original image carrier for searchable text; Based on the standardized material text set and the original image carrier, structural normalization processing is performed to obtain a set of structural elements; The structural normalization process includes identifying the set of region boundaries generated by the title area, body area, table area and appendix description area; parsing the table area into row and column units and generating a set of cell locations; and dividing the body area by row and merging paragraphs to generate a set of paragraph boundaries. The standardized material text set is subjected to text normalization processing to form a standardized element text set, and source location information is generated; The text normalization process includes time expression standardization, monetary expression standardization, place name level standardization, and unit name alias unification. The source location information is determined by the material package identifier, the regional boundary set, and the paragraph boundary set. Based on the set of structural elements, the standardized material text set is segmented into evidence fragments to generate candidate fragments. The scores of the evidence fragments are calculated, and the candidate fragments that meet the threshold conditions are selected and packaged into a set of evidence fragments. Generate a candidate set of fields and a candidate probability score based on a standardized set of element texts and a set of evidence fragments; The field candidate set is obtained by performing candidate extraction on each field according to a preset field system, including locating candidate segments based on field trigger words and structural anchoring rules, extracting field value segments from candidate segments based on sequence labeling models, and outputting question property candidates based on category discrimination models; Specifically, for each field value fragment, the sequence labeling model reads the positional confidence of the field label at the labeling position covered by the field value fragment and takes the average as the candidate probability score; for question nature candidates, the output confidence of the category discrimination model for the question nature category is read as the candidate probability score. Based on the candidate probability score, perform candidate normalization on the field candidate set and bind it with candidate evidence fragments, calculate the candidate confidence and encapsulate it with the corresponding evidence fragment set into candidate evidence fragments; Candidate normalization and candidate evidence fragment binding are achieved by merging candidates with the same value in the same field according to normalization values ​​and aggregating their source location information set and evidence fragment scores to generate candidate confidence scores. The candidate confidence scores are obtained by weighting and summing the candidate probability scores and the evidence fragment scores bound to the candidate according to preset fusion coefficients.

[0022] In this embodiment, the evidence fragment segmentation is performed in the order of priority for title area, priority for table row unit, and priority for body paragraph. The evidence fragment scoring is obtained by calculating the fragment completeness score, structural reliability score, and carrier quality score of the candidate fragments and then summing them according to preset non-negative weights. The fragment completeness score is obtained by statistically analyzing the key elements corresponding to the preset field system within the candidate fragments, and by weighting and synthesizing the normalized results of the number of hit elements, continuous coverage length, and proportion of missing placeholders. The structural reliability score is obtained by assigning basic weights to candidate segments based on their structural anchoring type, and then combining the normalized results of heading level consistency, table cell boundary integrity, and body paragraph boundary stability. The carrier quality score is obtained by selecting a set of quality indicators based on the carrier type label corresponding to the candidate fragment, and then weighting and synthesizing the normalized results of the mean character confidence score of the searchable text or the sharpness score, overexposure ratio score and occlusion ratio score of the original image carrier.

[0023] In this embodiment, the generation of the commitment conflict list and coverage gap results specifically includes: Read the evidence fragment set, the field candidate set, the candidate evidence fragment and the material package set, and construct the evidence commitment vector field template for each material in the material package set; The evidence commitment vector field template specifies that the evidence commitment vector includes hash aggregation commitment, field coverage summary, structural consistency checkable summary, carrier quality summary, source distribution summary, and candidate evidence fragment citation marker; Generate a hash aggregation commitment for each material in the material package set; The hash aggregation commitment obtains the text fingerprint by irreversibly mapping the main text of the material, and generates and aggregates fragment fingerprints by performing fragment fingerprint generation and aggregation mapping on the candidate evidence fragments associated with the material. The fragment fingerprint is jointly determined by the source location information of the candidate evidence fragment and the normalized value of the field value fragment. Generate a field overlay summary for each material in the material package set; The field coverage summary is obtained by counting the number of valid candidates and the number of candidate evidence fragments in the candidate set of the fields associated with the material field by field in the preset field system. Fields whose number of valid candidates meets the threshold condition are marked as covered fields, thus forming the field coverage summary. Generate a structurally consistent checkable summary for each material in the material package set; The structural consistency verifiable abstract is obtained by performing structural anchoring statistics on the source location information of the candidate evidence fragments associated with the material, including the proportion of the title area, the proportion of the table area, the proportion of the main text area, and the proportion of the appendix description area. The carrier quality stability index is obtained by statistically analyzing the carrier quality score based on the evidence fragment score, thus forming the structural consistency verifiable abstract. For each material in the material package set, generate a carrier quality summary and a source distribution summary; The carrier quality summary is obtained by statistically analyzing the carrier type marker of the material and the carrier quality score of the candidate evidence fragments, while the source distribution summary is obtained by statistically analyzing the material source marker of the material and the source location information of the candidate evidence fragments. Based on the evidence commitment vector field template, each material is assembled with hash aggregation commitment, field coverage summary, structural consistency checkable summary, carrier quality summary, source distribution summary and candidate evidence fragment citation mark to form an evidence commitment vector. The evidence commitment vector set is then generated by aggregating the material package sets. A list of commitment conflicts is obtained by performing pairwise consistency comparisons on the evidence commitment vectors of different materials within the same petition workspace. The consistency comparison includes text fingerprint inconsistency conflicts, fragment aggregation fingerprint inconsistency conflicts, field coverage summary difference conflicts, structural anchoring statistical difference conflicts, and carrier quality stability index difference conflicts. Each conflict record is bound to a corresponding candidate evidence fragment citation tag to form a traceable conflict entry, resulting in a committed conflict list. Based on the field coverage summary, the coverage field markings of all materials in the same petition workspace are summarized field by field in the preset field system to generate the coverage gap result; The gap coverage result is obtained by defining the field that is not marked as a coverage field by any material as a gap field, and encapsulating it together with the state where the corresponding candidate evidence fragment reference is empty.

[0024] In this embodiment, the generation of the field synthesis result set and the field conflict contribution decomposition result specifically includes: Read the evidence commitment vector set, commitment conflict list, field candidate set and candidate evidence fragment, and generate a commitment layer aligned input set for each evidence commitment vector in the evidence commitment vector set; The commitment layer alignment input set aligns the hash aggregation commitment, field coverage digest, structural consistency checkable digest, carrier quality digest, and source distribution digest in a unified field order to obtain a conflict pair index that is consistent with the commitment conflict list. Calculate the commitment conflict metric set based on the commitment layer aligned input set and the commitment conflict list; The commitment conflict metrics set includes fingerprint inconsistency conflict metrics, coverage difference conflict metrics, structural anchoring difference conflict metrics, and carrier quality stability difference conflict metrics, and generates conflict intensity values ​​and conflict hit tags for each type of conflict metric. The reliable posterior result of the source is calculated by performing a reliable posterior result of the source based on the set of commitment conflict metrics and the source distribution summary; The calculation of source reliability posterior results involves grouping the evidence commitment vectors corresponding to the same material source marker into the same source group, aggregating the conflict intensity values ​​within the source group, and applying priority weights to the coverage difference conflict measure and the carrier quality stability difference conflict measure to generate source risk scores and map them to source reliability posterior results. Dynamic discounting results are generated based on reliable posterior results from sources, a set of commitment conflict metrics, and coverage gap results. The dynamic discount result is a discount mapping from field to material. Different discount values ​​are generated for different fields of the same material. The discount trigger is determined only by the fingerprint inconsistency conflict measure and coverage difference conflict measure in the commitment conflict measure set, the coverage gap result, and the source reliable posterior result. The discount value corresponding to the gap field that hits the coverage gap result is downgraded, the discount value corresponding to the material that hits the fingerprint inconsistency conflict measure is downgraded, and the discount value corresponding to the material in the source group whose source reliable posterior result is lower than the reliable threshold is downgraded to obtain the dynamic discount result. Based on the dynamic discount results, perform field-level evidence synthesis on the candidate field set and output the synthesized field result set; The field-level evidence synthesis takes each field in the field candidate set as a unit, and the candidate evidence fragments for that field are grouped according to the material source. After discount suppression is performed on the candidate confidence based on the dynamic discount result, the candidate synthesis confidence is generated by aggregation. The candidates for the same field are sorted and the synthesis candidate set is selected and packaged into a field synthesis result set. Generate field conflict contribution decomposition results based on dynamic discount results and field synthesis result set; The field conflict contribution decomposition results are generated by calculating the difference in composite confidence before and after discount for the conflicting candidate pairs of the same field and tracing back to the conflict entries in the commitment conflict metric set and commitment conflict list. Conflict source contribution entries and conflict evidence fragment contribution entries are generated and aggregated according to field identifiers to form the field conflict contribution decomposition results.

[0025] In this implementation, the generation of the global field combination result, field consistency conclusion, and minimum repair set specifically includes: Read the set of field synthesis results, the list of commitment conflicts, the coverage gap results and candidate evidence fragments, and construct a minimum repair consistency inference factor graph; The minimum repair consistency inference factor graph includes a set of field variable nodes and a set of consistency constraint factors. The set of field variable nodes is established field by field by field by a preset field system and takes the synthesis candidate set in the field synthesis result set as the value domain. The set of consistency constraint factors includes commitment conflict list mapping factor, coverage gap factor and candidate evidence fragment consistency factor. Generate a candidate aligned input set based on the field synthesis result set and candidate evidence fragments; The candidate alignment input set is obtained by expanding the synthetic candidate set of each field into a candidate assignment set, and encapsulating the source location information set of candidate evidence fragments, evidence fragment scores and candidate confidence scores into candidate evidence input entries. A set of consistency constraint inputs is generated based on the mapping between the commitment conflict list and the coverage gap results; The consistency constraint input set maps the text fingerprint inconsistency conflict, fragment aggregation fingerprint inconsistency conflict, field coverage summary difference conflict, structural anchoring statistical difference conflict and carrier quality stability index difference conflict in the commitment conflict list to consistency constraint items respectively, and maps the gap field in the coverage gap result to supplementary evidence constraint items. Perform global consistency inference on the minimum repair consistency inference factor graph, calculate the consistency target quantity based on the candidate alignment input set and the consistency constraint input set, search for the field assignment combination with the minimum consistency target quantity, output the global field combination result and generate the field consistency conclusion; The consistency target quantity is obtained by weighting the candidate cost quantity jointly determined by the candidate confidence and evidence fragment score, the conflict penalty quantity determined by the hit rate and violation magnitude of the consistency constraint item, the gap penalty quantity determined by the hit rate of the supplementary evidence constraint item, and the source consistency penalty quantity determined by the material source mark and source location information set of the candidate evidence fragments, according to preset non-negative weights. The field consistency conclusions include three categories: consistent, pending further verification, and inconsistent. When the field consistency conclusion is inconsistent or requires supplementary verification, the constraint entries that lead to an increase in the consistency target quantity are extracted from the consistency constraint input set as a candidate repair entry set, and divided into a rule repair candidate set and a supplementary verification repair candidate set. The rule repair candidate set consists of consistency constraint entries mapped from the commitment conflict list, and the supplementary proof repair candidate set consists of supplementary proof constraint entries mapped from the coverage gap results. The minimum set of rule repair candidates and the minimum set of supplementary evidence repair candidates are each subjected to a minimum set of repairs.

[0026] In this embodiment, the minimumity decision takes the minimum number of repair items as the first minimumity criterion and the maximum decrease in the consistency target quantity as the second minimumity criterion. First, the minimum rule repair subset that satisfies the first minimumity criterion is selected from the rule repair candidate set and encapsulated into a rule repair set. Then, the minimum supplementary evidence repair subset that satisfies the first minimumity criterion is selected from the supplementary evidence repair candidate set and encapsulated into a supplementary evidence repair set. The rule repair set and the supplementary evidence repair set are encapsulated into a minimum repair set in a binary output form.

[0027] In this embodiment, the generation of key information results and closed loop specifically includes: Read the field consistency conclusions, field conflict contribution decomposition results, minimum repair set, global field combination results and candidate evidence fragments, perform field-based encapsulation on the global field combination results according to the preset field system, and generate key information results; The key information results include petition item identifier, field identifier, field normalized value, field source location information set, and candidate evidence fragment citation mark corresponding to the field; The generation of the explanation-repair package is triggered based on the field consistency conclusion. When the field consistency conclusion is consistent, an explanation-repair package containing a consistency confirmation mark and a set of field source location information is generated. When the field consistency conclusion is pending verification or inconsistent, the subsequent conflict explanation link and repair suggestion link are entered and an explanation-repair package is generated. A set of conflict explanation entries is generated based on the field conflict contribution decomposition results; The conflict interpretation entry set is organized by field identifier to collect mutually conflicting candidate pairs, and binds the conflict source contribution entry, conflict evidence fragment contribution entry and the traceable conflict entry in the committed conflict list to form a conflict interpretation entry set with conflict field identifier, conflict type identifier, conflict candidate pair identifier, conflict source set, conflict evidence fragment set and conflict entry traceability mark. Generate a set of repair suggestion items based on the minimum repair set; The set of repair suggestions maps the rule repair set in the minimum repair set to the set of rule repair actions, and maps the supplementary evidence repair set to the set of supplementary evidence repair actions. The set of rule repair actions is limited to three types of actions: downgrade, freeze, and rollback, and only applies to the consistency constraint factor weights in the minimum repair consistency inference factor graph. The set of supplementary evidence repair actions is limited to two types of actions: supplementary evidence request and review request, and only applies to the gap field corresponding to the gap coverage result. Based on the conflict interpretation item set and the repair suggestion item set, an interpretation-repair package is assembled and bound with key information results to generate an exportable result package. The explanation-repair package includes consistency confirmation markers, field consistency conclusions, conflict explanation entries, repair suggestion entries, and candidate evidence fragment reference markers. A closed loop is formed by triggering a reflow update based on the explanation-repair package; The backflow update includes performing four types of update actions on the source reliable posterior result based on the rule repair action set in the explanation-repair package, namely, adjusting, lowering, freezing, and rolling back, and generating an updated source reliable posterior result; performing four types of update actions on the dynamic discount result based on the conflict explanation entry set in the explanation-repair package, namely, adjusting, lowering, freezing, and rolling back, and generating an updated dynamic discount result; and performing four types of update actions on the minimum repair consistent inference factor graph based on the repair suggestion entry set in the explanation-repair package, namely, adjusting, lowering, freezing, and rolling back the consistency constraint factor weights, and generating an updated minimum repair consistent inference factor graph.

[0028] Example 1: To verify the feasibility of this invention in practice, it was applied to a joint handling scenario involving a petition reception center and district / county petition windows in a provincial capital city. In this scenario, the same petition often includes simultaneously the following: registered letters and visits, online form submissions, scanned copies and photos supplemented by the public, transcripts of telephone follow-ups, processing receipts, and historical related materials. The sources of these materials are scattered, and the versions differ significantly. Furthermore, the different materials contain inconsistencies in the description of time, location, involved units, and demands, which can easily lead to problems such as missing fields, field conflicts, and difficulties in information tracing. This results in repeated data entry and verification, inconsistent interpretations, and reliance on experience-based judgment for subsequent processing.

[0029] In this scenario, materials related to the same petition from different channels are first consolidated into a unified workspace and packaged into a single document set. The document sets are then standardized in terms of text and structure. Scanned documents and photos are converted into searchable text while maintaining page order; transcribed text and receipts are organized into paragraphs; and online forms are organized into fields. Time, amount, location hierarchy, and unit names are also standardized. Based on the standardized results, a set of evidence fragments is obtained. Evidence fragments are primarily extracted from title areas, table rows, and body paragraphs. Fragment scores are generated based on fragment completeness, structural reliability, and carrier quality. Then, a sequence labeling model and a category discrimination model are used to extract a set of candidate fields and bind them to candidate evidence fragments, ensuring that each candidate field corresponds to a verifiable evidence location and source.

[0030] When cross-material conflicts and coverage gaps exist, this invention further generates an evidence commitment vector set for each material. The evidence commitment vector includes a hash aggregation commitment, a field coverage summary, a structural consistency checkable summary, a carrier quality summary, a source distribution summary, and candidate evidence fragment citation tags. Consistency comparisons are performed on the evidence commitment vectors of different materials within the workspace to generate a commitment conflict list. Simultaneously, coverage field tags are summarized to generate coverage gap results. Subsequently, commitment-level consistency synthesis is performed on the evidence commitment vector set to obtain reliable posterior results from the source and dynamic discount results. Dynamic discounting applies differential suppression to different fields using a field-to-material mapping method. Field-level evidence synthesis is then performed on field candidates using this dynamic discount, outputting a set of field synthesis results and a decomposition result of field conflict contributions. This allows conflicts to be traced back to specific sources and specific evidence fragments, thus providing direct evidence for subsequent handling.

[0031] During the processing, for multiple key fields of the same matter, this invention constructs a minimum repair consistency inference factor graph based on the field synthesis result set, commitment conflict list, coverage gap result, and candidate evidence fragments. Global consistency inference yields the global field combination result and field consistency conclusion, generating a minimum repair set, which is output by splitting into a rule repair set and a supplementary evidence repair set. For cases where the field consistency conclusion is pending supplementary evidence, the supplementary evidence repair set explicitly points to the gap field and generates a supplementary evidence request or review request. For cases where the field consistency conclusion is inconsistent, the rule repair set provides suggestions for lowering, freezing, or reverting the weight of the consistency constraint factor, thereby preventing a single material or candidate from dominating the final conclusion in conflict situations. Finally, an integrated explanation-repair package is formed, exporting key information results. The explanation part provides conflict candidate pairs, conflict source contributions, and conflict evidence fragment contributions; the repair part provides rule repair actions and supplementary evidence repair actions, and the execution results are fed back to update the reliable posterior results, dynamic discount results, and minimum repair consistency inference factor graph. This ensures that subsequent processing of similar matters at the same location and time period exhibits more stable field consistency and fewer redundant verification operations.

[0032] Within the continuous processing cycle of the petition reception center in the provincial capital, the key information results after the materials were consolidated by the staff were easier to verify. Conflict items could be directly located to the materials and evidence fragments. The triggers for supplementary evidence and review were more focused on the missing fields. When there were inconsistencies in the standards of cross-materials, clear repair suggestions and a unified output standard could be obtained. At the same time, as the feedback and updates continued, the reliable a posteriori of the source and dynamic discount gradually converged to the suppression strategy consistent with actual handling experience. This continuously improved the accuracy, interpretability and controllability of key information extraction, meeting the handling requirements in scenarios with multiple sources of materials and frequent conflicts.

[0033] Table 1. Comparison of Key Information Extraction and Consistency Closed-Loop Effects from Multi-Source Petition Materials

[0034] In terms of extraction quality, this invention achieves a field extraction accuracy of 92.3% within the coverage area, an improvement of 5.6 percentage points compared to control A and 2.9 percentage points compared to control B; the field consistency conclusion accuracy reaches 90.5%, an improvement of 8.4 percentage points compared to control A and 4.3 percentage points compared to control B. This improvement does not stem from a single stronger model, but rather from the combination of "evidence commitment vector—commitment-layer consistency synthesis—dynamic discounting—field-layer evidence synthesis" in your claim chain: the commitment conflict list and coverage gap results explicitly reveal cross-material contradictions in advance, and the dynamic discounting results apply differentiated suppression to different fields and different materials, making the field synthesis result set less susceptible to being dragged by low-quality or conflict sources, thereby improving the stability of the final extraction and consistency determination.

[0035] From the perspective of "interpretability and resolvability," the hit rate of conflict localization to evidence fragments increased from 58.6% in control A to 88.9%, an increase of 30.3 percentage points. At the same time, the traceable output completeness rate reached 94.1%, significantly higher than 78.4% in control B. This is directly related to the output mechanism of "field conflict contribution decomposition results" and "interpretation-repair package" in the claims: the difference in confidence before and after discount is traced back to the committed conflict item, and then the candidate evidence fragment reference mark is bound, so that the conflict interpretation no longer stays at the field level of vague hints, but can stably fall to specific evidence fragments and source contributions. Therefore, it can significantly reduce the time cost of "finding materials and sentences" in manual review.

[0036] From the perspective of efficiency and closed-loop management, this invention reduces the average manual review time per item to 10.6 minutes, a reduction of 7.8 minutes compared to control A and 4.3 minutes compared to control B; the average number of reworks per item is reduced to 0.83 times, a reduction of 0.79 times compared to control A; and the closed-loop convergence period for inconsistent items is shortened to 2.4 days. This is because the minimum repair consistency inference factor graph outputs a binary separation of "rule repair set / supplementary evidence repair set": when the field consistency conclusion is pending supplementary evidence or inconsistency, the review action is compressed into a small number of constraint entries or gap fields pointed to by the minimum repair set, rather than a full review of the entire item; simultaneously, the backflow update continuously converges the weight configuration of the reliable posterior results, dynamic discount results, and the minimum repair consistency inference factor graph, enabling subsequent similar items to reach a consistent conclusion faster under the same conflict mode, reflected in the simultaneous decrease in convergence period and reversal rate (the reversal rate after review decreased from 12.6% to 7.1%).

[0037] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for analyzing petition information and extracting key information based on machine learning, characterized in that, Includes the following steps: Collect multiple sources of petition materials for the same petition matter and merge them into a petition matter workspace to generate a material package collection; The material package set is subjected to text and structure normalization processing to generate a set of evidence fragments, and a field candidate set is further generated and bound to the candidate evidence fragments; Based on the evidence fragment set and the material package set, an evidence commitment vector set is generated for each material, and a commitment conflict list and coverage gap result are generated; Perform commitment-layer consistency synthesis on the evidence commitment vector set to generate reliable posterior results and dynamic discount results. Based on the dynamic discount results, perform field-layer evidence synthesis on the field candidate set to output the field synthesis result set and the field conflict contribution decomposition result. Based on the set of field synthesis results, the list of commitment conflicts, the coverage gap results and candidate evidence fragments, a minimum repair consistency inference factor graph is constructed and global consistency inference is performed to obtain the global field combination results, field consistency conclusions and minimum repair set; Based on the field consistency conclusions, field conflict contribution decomposition results, and minimum repair set, an explanation-repair package is generated to export key information results. The source reliable posterior results, dynamic discount results, and minimum repair consistency inference factor graph are then fed back to form a closed loop.

2. The method for analyzing petition information and extracting key information based on machine learning according to claim 1, characterized in that, The generation of the material package set specifically includes: Obtain multiple sources of petition materials corresponding to the same petition matter, generate material source markers, collection time markers and carrier type markers, and obtain the original set of multiple sources of petition materials; The original set of multi-source petition materials is processed to be uniformly formatted and parsable, resulting in a unified formatted set of materials. For each document in the unified format document set, generate a document fingerprint set and a content summary set; Deduplication and version folding are performed based on the material fingerprint set to form an effective material set and version cluster set; The system merges the valid material set and version cluster set of petitions. Materials belonging to the same petition are grouped into the petition workspace according to the petition identifier. The materials are organized into material packages using the material source tag and carrier type tag as grouping keys, and the merge integrity is checked.

3. The method for analyzing petition information and extracting key information based on machine learning according to claim 1, characterized in that, The generation of the candidate field set and candidate evidence fragments specifically includes: Read the material text and carrier type tag of each material package in the material package set to obtain a standardized material text set; Based on the standardized material text set and the original image carrier, structural normalization processing is performed to obtain a set of structural elements; The standardized material text set is subjected to text normalization processing to form a standardized element text set, and source location information is generated; Based on the set of structural elements, the standardized material text set is segmented into evidence fragments to generate candidate fragments. The scores of the evidence fragments are calculated, and the candidate fragments that meet the threshold conditions are selected and packaged into a set of evidence fragments. Generate a candidate set of fields and a candidate probability score based on a standardized set of element texts and a set of evidence fragments; Based on the candidate probability score, candidate normalization is performed on the field candidate set and bound to the candidate evidence fragments. The candidate confidence is calculated and encapsulated into a candidate evidence fragment with the corresponding evidence fragment set.

4. The method for analyzing petition information and extracting key information based on machine learning according to claim 3, characterized in that, The evidence segmentation is performed in the order of priority for title area, priority for table row unit, and priority for body paragraph. The evidence segment score is obtained by calculating the fragment completeness score, structural reliability score and carrier quality score of the candidate fragments and summing them according to the preset non-negative weights.

5. The method for analyzing petition information and extracting key information based on machine learning according to claim 1, characterized in that, The generation of the commitment conflict list and coverage gap results specifically includes: Read the evidence fragment set, the field candidate set, the candidate evidence fragment and the material package set, and construct the evidence commitment vector field template for each material in the material package set; Generate a hash aggregation commitment for each material in the material package set; Generate a field overlay summary for each material in the material package set; Generate a structurally consistent checkable summary for each material in the material package set; For each material in the material package set, generate a carrier quality summary and a source distribution summary; Based on the evidence commitment vector field template, each material is assembled with hash aggregation commitment, field coverage summary, structural consistency checkable summary, carrier quality summary, source distribution summary and candidate evidence fragment citation mark to form an evidence commitment vector. The evidence commitment vector set is then generated by aggregating the material package sets. A list of commitment conflicts is obtained by performing pairwise consistency comparisons on the evidence commitment vectors of different materials within the same petition workspace. Based on the field coverage summary, the coverage field marks of all materials in the workspace of the same petition matter are summarized field by field in the preset field system to generate the coverage gap result.

6. The method for analyzing petition information and extracting key information based on machine learning according to claim 1, characterized in that, The generation of the field synthesis result set and the field conflict contribution decomposition result specifically includes: Read the evidence commitment vector set, commitment conflict list, field candidate set and candidate evidence fragment, and generate a commitment layer aligned input set for each evidence commitment vector in the evidence commitment vector set; Calculate the commitment conflict metric set based on the commitment layer aligned input set and the commitment conflict list; The reliable posterior result of the source is calculated by performing a reliable posterior result of the source based on the set of commitment conflict metrics and the source distribution summary; Dynamic discounting results are generated based on reliable posterior results from sources, a set of commitment conflict metrics, and coverage gap results. Based on the dynamic discount results, perform field-level evidence synthesis on the candidate field set and output the synthesized field result set; The results of decomposing field conflict contributions are generated based on the dynamic discount results and the field synthesis result set.

7. The method for analyzing petition information and extracting key information based on machine learning according to claim 1, characterized in that, The generation of the global field combination result, field consistency conclusion, and minimum repair set specifically includes: Read the set of field synthesis results, the list of commitment conflicts, the coverage gap results and candidate evidence fragments, and construct a minimum repair consistency inference factor graph; Generate a candidate aligned input set based on the field synthesis result set and candidate evidence fragments; A set of consistency constraint inputs is generated based on the mapping between the commitment conflict list and the coverage gap results; Perform global consistency inference on the minimum repair consistency inference factor graph, calculate the consistency target quantity based on the candidate alignment input set and the consistency constraint input set, search for the field assignment combination with the minimum consistency target quantity, output the global field combination result and generate the field consistency conclusion; When the field consistency conclusion is inconsistent or requires supplementary verification, the constraint entries that lead to an increase in the consistency target quantity are extracted from the consistency constraint input set as a candidate repair entry set, and divided into a rule repair candidate set and a supplementary verification repair candidate set. The minimum set of rule repair candidates and the minimum set of supplementary evidence repair candidates are each subjected to a minimum set of repairs.

8. The method for analyzing petition information and extracting key information based on machine learning according to claim 7, characterized in that, The minimumity criterion is based on having the fewest repair items as the first minimumity criterion and the largest decrease in the consistency target quantity as the second minimumity criterion. First, the minimum rule repair subset that satisfies the first minimumity criterion is selected from the rule repair candidate set and encapsulated into a rule repair set. Then, the minimum supplementary evidence repair subset that satisfies the first minimumity criterion is selected from the supplementary evidence repair candidate set and encapsulated into a supplementary evidence repair set. The rule repair set and the supplementary evidence repair set are encapsulated into a minimum repair set in a binary output form.

9. The method for analyzing petition information and extracting key information based on machine learning according to claim 1, characterized in that, The generation of the key information results and closed loop specifically includes: Read the field consistency conclusions, field conflict contribution decomposition results, minimum repair set, global field combination results and candidate evidence fragments, perform field-based encapsulation on the global field combination results according to the preset field system, and generate key information results; The generation of the explanation-repair package is triggered based on the field consistency conclusion. When the field consistency conclusion is consistent, an explanation-repair package containing a consistency confirmation mark and a set of field source location information is generated. When the field consistency conclusion is pending verification or inconsistent, the subsequent conflict explanation link and repair suggestion link are entered and an explanation-repair package is generated. A set of conflict explanation entries is generated based on the field conflict contribution decomposition results; Generate a set of repair suggestion items based on the minimum repair set; Based on the conflict interpretation item set and the repair suggestion item set, an interpretation-repair package is assembled and bound with key information results to generate an exportable result package. The explanation-repair package triggers a backflow update, forming a closed loop.