Method for dynamic comparison and difference insight of text content of multi-modal document
By extracting text content units and their structural location information from multimodal documents, calculating the weights of content reliability and structural importance, and identifying and clustering discrepancy locations, this method solves the problem of semantic modification intent recognition in multimodal document scenarios, and improves the accuracy and interpretability of document comparison.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU RADIO & TV UNIV
- Filing Date
- 2026-05-21
- Publication Date
- 2026-07-31
AI Technical Summary
Existing document comparison technologies cannot accurately identify semantic-level modification intentions and scope of impact in multimodal document scenarios, and are unable to identify substantial differences across documents based on the semantic vectors and structural importance weights of content units.
By acquiring text content units and their structural location information in multimodal documents, the content reliability coefficient and structural importance weight of the content units are determined. Based on semantic similarity and reliability coefficient, the difference significance index is calculated, and the units are clustered into difference groups. The comprehensive insight index is calculated, and the comparison results are output.
It achieves an intelligent upgrade from character-level difference detection to semantic-level modification intent insight, improving the accuracy and interpretability of document comparison, suppressing noise interference, reducing the probability of misjudgment, and identifying the scope of the spread and impact of modified content.
Smart Images

Figure CN122221835B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, specifically to a method for dynamic comparison and difference insight of text content in multimodal documents. Background Technology
[0002] With the deepening development of information-based office work, various documents (such as contracts, regulations, technical standards, and academic papers) often exist in multimodal forms, including text, images, tables, embedded charts, and other elements. In application scenarios such as version iteration, compliance review, and plagiarism detection, it is necessary to quickly and accurately compare the differences between different versions or documents from different sources, and to understand their semantic modification intentions. To meet this need, document comparison technology has become an important research direction in the field of information processing.
[0003] Existing document comparison technologies are usually limited to character-level or line-level comparison of plain text, and cannot identify substantial differences across documents based on the semantic vectors and structural importance weights of content units. This makes it difficult to accurately understand the true intent and scope of impact of modifications in multimodal document scenarios. Summary of the Invention
[0004] This invention provides a method for dynamic text content comparison and difference insight for multimodal documents to solve existing problems.
[0005] The present invention provides a method for dynamic text content comparison and difference insight for multimodal documents, which adopts the following technical solution: One embodiment of the present invention provides a method for dynamic text content comparison and difference insight for multimodal documents. The method includes: acquiring at least two documents to be compared, and extracting text content units and their structural position information from the documents to be compared; for each content unit, determining the content reliability coefficient of the content unit based on the positional relationship between the content unit and adjacent units across the document space, and determining the structural importance weight of the content unit based on the depth of the content unit in the document structure hierarchy; determining the difference significance index of cross-document unit pairs based on the semantic similarity between the content units, the structural importance weight, and the content reliability coefficient, and identifying the difference positions; clustering the content units corresponding to the difference positions into difference groups according to semantic relevance, and determining the comprehensive insight index of the difference groups based on the influence range of the difference groups and the semantic relevance across the groups; determining the credibility ranking based on the comprehensive insight index of each document to be compared, and outputting the comparison results.
[0006] Further, determining the content reliability coefficient of the content unit based on the positional relationship between the content unit and adjacent units across document space includes: identifying each adjacent unit spatially adjacent to the content unit in different documents, and obtaining the position number of the content unit and each adjacent unit in the document; determining the local positional correlation degree of the content unit relative to each adjacent unit based on the difference between the position numbers of the content unit and each adjacent unit, and determining a local contextual integrity score by comprehensively considering each local positional correlation degree; determining an error source penalty factor based on the error flag of the content unit during the extraction process and the error position count of the content unit; and determining the content reliability coefficient based on the local contextual integrity score and the error source penalty factor.
[0007] Further, determining the structural importance weight of the content unit based on its depth in the document structure hierarchy includes: obtaining the depth value of the content unit in the document structure tree and the maximum depth value of the structure tree, and determining the reliability proportion of the content unit; determining a structure hierarchy factor based on the depth proportion of the depth value and the maximum depth value, and the reliability proportion; wherein the structure hierarchy factor increases as the depth value decreases; determining a semantic feature factor based on the cumulative contribution of the feature weights of each semantic feature contained in the content unit relative to the feature weights of the root node; and determining the structural importance weight based on the fusion result of the structure hierarchy factor and the semantic feature factor.
[0008] Further, determining the significance index of differences between cross-document unit pairs based on the semantic similarity between the content units, the structural importance weights, and the content reliability coefficients includes: obtaining the semantic vectors of two content units from different documents to be matched, and calculating the semantic similarity between the two semantic vectors; determining the weight difference of the structural importance weights of the two content units; determining the basic difference based on the inverse mapping relationship between the semantic similarity and the weight difference; obtaining the content reliability coefficients of the two content units, and selecting the smaller one as the reliability constraint factor; and determining the significance index of differences based on the result of adjusting the basic difference by the reliability constraint factor.
[0009] Further, identifying the difference location includes: normalizing the significance index of the difference for all the cross-document unit pairs to obtain the normalized difference value corresponding to each significance index; comparing each normalized difference value with a preset threshold; and determining the cross-document unit pairs whose normalized difference value is lower than the preset threshold as difference locations.
[0010] Furthermore, the step of clustering the content units corresponding to the difference positions into difference groups according to semantic relevance includes: obtaining the semantic vector of each content unit corresponding to the difference position; and aggregating content units that are semantically similar and spatially adjacent into the same difference group based on the distribution density and proximity of each semantic vector in the semantic vector space.
[0011] Further, determining the comprehensive insight index of the difference group based on the influence range and cross-group semantic association includes: calculating the average value of the difference significance index of each content unit within the difference group to obtain the average difference intensity within the cluster; determining the character span of the difference group in the document, and determining the influence range factor based on the ratio of the character span to a preset span threshold; determining other difference groups that meet the preset semantic association conditions with the difference group, calculating the average value of the semantic similarity between the difference group and each of the other difference groups to obtain the cross-group semantic association degree; and determining the comprehensive insight index based on the comprehensive calculation result of the average difference intensity within the cluster, the influence range factor, and the cross-group semantic association degree.
[0012] Further, the step of determining the credibility ranking based on the comprehensive insight index of each of the documents to be compared and outputting the comparison results includes: ranking the comprehensive insight index of each difference group, extracting a preset number of difference groups with the highest ranking as high-value difference groups; calculating the sum of the comprehensive insight index of all difference groups in each of the documents to be compared, determining the credibility ranking of each document based on the sum, and determining the document with the highest ranking as the benchmark document; and generating a structured comparison result containing difference location identifiers and semantic level difference descriptions based on the high-value difference groups.
[0013] Furthermore, the method also includes: generating semantic-level difference description information based on the high-value difference group using a large language model; wherein the difference description information is used to characterize the modification intention and semantic impact range of the modified content corresponding to the difference group.
[0014] Furthermore, the modality of the document to be compared includes a digitized document modality and / or an image modality; The step of extracting text content units and their structural location information from the document to be compared includes: for the document to be compared in the digital document modality, directly extracting the content units and their structural location information from the document to be compared; and for the document to be compared in the image modality, performing optical character recognition to extract the text content units and their structural location information therein.
[0015] The beneficial effects of the technical solution of the present invention are: In this embodiment of the invention, at least two documents to be compared are acquired, and text content units and their structural location information are extracted from the documents to be compared. For each content unit, the content reliability coefficient of the content unit is determined based on the positional relationship between the content unit and adjacent units across the document space, and the structural importance weight of the content unit is determined based on the depth of the content unit in the document structure hierarchy. Based on the semantic similarity, structural importance weight, and content reliability coefficient between content units, the difference significance index of cross-document unit pairs is determined, and the difference positions are identified. The content units corresponding to the difference positions are clustered into difference groups according to semantic relevance, and the comprehensive insight index of the difference groups is determined based on the influence range of the difference groups and the semantic relevance across the groups. The credibility ranking is determined based on the comprehensive insight index of each document to be compared, and the comparison results are output.
[0016] This invention achieves an intelligent upgrade from character-level difference detection to semantic-level modification intent insight through a progressive processing flow of structured unified representation of multimodal content, reliability coefficient quantification, structural weight evaluation, and semantic-level difference clustering. This improves the accuracy and interpretability of document comparison. Furthermore, by calculating the content reliability coefficient based on the positional relationship of adjacent units across document space and combining it with error sources such as OCR confidence for penalty constraints, noise interference in the image-modal document extraction process is effectively suppressed, enhancing the objectivity of content unit quality assessment in multimodal scenarios. Moreover, by integrating the difference significance calculation mechanism of semantic vector similarity and structural importance weights, the inverse mapping of weight differences amplifies the impact of structural distribution inconsistencies. Simultaneously, using the minimum reliability coefficient as a constraint condition reduces the probability of misjudgment caused by text order changes due to format adjustments. Finally, by clustering difference units into difference groups and calculating a comprehensive insight index, combined with the comprehensive calculation of average difference intensity within clusters, character span influence range, and cross-group semantic correlation, modification intent with associative characteristics can be identified, revealing the propagation influence range of modified content in the document semantic network. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating the method for dynamic text content comparison and difference insight for multimodal documents provided in this application embodiment; Figure 2 This is a schematic diagram of the component syntax tree provided in an embodiment of this application. Detailed Implementation
[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following detailed description, in conjunction with the accompanying drawings and preferred embodiments, provides a specific implementation method, structure, features, and effects of the invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0021] The following description, in conjunction with the accompanying drawings, details a specific solution provided by the present invention.
[0022] like Figure 1 As shown in the embodiments of this application, a method for dynamic text content comparison and difference insight for multimodal documents is provided, including: Step S110: Obtain at least two documents to be compared, and extract the text content units and their structural location information from the documents to be compared.
[0023] The aforementioned documents to be compared refer to a set of related documents that require content consistency verification or change tracking in specific business scenarios. These include, but are not limited to, version iterations of the same text object at different points in time (such as revised and final versions of contracts and agreements), similar content files obtained from different sources (such as the main body and appendices of technical standards), or comparison objects requiring verification of originality during the review process (such as submitted drafts of academic papers and publicly available literature). During information processing, these documents are defined as semantically related data sources to be analyzed. By extracting their internal discretized text content units and their positional attributes within the document's logical structure (such as chapter level, paragraph number, and header / footer identifiers), a standardized input set is formed for subsequent difference calculations.
[0024] Optionally, the modality of the document to be compared includes a digital document modality and / or an image modality. In this case, step S110 may include: for the digital document modality document to be compared, directly extracting the content units and their structural location information in the document to be compared; for the image modality document to be compared, performing optical character recognition to extract the text content units and their structural location information therein.
[0025] The aforementioned digital document modality refers to a collection of electronic documents stored in digital encoding form and conforming to standard file format specifications. Its data structure can be directly parsed by computing devices without requiring modality conversion processing such as optical character recognition (OCR). This includes, but is not limited to, Word documents conforming to the Office Open XML specification, PDF documents conforming to the ISO 32000 standard, web page documents described by Hypertext Markup Language (HTML), and presentation formats based on the ECMA-376 standard. The aforementioned image modality refers to a collection of two-dimensional visual data acquired or scanned by optical imaging devices and stored in pixel matrix form. Its text content exists as raster images rather than stored in character encoding form. This mainly includes scanned copies of paper documents, screenshots of digital photographs, and embedded bitmap objects, requiring optical character recognition technology to convert visual symbols into machine-readable text encoding sequences.
[0026] For digital document modalities, differentiated parsing strategies can be adopted based on specific format types: For Word documents (DOCX format) based on ZIP archive structures, the Python-docx library or Apache POI component can be used to parse its document.xml part to extract paragraph nodes; for PDF documents using object flow structures, the PyMuPDF or PDFBox engine can be used to parse the page object tree to locate text blocks; for HTML documents described by markup languages, BeautifulSoup or lxml parsers can be used to traverse the DOM node tree; for presentations based on slideshow structures, the python-pptx interface can be used to extract text frames within shape objects. The content units extracted from digital document modalities include: paragraph text blocks with hierarchical relationships and their position numbers in the chapter structure; table cells defined by row and column coordinates and their cross-cell merging attributes; text box objects in floating layers and their spatial coordinates relative to the page; metadata tags for embedded image objects; and font styles, font size weights, and heading level identifiers that represent format semantics.
[0027] For image-modal documents, an optical character recognition engine can be used for text extraction: deploying the Tesseract open-source OCR engine based on an LSTM neural network model to recognize multilingual characters, or using the PaddleOCR system combined with DBNet text detection and CRNN text recognition algorithms to extract the bounding box coordinates of text regions in the image, recognition confidence scores, and corresponding character sequences. The content units extracted from the image modality include: the text character sequence generated by OCR recognition; the two-dimensional bounding box coordinates of each character or word in the original image coordinate system (usually the coordinates of the upper left and lower right corners of a rectangular area); and a confidence score value representing the recognition quality, which is calculated by the OCR engine based on the posterior probability of the character and is used to quantify the reliability of the extracted content.
[0028] It is understood that the parsing tools, parsing algorithms / models, parsing systems, and other technologies for various document formats are all mature existing technologies. For the implementation methods and working principles of each technology, please refer to the relevant technologies. The embodiments in this application will not be repeated here.
[0029] Step S120: For each content unit, determine the content reliability coefficient of the content unit based on the positional relationship between the content unit and adjacent units across the document space, and determine the structural importance weight of the content unit based on the depth of the content unit in the document structure hierarchy.
[0030] The reliability coefficient mentioned above is a normalized quantitative evaluation index, and its values are distributed as follows: The confidence level coefficient is used to characterize the fidelity and reliability of specific text content units extracted from multimodal source documents during the information conversion process. This coefficient establishes a quality mapping relationship between the extraction result and the original document, assigning a calculable confidence weight to each discrete semantic unit, thus providing an objective quality constraint benchmark in the subsequent difference detection process. It is understandable that, due to the inherent errors in the recognition of raster images by optical character recognition engines and the potential deviations in the structural reconstruction of complex document layouts by parsers, the calculation of the content reliability coefficient establishes a screening mechanism to distinguish between high-fidelity extraction units and noisy interference units. By assigning a confidence parameter reflecting the extraction quality to each content unit, this mechanism can prioritize the use of reliable semantic information during cross-document comparison, while suppressing the interference of abnormal data introduced by OCR misidentification, layout parsing errors, or image noise on the difference judgment results, thereby improving the robustness and accuracy of the overall comparison process.
[0031] In the embodiments of this application, the content reliability coefficient of a content unit can be determined by at least one of the following methods: The first approach: adopts a composite computational model that combines cross-document neighborhood position continuity evaluation with error penalty; Optionally, determining the content reliability coefficient of a content unit based on its positional relationship with adjacent units across document spaces may include: identifying each adjacent unit spatially adjacent to the content unit in different documents, and obtaining the positional indices of the content unit and each adjacent unit in the document; determining the local positional correlation of the content unit relative to each adjacent unit based on the difference between the positional indices of the content unit and each adjacent unit, and determining a local contextual integrity score by combining the local positional correlations; determining an error source penalty factor based on the error flags of the content unit during the extraction process and the error position count of the content unit; and determining the content reliability coefficient based on the local contextual integrity score and the error source penalty factor.
[0032] The aforementioned cross-document spatially adjacent units refer to the set of corresponding text units that, during multimodal document parsing, occupy the same or similar logical spatial positions as the target content unit and originate from different documents to be compared. This adjacency relationship is established based on the spatial proximity criterion of document structure, typically determined through page number identifiers, paragraph numbers, or coordinate distance thresholds, forming the local neighborhood topology of the target unit in a cross-document environment. The position number is a unique integer coordinate assigned to each content unit in the linear sequence of the document, representing the unit's absolute or relative position in the document's reading order. It is usually encoded incrementally according to the natural reading flow (top-down, left-to-right) to quantify the spatial distance between units.
[0033] Local positional relevance is a quantitative indicator that measures the spatial consistency between a target content unit and its specific neighboring units. It is obtained by calculating the absolute value of the difference between their position indices and substituting it into an exponential decay function. Local positional relevance follows the basic principle that the closer the spatial distance, the stronger the relevance, and uses a negative exponential function to map the non-linear relationship between distance and similarity. Local contextual integrity score is a comprehensive assessment of the degree to which the target unit maintains its structure across the entire cross-document neighborhood. It is obtained by aggregating the local positional relevance of all neighboring units and calculating the arithmetic mean, reflecting the layout stability of the text fragments surrounding the target unit in different data sources.
[0034] Error flags are binary identifiers set for quality defects generated during the extraction process of content units. They are set to a valid state when the optical character recognition confidence level falls below a preset threshold or parsing verification fails, marking the presence of extraction noise or recognition errors in the unit. Error location counts the number of locations within a unit marked as error-prone. To avoid division-by-zero anomalies when there are no errors within a unit, this count is set to a single value. The error source penalty factor is a normalized penalty coefficient constructed based on the above error statistics. It uses a linear mapping of one minus the error flag to the error location count, quantifying the level of uncertainty introduced during the extraction process.
[0035] The method for calculating the content reliability coefficient is as follows: ; in, Indicates the current target content unit; Representation and Unit Spatially adjacent units across document environments A set; This represents the cardinality of the set (i.e., the total number of adjacent units). and Representing units respectively and adjacent units The position number in the document. Represents the absolute value of the spatial distance between the two; This is the location attenuation constant, which is usually set to a fixed value (such as 10) to control the attenuation rate of distance sensitivity; For unit Error flags (1 for OCR confidence below the threshold, 0 otherwise. Because a low OCR confidence level forces the output of the most plausible result, these words should be avoided in reliability assessment). For unit Error location count (set to 1 when there are no errors to avoid the denominator being zero).
[0036] The calculation logic of the above formula is as follows: the first half of the formula The local context integrity score is calculated by taking the exponentially weighted average of the similarity between adjacent units, thus evaluating the cross-document preservation of the target unit's local structure. (Second half) This constitutes an error source penalty factor, which is reduced to penalize low-quality data when an extraction error occurs in a cell. The two parts are then merged through a product operation to form an original reliability assessment value that comprehensively considers both structural stability and extraction quality.
[0037] To further normalize the reliability coefficient to a standard numerical range, a Sigmoid activation function is used to perform a nonlinear transformation on the original calculation results, resulting in a final content reliability coefficient. The coefficient converges within the closed interval [0,1] and serves as the standardized confidence parameter for subsequent weighted calculations of difference detection. This coefficient performs a quality screening function in the workflow, effectively suppressing the interference of optical character recognition false alarms, layout parsing bias, and image noise on document comparison accuracy by reducing the weight contribution of low-reliability units. This ensures that cross-document difference detection prioritizes comparison and analysis based on high-fidelity semantic information.
[0038] The second approach is to employ a stability assessment mechanism based on the degree of dispersion of location distribution. In this approach, the set of neighboring units associated with the target content unit across documents is first compiled, and the positional coordinates of each neighboring unit in the document space coordinate system are obtained. Then, statistical dispersion indices of this position set, such as variance, standard deviation, or interquartile range, are calculated to quantify the degree of clustering in the spatial distribution of neighboring units. If the positional distribution of neighboring units exhibits low dispersion, it indicates that the local structure of the target unit remains stable in the cross-document environment, thus assigning a higher content reliability coefficient. Conversely, if the positional distribution is highly dispersed, it indicates potential layout parsing anomalies or cross-document alignment errors, correspondingly reducing the reliability coefficient. This approach replaces the exponential decay model with statistical distribution hypothesis testing, providing a confidence interval estimate based on probability distribution for neighborhood structure evaluation.
[0039] The third approach: Employing a neighborhood cross-validation mechanism based on semantic vector consistency; In this approach, based on establishing cross-document spatial adjacency relationships, the semantic feature vectors of each adjacent unit are further extracted, and the cosine similarity or Euclidean distance between the target unit and each adjacent unit in the semantic vector space is calculated. By analyzing the distribution characteristics of semantic similarity, such as calculating the average similarity or constructing a similarity entropy index, the semantic coherence between the target unit and its neighboring units is evaluated. When the target unit maintains high semantic similarity with most adjacent units, it indicates that the unit has semantic consistency support in the cross-document context, thus assigning a higher reliability coefficient; if semantic coherence is broken or abnormal, the reliability evaluation value is lowered accordingly. This approach combines spatial location relationships with semantic feature distribution, and enhances the semantic sensitivity of reliability judgment through semantic consistency cross-validation.
[0040] The aforementioned structural importance weight is a normalized evaluation parameter based on the document's logical hierarchy and semantic feature distribution, used to quantify the relative importance of a specific text content unit within the overall document semantic network. The structural importance weight comprehensively considers the content unit's topological position in the document's tree structure and the cumulative contribution of its semantic feature set. By fusing physical location attributes with semantic value attributes, it forms a calculable index that reflects the information value of the unit. In a multimodal document processing framework, the structural importance weight serves as a quantitative benchmark for distinguishing core semantic carriers from peripheral content, providing a structured priority evaluation basis for difference detection algorithms.
[0041] It is understandable that in complex document structures, the contribution of text units at different locations to the overall semantic expression varies significantly. For example, chapter titles typically carry summary core information, while functional words in the main body paragraphs are semantic auxiliary components. By calculating structural importance weights, a semantic value assessment system based on the document's logical hierarchy can be established, assigning higher weights to high-value units located in shallow structures (such as titles and abstracts) and relatively lower weights to transitional content in deeper structures. This differentiated weight allocation mechanism ensures that changes to key structural content are prioritized and evaluated during cross-document comparisons, avoiding the obscuring of substantive semantic modifications by formatting disturbances or minor adjustments to peripheral content. The introduction of structural importance weights provides structural dimension constraints for calculating the significance of differences, enabling the difference detection results to reflect the importance level of the modified content within the document's semantic system. When corresponding units in two documents being compared show significant differences in structural importance weights, it indicates that a structural adjustment involving core semantics (such as a change in title level or replacement of key terms) may have occurred at that location, rather than a simple character-level editing operation. This structural weight-based differential processing mechanism effectively improves the document comparison system's sensitivity to substantial semantic changes and its robustness to format noise, making the output results more consistent with human readers' cognitive patterns of document importance.
[0042] In the embodiments of this application, the structural importance weight of a content unit can be determined by at least one of the following methods: The first approach is to adopt a computational framework that uses a hierarchical evaluation based on the coupling of depth proportion and reliability proportion, and combines the cumulative contribution of semantic features for weighted fusion. Optionally, the above method of determining the structural importance weight of a content unit based on its depth in the document structure hierarchy includes: obtaining the depth value of the content unit in the document structure tree and the maximum depth value of the structure tree, and determining the reliability proportion of the content unit; determining the structure hierarchy factor based on the depth proportion and reliability proportion of the depth value and the maximum depth value; wherein the structure hierarchy factor increases as the depth value decreases; determining the semantic feature factor based on the cumulative contribution of the feature weights of each semantic feature contained in the content unit relative to the feature weights of the root node; and determining the structural importance weight based on the fusion result of the structure hierarchy factor and the semantic feature factor.
[0043] The document structure tree described above is a hierarchical topological data structure generated by formally modeling the logical hierarchy of the document using constituent parsing (CON tool). Its root node corresponds to the document's top-level heading or topic identifier, intermediate nodes represent the nesting relationships between chapters or subheadings at various levels, and leaf nodes correspond to paragraph units carrying specific text content. Each node in the tree establishes a one-to-one mapping relationship with a specific content unit, and the parent-child connections between nodes reflect the document's outline hierarchy. The depth value is an integer parameter assigned to the hierarchical position of each node (i.e., the corresponding content unit) in the tree. It starts at zero with the depth value of the root node and increments by one for each edge encountered while traversing down the tree branches. This parameter quantifies the relative hierarchical position of the unit within the document's logical structure. The maximum depth value is the supremum of the set of depth values for all nodes in the document structure tree, representing the deepest level of the document structure tree, and is used to normalize absolute depth values to a relative depth percentage.
[0044] Constituent syntax trees are tree-like data structures generated by hierarchically modeling natural language sentences through constituency parsing. They reveal the bottom-up combinatorial relationships of lexical units through recursively nested phrase structures. For example... Figure 2As shown, the constituency syntactic tree takes the sentence "HanLP is a natural language processing toolkit for production environments." as the analysis object. In the figure, the Token column on the left lists the smallest grammatical units that make up the sentence (including the lexical items "HanLP", "is", "for", and the punctuation mark "."). The PoS column annotates the词性 categories of each lexical item (e.g., NR marks proper nouns, VC marks copulative verbs, NN marks nouns, and PU marks punctuation marks). The right tree diagram shows the hierarchical merger process from words to phrases and from phrases to clauses by connecting each node with directed line segments. The non-terminal nodes in the tree represent the grammatical function categories of phrases, including noun phrases (NP), verb phrases (VP), complement phrases (CP), and sentence-level components (IP), etc. Taking the noun phrase "natural language processing toolkit" in the figure as an example, the words "natural", "language", "processing", and "toolkit" are all terminal nodes of the NN词性. Through hierarchical merger, they form a lower-level NP node, and then combine with the marker word "的" (DEC) to construct a higher-level NP structure, which is finally embedded under the IP sentence root node with "HanLP" as the subject and "is" as the copulative verb. This hierarchical nesting relationship establishes the parent-child dependency relationship between nodes: the IP node is the root node with a depth value of zero; its direct child nodes NP (subject) and VP (predicate) have a depth value of one; the IP (object clause / complement) node subordinate to VP has a depth value of two, and so on. The depth value of the leaf node (lexical item) is determined by the number of path edges from the root node to this leaf node. The document structure tree is isomorphic in form to the above constituency syntactic tree, and both use tree topology to represent hierarchical composition relationships. In the structure tree at the document level, the root node corresponds to the top-level title or theme identifier of the document (equivalent to the IP root node in the syntactic tree), the intermediate nodes represent each level of chapters (equivalent to phrase nodes such as NP / VP), and the leaf nodes correspond to the paragraph units carrying specific text content (equivalent to lexical terminal nodes). The quantization rule for the depth value of the document structure tree corresponding to the syntactic tree depth calculation method shown in the attached figure (the path length from the root node IP to the leaf node "."): The depth of the root node is defined as zero, and the depth value increases by one for each parent-child edge traversed downward along the tree branch. This hierarchical depth parameter provides a topological position benchmark for calculating the structural importance weight, enabling nodes in the shallow layer of the structure tree (such as the direct child nodes of IP in the attached figure) to obtain a higher importance weight assignment in document-level analysis, while deep nodes (such as lexical items nested within multiple phrases) obtain a relatively lower weight, thus achieving a differential evaluation of core structural content and marginal subsidiary content during document comparison.
[0045] The aforementioned reliability ratio is the ratio between the content reliability coefficient of a content unit and the content reliability coefficient of the root node, reflecting the relative level of the target unit's quality reliability compared to the document's baseline reference point. When this ratio approaches one, it indicates that the extraction quality of the target unit is comparable to that of the document's top-level structural units, possessing a high level of reliability. When this ratio is significantly less than one, it indicates that the target unit has relatively more extraction noise or recognition uncertainty. The depth ratio is the ratio of the content unit's depth value to the maximum depth value of the tree structure, representing the relative position depth of the unit within the overall document hierarchical structure. Its value ranges from zero to one in a closed interval; a smaller value indicates that the unit is located in a shallower position within the tree structure.
[0046] The aforementioned structural hierarchy factor is a normalized evaluation parameter calculated based on the coupling of depth and reliability proportions and mapped using a nonlinear activation function. Its mathematical characteristic is that it monotonically increases as the depth value decreases; that is, units in the shallow layers of the tree structure (such as first-level headings and chapter topics) receive higher factor values, while units in the deeper layers (such as body paragraphs and footnotes) receive lower factor values. The calculation of the structural hierarchy factor can be performed using a Sigmoid activation function, taking the product of depth and reliability proportions as the input parameter. This is mapped onto an open interval of zero to one using an S-curve, preserving the relative differences in the input parameters while achieving a smooth constraint on the numerical range.
[0047] The aforementioned semantic feature factors are weighted sums calculated based on the cumulative contribution of each semantic feature contained in a content unit. Semantic features refer to lexical or phrase units with specific semantic functions or domain significance extracted from the text content of a unit, including but not limited to professional terms, named entities, keywords, and fixed collocations. Each semantic feature is assigned a feature weight representing its informational importance or domain specificity. This weight can be pre-determined based on domain dictionaries, corpus statistical frequencies, or attention weights of pre-trained language models. The root node feature weight is the sum of the feature weights of the semantic features contained in the top-level structural unit of the document, serving as a normalization benchmark. The cumulative contribution of the feature weights relative to the root node feature weights is obtained by summing the sums of each feature weight divided by the root node feature weight, reflecting the richness of the target unit in terms of semantic information density relative to the core theme of the document.
[0048] The method for calculating structural importance weights is as follows: ; in, Indicates the current content unit; Representation unit Depth value in the document structure tree (root node depth is defined as 0); This represents the maximum depth value of the current tree structure, used for depth normalization. This represents the reliability coefficient of the root node's content. Representation unit Content reliability coefficient, ratio The proportion of reliability; Representation unit The set of semantic features it contains; Representation of features The sum of the weights of each word node (for example, the feature "production environment" may contain the sum of the weights of nodes with "production environment" = 0.8, "production" = 0.5, and "oriented" = 0.3). The feature weights of the root node are represented by the summation term. Characteristic factors that represent semantic features.
[0049] The antecedent of the formula As a structural hierarchy factor, its working principle is based on the following logic: when the unit is in the shallow layer of the structure tree ( (Smaller value) and higher reliability ( near When inputting parameters When the value approaches zero or a small positive value, the output of the Sigmoid function approaches 0.5 or higher, indicating that the element has high structural importance; when the element is in a deep layer ( When the value is large or the reliability is low, the absolute value of the input parameter increases, and the Sigmoid output decreases accordingly. (The following term...) As a semantic feature factor, the amount of semantic information carried by a unit is quantified by accumulating the normalized weight contributions of each semantic feature. Two factors are fused through a product operation to form a structural importance weight that comprehensively considers both topological hierarchical position and semantic content value. The numerical distribution reflects the relative importance level of the unit in the document semantic network, providing a quantitative basis for distinguishing core content from peripheral content in subsequent difference detection.
[0050] The second approach is to use a weighting mechanism based on heuristic rules and style feature extraction. In this approach, the formatting features of content units can be directly extracted by parsing the document's style description layer (such as styles.xml in a Word document or a font description dictionary in a PDF). These features include, but are not limited to, heading level tags (Heading 1, Heading 2, etc.), font weight attributes (Bold weight), font size parameters, and color identifiers. Based on a pre-defined heuristic rule base, fixed weight bases are assigned to different combinations of formatting features. For example, the highest weight coefficient is assigned to first-level headings, the second highest weight coefficient is assigned to bold emphasis text, and a baseline weight coefficient is assigned to regular body text. Furthermore, by combining the unit's physical position in the document (such as whether it is located at the beginning, end, or top of the page), the weight base is adjusted using a position correction factor to generate the final structural importance weights.
[0051] The third approach: adopting an unsupervised evaluation mechanism based on information entropy and statistical significance; In this approach, word frequency statistics are first performed on all content units to calculate the information entropy value of each unit, assessing the dispersion of its word distribution. Units with high information entropy (containing diverse technical terms) are considered to have higher semantic importance. Simultaneously, based on the positional distribution of all units within the document, a statistical significance score for the unit's position is calculated (e.g., inverted frequency based on paragraph position), with units located at the beginning of the document receiving higher significance scores. Then, the information entropy values and positional significance scores are weighted and summed or multiplied to construct structural importance weights. This approach does not rely on predefined tree structures or style rules and is suitable for free documents with variable structures or non-standard formats, automatically identifying key content units with high information content through statistical regularities.
[0052] Step S130: Based on the semantic similarity, structural importance weight, and content reliability coefficient between content units, determine the significance index of differences between cross-document unit pairs and identify the location of differences.
[0053] The aforementioned significance index is a comprehensive metric used to quantify the degree of content deviation between two content units from different documents to be compared. By integrating similarity measures in the semantic vector space, importance weight differences in the document's logical structure, and quality reliability constraints during information extraction, the significance index constructs a multi-dimensional evaluation system to characterize the overall deviation level of the corresponding unit's content consistency and structural alignment in a cross-document environment. After normalization, the index is mapped to a specific numerical range, and its value directly reflects the degree of difference between the unit pairs—a higher value indicates a more significant semantic deviation and structural mismatch, while a lower value indicates a greater tendency towards content consistency. By setting a judgment threshold, the index is further transformed into a basis for identifying the location of differences, realizing a mapping from continuous measurement to discrete classification.
[0054] In the above scheme, the core purpose of determining the significance index is to overcome the limitations of traditional character-level or line-level comparison methods and achieve accurate identification and noise filtering of substantial semantic differences between documents. Traditional methods rely only on surface matching of strings and cannot penetrate the lexical surface to capture deep semantic equivalence relationships (such as synonym replacement or sentence restructuring), nor can they distinguish between text order changes caused by format adjustments and true semantic modifications. By introducing a similarity measure based on semantic vectors, the significance index can identify semantically equivalent unit pairs with different expressions, avoiding misjudging minor wording adjustments as substantial changes, and capturing pseudo-consistency phenomena of similar characters but semantic divergence, thereby improving the semantic sensitivity of difference detection. Introducing structural importance weights as an evaluation dimension allows the significance index to differentiate units based on their importance in the document's logical hierarchy. In cross-document comparison, changes in units located in the shallow layers of the tree structure (such as chapter titles) usually mean modifications to the core theme of the document, while changes in deep units (such as body text annotations) may only involve adjustments to peripheral information. By incorporating differences in structural weights into the index calculation, structural changes involving core semantic carriers can be prioritized, reducing the sensitivity to format disturbances in peripheral content. This makes the difference detection results more consistent with the semantic importance hierarchy of documents, enhancing the interpretability and business relevance of the comparison results.
[0055] Furthermore, by incorporating the content reliability coefficient as a constraint parameter, the difference significance index possesses the ability to suppress interference from low-quality extracted data. During multimodal document parsing, optical character recognition errors or layout parsing deviations may generate noisy units, which would lead to false judgments of differences if directly compared. By using the reliability coefficient as a modulating factor, this index reduces the weight contribution of units with insufficient source credibility, ensuring that difference determination is primarily based on high-fidelity semantic information for comparative analysis. This effectively avoids attributing extraction errors to actual content differences between documents, thereby improving the robustness and accuracy of cross-modal comparison.
[0056] In the embodiments of this application, the significance index of differences across document unit pairs can be determined by at least one of the following methods: The first approach is to use a composite evaluation model that combines the basic difference calculation based on the inverse mapping between semantic vector space similarity and structural weight difference with the minimum reliability coefficient constraint. Optionally, the above-mentioned determination of the significance index of differences between cross-document unit pairs based on semantic similarity, structural importance weight, and content reliability coefficient between content units includes: obtaining the semantic vectors of the two content units from different documents to be matched, and calculating the semantic similarity between the two semantic vectors; determining the weight difference of the structural importance weights of the two content units; determining the basic difference based on the inverse mapping relationship between semantic similarity and weight difference; obtaining the content reliability coefficients of the two content units, and selecting the smaller one as the reliability constraint factor; and determining the significance index of differences based on the result of adjusting the basic difference by the reliability constraint factor.
[0057] The aforementioned semantic vectors map discrete text words to a high-dimensional mathematical representation in a continuous real-valued vector space using a distributed word embedding model. Each dimension reflects the distribution strength of a word on a specific semantic attribute. In cross-document comparison scenarios, Word2vec tools (including CBOW or Skip-gram architectures) are used to vectorize the text sequence of content units, generating dense vectors with fixed dimensions (typically 384, 768, or 1024 dimensions). and This representation preserves the semantic connections between words, allowing semantically similar units to have a small spatial distance in the vector space.
[0058] Semantic similarity is a metric that quantifies the directional consistency between two semantic vectors in the feature space, and it uses the cosine similarity function. The calculation is performed, and its mathematical definition is the ratio of the dot product of two vectors to the product of their respective magnitudes. Semantic similarity values range from the closed interval [−1, 1]. Values closer to 1 indicate greater semantic similarity between the two units, values closer to -1 indicate opposite semantics, and values close to 0 indicate semantic irrelevance. In document difference detection applications, high cosine similarity corresponds to semantic equivalence between content units, even if there are differences in wording or sentence restructuring.
[0059] The weight difference is an absolute measure of the degree of deviation between the importance levels of two content units in the document's logical structure. It is calculated by taking the absolute value of the difference between the structural importance weights of the two units. This difference quantifies the asymmetry in hierarchical depth and semantic value distribution between corresponding units across documents. A difference of zero indicates that the two units are at the same level of structural importance, while an increasing difference indicates a structural mismatch (such as a mismatch between a title and body text). During the calculation process... and These represent the units from document A. and units from document B The structural importance weights are calculated based on the depth of the tree structure and the contribution of semantic features, as described in the preceding steps.
[0060] The basic dissimilarity score is a preliminary assessment that integrates semantic similarity and structural weight differences. Its calculation follows an inverse mapping relationship: the higher the semantic similarity and the smaller the structural weight differences, the lower the basic dissimilarity score. This mapping is achieved by dividing the cosine similarity by the sum of the weight differences, i.e. The addition of one in the denominator ensures numerical stability and controls the upper bound of the basic difference. When the weight difference is zero, the basic difference equals the semantic similarity. When the weight difference increases, the denominator grows linearly, causing the basic difference to decay, thereby suppressing false difference judgments caused by structural hierarchical misalignment.
[0061] The reliability constraint factor is a parameter characterizing the assurance of overall information quality across document units. It is determined by selecting the smaller value between the content reliability coefficients of two content units. Confirmed. This selection strategy follows the principle of the weakest link, meaning that the quality reliability of a matching pair is limited by the less reliable unit. Even if a single unit has an extremely high content reliability coefficient, the overall matching result should still be subject to quality constraints if its paired unit has extraction noise or recognition errors. and Representing units respectively and The content reliability coefficient is calculated based on cross-document neighborhood position continuity and error penalty in the aforementioned steps, and its value ranges from [value missing]. .
[0062] The method for calculating the significance index of difference is as follows: ; in, Indicates a cell from document A With the unit from document B The significance index of the difference; and These represent the semantic vectors of the two units respectively; Cosine similarity between two semantic vectors; and These represent the structural importance weights of the two units, respectively. Represents the absolute difference in structural importance weights; and These represent the content reliability coefficients of the two units, respectively. This means taking the smaller of the two reliability coefficients as the constraint factor.
[0063] In the formula for calculating the significance index of the differences mentioned above, the cosine similarity of the numerator reflects the degree of semantic equivalence, the denominator suppresses structural misjudgments caused by format adjustments through the amplification effect of structural weight differences, and the minimum reliability constraint of the multiplier ensures that low-quality extraction units do not dominate the difference judgment results. When two units are highly similar in semantics, have the same structural hierarchy, and both have high reliability, A value close to 1 indicates extremely low significance (content consistency); when there is semantic deviation, structural misalignment, or insufficient reliability, The values are reduced accordingly, and if they fall below a preset threshold after normalization, they are identified as differences. This computational architecture realizes a paradigm shift in difference detection from the character surface level to the semantic depth level, and from single-dimensional to multi-dimensional constraints, effectively distinguishing between substantive content changes and format noise interference.
[0064] The second approach: adopting a difference measurement mechanism based on weighted multidimensional feature space distance; In this approach, semantic similarity, structural importance weights, and content reliability coefficients can be mapped to a unified multidimensional evaluation space, constructing feature components for semantic, structural, and reliability dimensions respectively. In the semantic dimension, cosine similarity or Euclidean distance is used to measure the spatial proximity of two semantic vectors; in the structural dimension, normalized weighted differences or relative entropy are used to measure the deviation in importance levels; and in the reliability dimension, the harmonic mean or geometric mean of the reliability coefficients is used to measure the quality assurance level of the matching pair. Furthermore, preset weight coefficients are assigned to each dimension component to reflect its relative importance in difference assessment. A comprehensive difference value is calculated through weighted summation or weighted Euclidean distance, and finally, a sigmoid or linear normalization function is used to map the comprehensive difference value to the zero-to-one interval to obtain a difference significance index. This implementation provides flexibility in the evaluation strategy through explicit dimension weight configuration, allowing the priority of semantic matching and structural alignment to be adjusted according to the application scenario.
[0065] The third approach: employing an uncertainty reasoning mechanism based on a probabilistic graphical model; In this approach, significant differences can be modeled as latent variables, constructing a Bayesian network structure that includes semantic similarity nodes, structural weight difference nodes, reliability coefficient nodes, and significant difference nodes. A conditional probability table describes the probabilistic dependencies between each observation node and the latent variables, such as setting a low probability of high significant differences under high semantic similarity conditions. Based on the actual observations of unit pairs, Bayesian inference is used to calculate the posterior probability distribution of the significant difference latent variables, and then the expected value or maximum posterior probability value is taken as the output of the significant difference index. This implementation introduces an uncertainty quantification mechanism, which can generate difference assessment results with confidence intervals rather than deterministic point estimates. It is suitable for high-precision comparison scenarios that are sensitive to the risk of difference determination, and naturally integrates the uncertainty propagation of multi-source heterogeneous features through probabilistic inference.
[0066] Optionally, the above-mentioned identification of difference locations includes: normalizing the significance index of all cross-document unit pairs to obtain the normalized difference value corresponding to each significance index; comparing each normalized difference value with a preset threshold; and determining cross-document unit pairs with normalized difference values lower than the preset threshold as difference locations.
[0067] The normalized difference value mentioned above is a standardized evaluation parameter obtained by performing a numerical range scaling transformation on the difference significance index. Its value is mapped to a specific unified interval (usually a closed interval of zero to one). This transformation process aims to eliminate the inconsistency of dimensions caused by factors such as differences in semantic vector distribution and structural weight cardinality in different pairs of documents or different batches of calculations, making the difference evaluation results across documents and batches comparable. In implementation, a linear normalization method based on extreme values can be used, which calculates the ratio of the difference between the current difference significance index and the minimum value of all indices to the range of all indices (the difference between the maximum and minimum values) to achieve linear compression of the numerical interval; or a standard score transformation based on statistical distribution can be used to map the original index to a numerical domain that follows a standard normal distribution; alternatively, nonlinear normalization based on the Sigmoid function can be used, utilizing the saturation characteristics of the S-curve to smooth extreme values and enhance the sensitivity of intermediate segmentation.
[0068] The preset threshold is a pre-set decision boundary parameter for achieving binary classification of differences. This parameter serves as the critical value for distinguishing between pairs of identical and differing content units. Its value can be set based on the statistical distribution characteristics of the validation dataset (such as the mean minus a certain number of standard deviations, or the optimal split point determined based on the ROC curve), or it can be manually calibrated according to the sensitivity requirements of difference detection in specific business scenarios (for example, setting a higher threshold in a precision-priority scenario to suppress false positives, or setting a lower threshold in a recall-priority scenario to avoid missed detections).
[0069] The determination of discrepancy locations follows a threshold-based classification logic, identifying cross-document unit pairs with normalized difference values below a preset threshold as discrepancy locations. This determination rule works based on the inherent semantics of the difference significance index: this index essentially characterizes the comprehensive similarity of two units in semantic vector space, encompassing both proximity and structural alignment. Higher values indicate greater content consistency, while lower values indicate more significant semantic deviations or structural mismatches. Therefore, when the normalized difference value is below the preset threshold, it indicates that the comprehensive similarity of the unit pair is insufficient to support the assumption of content consistency, thus marking it as a discrepancy location and incorporating it into subsequent clustering analysis and insight index calculation processes. Unit pairs with values above or equal to the threshold are filtered out as non-discrepancy content, thereby achieving a transformation from continuous measurement to discrete difference labels, effectively filtering out substantive change points requiring manual review or further processing.
[0070] Step S140: Cluster the content units corresponding to the difference positions into difference groups according to semantic relevance, and determine the comprehensive insight index of the difference groups based on the influence range of the difference groups and the semantic relevance across groups.
[0071] The aforementioned Comprehensive Insight Index is a comprehensive metric used to quantify the importance and scope of influence of differing groups in the document's semantic network. This index constructs a multi-dimensional value assessment system by integrating the mean difference intensity within a group, the group's character span proportion in the document's physical space, and the semantic association density between the group and other differing groups. Unlike examining isolated changes to individual characters or words, the Comprehensive Insight Index aggregates discrete differing units into semantically complete group entities, assessing their overall impact on the document's thematic consistency, logical coherence, and information integrity, thereby characterizing the profound influence of specific modification operations on the document's semantic architecture.
[0072] In the aforementioned approach, the core purpose of calculating the comprehensive insight index is to achieve a paradigm shift from listing character-level differences to understanding semantic intent, accurately identifying modifications with substantial business value. In complex document comparison scenarios, scattered difference units may originate from the same editorial intent (such as terminology uniformity replacement, clause reorganization, or theme migration). Traditional methods only present isolated lists of differences, failing to reveal the systemic connections and underlying motivations of the modifications. By introducing the comprehensive insight index, we can identify groups of differences that are semantically mutually supportive and spatially continuous, distinguishing between accidental typo corrections and strategically intended structural adjustments. This allows the comparison results to focus on the set of changes that truly impact the core meaning of the document, enhancing the business interpretability and decision support value of the difference detection report.
[0073] Optionally, the above-mentioned clustering of content units corresponding to the difference positions into difference groups according to semantic relevance includes: obtaining the semantic vector of each content unit corresponding to the difference position; and aggregating content units that are semantically similar and spatially close into the same difference group based on the distribution density and proximity of each semantic vector in the semantic vector space.
[0074] The aforementioned difference groups are semantic association sets formed by aggregating content units corresponding to isolated difference locations through cluster analysis. They represent a set of modification operations with spatial continuity or semantic consistency within a document. Units within this group share similar semantic feature distributions and exhibit proximity in the document's physical location, reflecting the editor's intention to make batch modifications to the same or related semantic units, rather than isolated character-level errors or minor formatting adjustments. By aggregating scattered difference units into difference groups, patterns and structures of document modifications can be identified at a higher level of abstraction, laying the foundation for subsequent assessment of the systematic impact of modifications.
[0075] Each content unit corresponding to a difference location has a semantic vector representation pre-computed by the Word2vec tool. This vector is a high-dimensional real vector (the dimension is usually set to 384, 768, or 1024), and each dimension component encodes the distribution characteristics of the unit in terms of lexical semantics, syntactic structure, and contextual relevance. In the clustering preparation stage, these semantic vectors are obtained to construct a feature matrix, with each row corresponding to a difference unit and each column corresponding to a dimension of the semantic space, forming a standardized input dataset for density clustering analysis.
[0076] In this embodiment, clustering is performed based on the distribution density and proximity relationships in the semantic vector space. The density-based spatial clustering algorithm (DBSCAN) can be used as an implementation. This algorithm identifies high-density core points and their density-reachable regions in the semantic vector space by defining a neighborhood radius parameter (Epsilon) and a minimum sample point parameter (MinPts). For any dissimilar unit, if the number of semantic vectors contained within its neighborhood radius is not less than the minimum sample point threshold, the unit is marked as a core point. Non-core points located within the neighborhood radius of a core point are merged into the same cluster. Units that are neither core points nor reachable by the density of core points are determined as noise points. The DBSCAN clustering mechanism does not predetermine the number of clusters and can adaptively discover dense regions of arbitrary shapes in the semantic space, effectively handling the problem of assigning dissimilar units with ambiguous boundaries. It is understood that the above-mentioned DBSCAN clustering is a mature clustering technology; its specific implementation and working principle can be found in related technologies, and will not be elaborated further in this embodiment.
[0077] Optionally, the above-mentioned comprehensive insight index for determining the difference group based on the influence range of the difference group and the semantic association across groups includes: calculating the average value of the difference significance index of each content unit within the difference group to obtain the average difference intensity within the cluster; determining the character span of the difference group in the document, and determining the influence range factor based on the ratio of the character span to a preset span threshold; identifying other difference groups that meet the preset semantic association conditions with the difference group, calculating the average semantic similarity between the difference group and each other difference group to obtain the cross-group semantic association degree; and determining the comprehensive insight index based on the comprehensive calculation results of the average difference intensity within the cluster, the influence range factor, and the cross-group semantic association degree.
[0078] The aforementioned intra-cluster average difference intensity is a statistical measure characterizing the central tendency of deviation among content units within a difference group. It is obtained by calculating the arithmetic mean of the significance indices of differences among all units within the group. The intra-cluster average difference intensity reflects the content deviation intensity exhibited by a specific difference group as a whole in cross-document comparison. A higher average value indicates that the semantic deviation or structural mismatch among units within the group is more consistent and significant, suggesting that the group corresponds to a substantial modification operation with a clear intent rather than random noise. In addition to the arithmetic mean, a weighted average (with the content reliability coefficient as the weight) can be used to enhance the contribution of high-fidelity units, or the median can be used to suppress the interference of extreme outliers.
[0079] Character span is a linear metric that quantifies the extent to which a group of differences covers the physical space of a document. It is defined as the number of characters corresponding to the difference between the maximum and minimum position numbers of all content units within the group in the document's reading order. Character span quantifies the vertical extension of modifications within a document; groups with larger spans typically correspond to chapter-level or paragraph-level structural adjustments, while groups with smaller spans may correspond to local word replacements or sentence corrections. In its calculation, character span can be based on either raw character counts (including punctuation and spaces) or effective word counts (the number of tokens after word segmentation), depending on the document type and the granularity requirements of the business scenario.
[0080] The scope of influence factor is a relative metric obtained by normalizing the character span to a preset span threshold, which is typically set based on document layout characteristics (such as the character capacity of a standard page or the total vocabulary of a fixed-length document). The calculation of the scope of influence factor follows a linear proportional mapping principle: when the character span approaches or exceeds the preset threshold, the factor tends to be one, indicating that the scope of influence of this difference group reaches or exceeds the typical page size; when the character span is small, the factor decreases accordingly, reflecting the locality of the modification. In addition to linear normalization, a logarithmic function-based compression mapping can be used to adapt to the evaluation needs of very long documents, or the continuous span can be discretized into a finite number of levels (such as local, paragraph, chapter, and full-text levels) based on quantile level division.
[0081] Preset semantic association criteria are used to determine whether a semantic correspondence exists between two dissimilar groups. These are typically set as follows: the semantic similarity between the central vectors of the two groups is higher than a specific threshold (e.g., cosine similarity greater than 0.6), or based on density reachability (e.g., the core unit of one group is located within the neighborhood radius of another group). Groups that meet the preset semantic association criteria are included in the semantic association set of the current group. This indicates that these groups may originate from different manifestations of the same editorial intent (e.g., contextualized modifications triggered by terminology replacement) or a series of adjustments with logical dependencies (e.g., simultaneous revisions of preconditions and conclusion clauses). The determination of this criterion can employ a hard threshold dichotomy or a probability-based fuzzy membership calculation, allowing the association strength to be continuously distributed between zero and one.
[0082] Cross-group semantic relevance is a high-order measure of the semantic tightness between sets of discriminative and related groups. It is obtained by calculating the arithmetic mean of the cosine similarity between the current group and the central vectors of each related group, and applying a gain offset (such as adding one) to the average to reflect the positive contribution of the association itself. Cross-group semantic relevance can reflect the systematic impact of a specific modification operation on the document's semantic network. A higher relevance indicates that the change forms a semantically resonant network with other modifications, which may involve deeper editorial intentions such as topic migration or conceptual system reconstruction. In optional implementations, in addition to the arithmetic mean, a weighted average (with the comprehensive insight index of the related group itself as the weight) can be used to highlight the impact of high-value associations, or a maximum value operation can be used to characterize the constraining effect of the strongest association.
[0083] The center vector is the centroid vector representing the overall semantic feature distribution of a dissimilar group. It is obtained by calculating the arithmetic mean of the semantic vectors of all content units within the group. This vector is located at the geometric center of the high-dimensional semantic space and can represent the comprehensive features of the group in terms of lexical semantics, syntactic structure, and contextual association. As the basic representation for calculating semantic similarity between groups, the center vector effectively aggregates the individual semantic attributes of group members, eliminates internal variation noise, and provides a stable feature anchor for determining cross-group semantic association. In alternative implementations, the center vector can also be a density-based core point vector (core samples in the DBSCAN algorithm) or a median vector based on the geometric median to enhance robustness to anomalous units within the group.
[0084] The calculation method for the Comprehensive Insight Index is as follows: ; in, Indicates a differential group (differential cluster); Representation unit The significance index of the difference; Indicates group Total number of contained units, summation terms This represents the cumulative value of the significance index of all units within the group, and its ratio to the total number of units constitutes the average difference intensity within the cluster. Indicates group The character span in the document (the number of characters obtained by subtracting the start position from the end position). This represents the preset span threshold (set according to the length of the text, such as 200 words per page), and the ratio of the two constitutes the influence range factor. Indicates group Other differential groups that satisfy the preset semantic association conditions This represents the cardinality of the set (i.e., the number of associated groups). and Representing groups and related groups The center vector (mean of the unit vector within the cluster). The summation term represents the cosine similarity between two center vectors. This represents the average semantic similarity across groups. Adding one results in the semantic association degree across groups.
[0085] The calculation logic of the above formula includes: the first term The local intensity attribute reflecting the modification characterizes the degree to which the group as a whole deviates from the baseline document; the second item The spatial breadth attribute reflects the extent of the modification, representing the scope of the change in the document's physical space; the third item The network effect attributes of the modification reflect the semantic correspondence between the change and other modifications in the document. The three factors are coupled through multiplication, making it difficult for groups with outstanding differences in only a single dimension to obtain a high comprehensive insight index, while groups that perform well in all three aspects of intensity, scope, and relevance obtain exponentially higher evaluation values. This allows for the accurate identification of systematic modification intentions with high business value, while filtering out accidental formatting noise or local typos.
[0086] Step S150: Determine the credibility ranking based on the comprehensive insight index of each document to be compared, and output the comparison results.
[0087] Optionally, step S150 may include: sorting the comprehensive insight index of each difference group, extracting a preset number of difference groups at the top of the sort as high-value difference groups; calculating the sum of the comprehensive insight index of all difference groups in each document to be compared, determining the credibility ranking of each document based on the sum, and determining the document at the top of the sort as the benchmark document; and generating a structured comparison result containing difference location identifiers and semantic level difference descriptions based on the high-value difference groups.
[0088] The aforementioned comprehensive insight index ranking is a filtering operation that sorts each difference group in descending order based on its index value, aiming to identify the modification clusters with the highest business focus from the entire difference set. The preset quantity parameter (usually denoted as K) can be set based on a fixed threshold (e.g., extracting the top 10% or top 20 groups) or based on cumulative contribution rate (e.g., extracting the set of groups with a cumulative index share exceeding 80%), achieving tiered filtering from high-value to marginal-value by truncating the ranking sequence. High-value difference groups are defined as the set of difference groups ranking in the top K positions in the ranking sequence. These groups perform excellently in three aspects: intra-cluster difference strength, document impact range, and cross-group semantic relevance, representing a systematically planned modification intent or one with substantial semantic impact in the document, rather than accidental formatting adjustments or noise identification.
[0089] Document-level credibility assessment is achieved through aggregation analysis, calculating the sum of the comprehensive insight indices of all difference groups contained in each document to be compared. This sum quantifies the overall semantic deviation or the systematic level of modification of a specific document relative to other versions. A credibility sequence is generated by sorting the sums in descending order, with the document at the top of the list designated as the benchmark document, serving as the reference point for subsequent visualization comparisons and difference descriptions. In optional implementations, in addition to simple summation, a weighted summation strategy (using the character span of the group as the weight) can be used to highlight the impact of large-scale modifications, or a dispersion calculation based on information entropy can be used to assess the concentration of modification distribution. The logic for determining the benchmark document can be adjusted according to the business scenario; for example, in a version iteration scenario, the version with the most comprehensive modifications can be selected as the benchmark, while in an originality review scenario, the version with the smallest sum of differences can be selected as the benchmark.
[0090] The structured comparison results described above are a multimodal output set constructed based on high-value difference groups, including machine-readable data formats and human-readable visual presentations. Difference location identification is achieved through visual markers, highlighting, underlining, or annotating boxes in the document view to mark the text areas corresponding to high-value difference groups, along with group numbers and index values. Semantic difference descriptions utilize large language models to analyze the semantic vector changes and contextual information of high-value difference groups, generating natural language explanations representing modification intentions (such as "terminology standardization replacement" or "changes in rights and obligations") and the scope of semantic impact. The output format uses JSON or XML structured encoding, including difference group identifiers, location coordinates, comprehensive insight index, a list of related groups, and semantic description fields, enabling the programmatic storage, transmission, and further analysis and processing of difference data.
[0091] Optionally, step S150 may further include: generating semantic-level difference description information based on high-value difference groups using a large language model; wherein the difference description information is used to characterize the modification intent and semantic impact range of the modified content corresponding to the difference group.
[0092] The aforementioned large language model is a pre-trained generative neural network based on the Transformer architecture. It models long-range dependencies in text sequences through a self-attention mechanism, enabling it to understand natural language semantics and generate coherent text. In difference insight application scenarios, the deployed large language model (such as the GPT series, LLaMA, or domain-specific models) receives structured features of high-value difference groups as input cues and outputs difference descriptions in natural language form. This information represents the deep semantic features and business meaning of the modified content corresponding to the difference groups, achieving semantic conversion from numerical indicators to interpretable text. It is understood that the aforementioned GPT series, LLaMA, and domain-specific models are all mature existing technologies. For model architecture and working principles, please refer to relevant technologies; this application's embodiments will not elaborate further.
[0093] Difference description information is semantically interpreted text encoded in natural language, containing a formal representation of the editing intent and the scope of semantic impact. Editing intent refers to the editor's purpose in performing a specific change, including but not limited to standardized terminology replacement (e.g., unifying "Party A" with "Client"), changes to the subjects of rights and obligations (e.g., systematic modification of party names), upgrades to technical standards (e.g., iterative updates to specification numbers), or logical restructuring (e.g., adjustments to the order of clauses). The scope of semantic impact refers to the boundaries by which a specific change affects the overall semantic coherence, logical consistency, and information integrity of the document, indicating whether the change is limited to local vocabulary or involves adjustments to thematic relevance across paragraphs or chapters.
[0094] The vectorized representations of high-value difference groups, along with contextual information, constitute the input feature set of a large language model. In implementation, a zero-shot generation strategy based on prompt engineering can be adopted, constructing structured prompt templates containing group center vector semantic labels, structural location descriptions, associated group identifiers, and original text fragments to guide the model in generating difference descriptions conforming to specific format specifications. Alternatively, a contextual example strategy based on few-shot learning can be used, embedding manually annotated difference description examples into the prompts to capture pattern features of the description style using the model's contextual learning capabilities. A hybrid strategy based on retrieval-augmented generation (RAG) can also be employed, retrieving similar historical cases from a pre-built difference description knowledge base as generation references to enhance the domain adaptability and accuracy of the description content.
[0095] In the above scheme, the large language model maps the semantic vectors, structural weights, and cross-group association features of high-value difference groups to the latent state space through the encoding layer, capturing the deep semantic patterns of modification operations. Then, the decoding layer autoregressively generates descriptive word sequences, calculating the probability distribution of candidate words based on the preceding context and input constraints at each generation step, and sampling the output word units that best conform to semantic coherence. For identifying modification intent, the model utilizes semantic reasoning capabilities learned in the pre-training phase to analyze the semantic domain mapping relationships of word replacements within groups (such as the migration from general vocabulary to specialized terms) to infer the editing purpose. For assessing the semantic impact scope, the model combines the distribution density and spatial span information of related groups to determine whether the modification is an isolated event or part of a systemic change, ultimately generating a natural language description that combines technical accuracy with business readability.
[0096] To facilitate understanding of the working principle of the above-described method for dynamic text content comparison and difference insight for multimodal documents, this application also provides an application example of this method in a specific scenario. In this application scenario, the above-described method for dynamic text content comparison and difference insight for multimodal documents mainly includes: Step 1: Extract multimodal documents and clean them to obtain initial content; The documents to be compared include Word versions (DOCX format) and scanned versions (PDF image format) of the same technical cooperation agreement. For the Word documents, the Python-docx library is used to parse the document.xml component within the ZIP archive structure, extracting paragraph text blocks, table cells, and text box objects, and recording the position number of each unit in the chapter hierarchy (e.g., "Chapter 1, Section 3, Paragraph 2") and font style features (e.g., Heading 1, body text, bold, etc.). For the scanned versions, the PaddleOCR engine combined with DBNet text detection and CRNN text recognition algorithms is used to extract the bounding box coordinates (two-dimensional coordinates of the top left and bottom right corners), recognition confidence scores (0 to 1 range), and corresponding character sequences of the text regions in the image. All extracted text units are uniformly labeled as content units, establishing an initial content set containing spatial location attributes (page number, paragraph number) and format features to provide standardized input for subsequent steps.
[0097] Step 2: For each content unit, determine the content reliability coefficient of the content unit based on its positional relationship with adjacent units across the document space, and determine the structural importance weight of the content unit based on its depth in the document structure hierarchy. Extract all text content units (paragraphs, table cells, image OCR text) from documents of different formats (PDF, DOCX, images, etc.) and record their structural positions (chapter, header, footnote, etc.) in the original documents. Extract text from images using the OCR module to obtain the corresponding text sequences. Further, for the text sequences extracted from each data source, use the jieba word segmentation tool to split them, thus obtaining the vocabulary sets of each data source. Then, evaluate the overall reliability of each data source in representing text information, serving as a credibility reference for subsequent text content comparison across different data sources.
[0098] Because OCR recognition and PDF parsing may contain errors, a content reliability coefficient needs to be calculated for each content unit. : ; By combining units, text structures of different granularities, such as sentences and phrases, can be obtained. These text structures represent a set of multi-granularity semantic units. Within these structures, different units have varying importance (e.g., based on / self-developed / 3D / rendering / engine / Cetus3D / built / zero-code / digital / twin / development / platform; where units like "self-developed" and "platform" should be more important than units like "built" and "based on"). Simultaneously, the document structure tree (chapter level, font style) is used to calculate the structural importance weight for each unit. : ; Step 3: Based on the semantic similarity, structural importance weight, and content reliability coefficient between content units, determine the significance index of differences between cross-document unit pairs and identify the locations of differences; All semantic units of the two documents to be compared (e.g., docx version A and pdf version B) are matched pairwise to calculate semantic similarity. The structural importance weights and the reliability of the content are then combined to derive a significance index for the difference between each matching pair. : ; Thus, by traversing the two documents to be compared, the significance D of the difference between each pair of matches is obtained. Therefore, in actual difference detection, it is necessary to filter out valid and invalid pairs of matches. Thus, the significance D calculated for all current pairs of matches is processed by maximum-minimum normalization to obtain a normalized result with a value range of [0,1]. A threshold is then set. ,like Then it is determined to be a position of difference.
[0099] Finally, a confidence assessment is performed on the comparison between documents. The most credible reference document is extracted from the multimodal document content, and the differences between documents are filtered based on the confidence level.
[0100] Step 4: Cluster the content units corresponding to the differences into difference groups according to their semantic relevance, and determine the comprehensive insight index of the difference groups based on the influence range of the difference groups and the semantic relevance across the groups; Scattered discrepancy units (sentences, phrases) often belong to the same editing intent (e.g., corresponding modifications made by editors of different documents). It is necessary to cluster semantically similar and geographically adjacent discrepancy units into discrepancy clusters and calculate the comprehensive insight index of these clusters. It extracts reliable text structure from documents, enabling accurate matching across different document sources even when the editor's modifications are unknown. The semantic vectors are embedding vectors extracted by the Word2vec tool, typically with dimensions of 384, 768, or 1024 (optional; 1024 is used in this example). Each component is a semantic dimension, reflecting the distribution characteristics of units in terms of word meaning, syntax, context, etc. Therefore, DBSCAN is used as a clustering algorithm to extract various clusters from the document, and the comprehensive insight index of each cluster is extracted: ; This completes the calculation of the comprehensive insight index of the words in the current text, and further generates matching results between text content based on the comprehensive insight.
[0101] Step 5: Determine the credibility ranking based on the comprehensive insight index of each document to be compared, and output the comparison result; According to the comprehensive insight index of the difference cluster Sort all clusters in descending order, and extract the top high-value clusters. Then, rank the credibility of the document data sources in descending order according to the sum of the insight indices I of the clusters where all their units are located (the indices of the clusters where each vocabulary is located are added together), so as to generate a structured report. The report content includes the visual marking of the modified content (the one with the highest score is used as the original text, and the rest are used as the modified text), and a short description is generated through a large language model (such as "changing 'Party A' to 'Entrusting Party' affects the main body of rights and obligations throughout the text"). Finally, output the comparison result in JSON or XML format, and attach a visual comparison view, marking the key difference clusters to achieve the intelligent upgrade from "character change" to "semantic insight".
[0102] So far, the present invention is completed.
[0103] In summary, in this embodiment of the invention, at least two documents to be compared are obtained, and text content units and their structural location information are extracted from the documents to be compared. For each content unit, the content reliability coefficient of the content unit is determined based on the positional relationship between the content unit and adjacent units across the document space, and the structural importance weight of the content unit is determined based on the depth of the content unit in the document structure hierarchy. Based on the semantic similarity, structural importance weight, and content reliability coefficient between content units, the difference significance index of cross-document unit pairs is determined, and the difference positions are identified. The content units corresponding to the difference positions are clustered into difference groups according to semantic relevance, and the comprehensive insight index of the difference groups is determined based on the influence range of the difference groups and the semantic relevance across the groups. The credibility ranking is determined based on the comprehensive insight index of each document to be compared, and the comparison results are output. This invention achieves an intelligent upgrade from character-level difference detection to semantic-level modification intent insight through a progressive processing flow of structured unified representation of multimodal content, reliability coefficient quantification, structural weight evaluation, and semantic-level difference clustering, thereby improving the accuracy and interpretability of document comparison. On the other hand, by calculating the content reliability coefficient based on the positional relationship of adjacent units across the document space and combining it with error sources such as OCR confidence for penalty constraints, it effectively suppresses noise interference in the image-modal document extraction process, enhancing the objectivity of content unit quality assessment in multimodal scenarios. Furthermore, by integrating the difference significance calculation mechanism of semantic vector similarity and structural importance weights, it amplifies the impact of structural distribution inconsistencies using the inverse mapping of weight differences, while using the minimum reliability coefficient as a constraint condition to reduce the probability of misjudgment caused by text order changes due to format adjustments. Finally, by clustering difference units into difference groups and calculating a comprehensive insight index, combined with the comprehensive calculation of the average difference intensity within the cluster, the influence range of character span, and the semantic correlation degree across groups, it can identify modification intents with associative characteristics, revealing the propagation influence range of modified content in the document semantic network.
[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for multi-modal document oriented text content dynamic comparison and difference insight, characterized in that, The method includes: Obtain at least two documents to be compared, and extract the text content units and their structural location information from the documents to be compared; For each content unit, the content reliability coefficient of the content unit is determined based on the positional relationship between the content unit and adjacent units across the document space, and the structural importance weight of the content unit is determined based on the depth of the content unit in the document structure hierarchy. Based on the semantic similarity between the content units, the structural importance weight, and the content reliability coefficient, the significance index of differences across document unit pairs is determined, and the location of differences is identified. The content units corresponding to the differences are clustered into difference groups according to semantic relevance, and the comprehensive insight index of the difference groups is determined based on the influence range of the difference groups and the cross-group semantic relevance. The credibility ranking is determined based on the comprehensive insight index of each of the documents to be compared, and the comparison results are output. The step of determining the structural importance weight of a content unit based on its depth in the document structure hierarchy includes: obtaining the depth value of the content unit in the document structure tree and the maximum depth value of the structure tree, and determining the reliability proportion of the content unit; determining a structure hierarchy factor based on the depth proportion of the depth value and the maximum depth value, and the reliability proportion; wherein the structure hierarchy factor increases as the depth value decreases; determining a semantic feature factor based on the cumulative contribution of the feature weights of each semantic feature contained in the content unit relative to the feature weights of the root node; and determining the structural importance weight based on the fusion result of the structure hierarchy factor and the semantic feature factor. The determination of the comprehensive insight index of the difference group based on the influence range and cross-group semantic relevance of the difference group includes: calculating the average value of the difference significance index of each content unit within the difference group to obtain the average difference intensity within the cluster; determining the character span of the difference group in the document, and determining the influence range factor based on the ratio of the character span to a preset span threshold; determining other difference groups that meet the preset semantic relevance conditions with the difference group, calculating the average semantic similarity between the difference group and each of the other difference groups to obtain the cross-group semantic relevance; and determining the comprehensive insight index based on the comprehensive calculation result of the average difference intensity within the cluster, the influence range factor, and the cross-group semantic relevance.
2. The multi-modal document oriented text content dynamic alignment and difference insight method of claim 1, wherein, Determining the content reliability coefficient of a content unit based on its positional relationship with adjacent units across document space includes: Identify the adjacent units that are spatially adjacent to the content unit in different documents, and obtain the position number of the content unit and each of the adjacent units in the document; Based on the difference between the position numbers of the content unit and each of the adjacent units, the local positional correlation degree of the content unit relative to each of the adjacent units is determined, and the local contextual integrity score is determined by combining the local positional correlation degrees. Based on the error flags of the content units during the extraction process and the error location counts of the content units, an error source penalty factor is determined; The content reliability coefficient is determined based on the local context integrity score and the error source penalty factor.
3. The multi-modal document oriented text content dynamic comparison and difference insight method of claim 1, wherein, The determination of the significance index of differences across document unit pairs based on the semantic similarity between the content units, the structural importance weight, and the content reliability coefficient includes: Obtain the semantic vectors of two content units from different documents to be matched, and calculate the semantic similarity between the two semantic vectors; Determine the weight difference of the structural importance weights of the two content units respectively; The basic difference degree is determined based on the inverse mapping relationship between the semantic similarity and the weight difference. Obtain the content reliability coefficients of the two content units, and select the smaller one as the reliability constraint factor; The significance index of the difference is determined based on the result of adjusting the basic difference degree by the reliability constraint factor.
4. The multi-modal document oriented text content dynamic comparison and difference insight method of claim 1, wherein, The identification of the difference location includes: The significance index of the differences between all the cross-document unit pairs is normalized to obtain the normalized difference value corresponding to each significance index. Each of the normalized difference values is compared with a preset threshold. Cross-document unit pairs whose normalized difference values are lower than the preset threshold are identified as difference locations.
5. The multi-modal document oriented text content dynamic comparison and difference insight method of claim 1, wherein, The step of clustering the content units corresponding to the difference positions into difference groups according to semantic relevance includes: Obtain the semantic vector of each content unit corresponding to the difference position; Based on the distribution density and proximity of each semantic vector in the semantic vector space, content units that are semantically similar and spatially adjacent are aggregated into the same difference group.
6. The method for dynamic text content comparison and difference insight for multimodal documents according to any one of claims 1-5, characterized in that, The process of determining the credibility ranking based on the comprehensive insight index of each of the documents to be compared, and outputting the comparison results, includes: The comprehensive insight index of each difference group is sorted, and a preset number of the difference groups with the highest ranking are extracted as high-value difference groups. Calculate the sum of the comprehensive insight index of all difference groups in each of the documents to be compared, determine the credibility ranking of each document based on the sum, and determine the document with the highest ranking as the benchmark document; Based on the high-value difference groups, structured comparison results are generated that include difference location identifiers and semantic difference descriptions.
7. The multi-modal document oriented text content dynamic comparison and difference insight method of claim 6, wherein, The method further includes: Based on the high-value difference groups, semantic-level difference description information is generated using a large language model; wherein, the difference description information is used to characterize the modification intention and semantic impact range of the modified content corresponding to the difference groups.
8. The multi-modal document oriented text content dynamic alignment and difference insight method according to any one of claims 1-5, characterized in that, The modalities of the documents to be compared include digitized document modalities and / or image modalities; The step of extracting text content units and their structural location information from the document to be compared includes: For the document to be compared in the digital document modality, the content units and their structural location information in the document to be compared are directly extracted; For the document to be compared in the image modality, optical character recognition is performed to extract the text content units and their structural location information.