Text semantic element automatic extraction system based on multi-modal large model

The text semantic element automatic extraction system based on multimodal large model solves the problems of attached object drift and structured cascading distortion caused by the migration of anchor points of exception clauses in complex specification documents. It achieves high accuracy and data confidence of process semantic results and is applicable to financial compliance control, government approval process and industrial manufacturing procedures.

CN122491285APending Publication Date: 2026-07-31CLOUD INTERNAL CONTROL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CLOUD INTERNAL CONTROL TECH CO LTD
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing document intelligence and automatic text semantic element extraction systems, when processing complex and standardized documents, fail to perceive the anchor point migration mechanism of local exception clauses, resulting in element attachment object drift and structured cascading distortion.

Method used

The automatic text semantic element extraction system based on a multimodal large model generates a local semantic carrier set through a carrier parsing module, identifies exceptional trigger fragments through a trigger recognition module, calculates anchor point migration scores through an anchor point mapping module, and generates an anchor point constraint mapping set through a constraint extraction module, ensuring that exceptional elements are accurately attached to the actual process sub-branches.

Benefits of technology

It effectively solves the problems of object drift and structured cascading distortion in the exception and restriction clauses of complex specification documents, ensuring the data confidence and graph logic accuracy of the semantic results of the process, and is applicable to fields such as financial compliance control, government approval process and industrial manufacturing procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491285A_ABST
    Figure CN122491285A_ABST
Patent Text Reader

Abstract

This invention discloses an automatic text semantic element extraction system based on a multimodal large model, relating to the fields of natural language processing and document intelligence. Addressing the problem of element extraction distortion caused by anchor point migration in local exception clauses of complex normative documents, this system first segments the normative document to generate a set of local semantic carriers and extracts multi-dimensional structured attributes such as text and page layout. Second, by calculating semantic tension and combining it with spatial page layout features, it accurately identifies and extracts migration trigger fragments that have experienced attachment drift. Then, it calculates the total migration score based on spatial subordination, hierarchical subordinate relationships, and semantic contraction relationships, mapping the trigger fragments to real constraint anchor points. Finally, it aggregates the real anchor points and their bound trigger fragments into a constraint context, which is then input into a multimodal large model for restricted structured decoding and parsing. This invention effectively overcomes the hidden dangers of cross-page element mismatch and achieves high-precision structured extraction of semantics from complex processes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and document intelligence, specifically to an automatic text semantic element extraction system based on a multimodal large model. Background Technology

[0002] In professional fields such as financial compliance management, government approval processes, and industrial manufacturing procedures, the daily operations of enterprises and institutions heavily rely on massive amounts of cross-page, long-chain process specification documents. These normative texts carry rigorous business flow logic, and their layout and expression exhibit a highly complex cross-modal nature. Complete process semantics not only include conventional stage divisions and main actions but also involve intricate triggering conditions and exceptional cases. These process semantic elements often do not appear centrally and continuously in a single linear text but are scattered and discretely embedded in main clauses, cross-page tables, diagram nodes, and even footnotes and local annotations by the layout mechanism. This non-linear layout paradigm makes the extraction of semantic elements from documents a complex systems engineering project. The system needs to comprehensively consider text content, page position, structural hierarchy, and local spatial attachment relationships to accurately reconstruct discrete information into a structured process semantic map.

[0003] Existing document intelligence and automatic text semantic element extraction systems generally employ a conventional operating mode of linear segmentation combined with isolated extraction and proximity merging when processing complex specification documents. The underlying logic of this mode is based on the idealized assumption that linear adjacency of text equates to logical dependence. However, actual specification text compilation and typesetting practices often involve the hidden phenomenon of semantic anchor point migration in local exception clauses. To maintain the visual simplicity of the main clauses, many semantic constraint fragments with restrictive, exceptional, or supplementary characteristics are usually extracted from the main text and marginalized in footnotes, table notes, or cross-page table footers. From the perspective of visual typesetting and text parsing sequence, these fragments appear to be generalized supplementary explanations of the nearest higher-level text, but from the perspective of actual process semantic logic analysis, they often have a strong contractile nature, strictly constraining only a specific sub-branch of the process. Existing technologies lack a quantitative perception and control mechanism for the anchor point migration mechanism. During the extraction of execution elements, they mechanically force exception constraint elements located at the visual edge to the nearest linearly distant superior main clause or adjacent process node. The aforementioned attachment object drift problem directly leads to irreversible structured cascading distortion of the stages, subjects, conditions, and exception elements extracted by the system, resulting in the final generated process semantic results failing to meet the practical application needs of complex compliance review and logical reasoning scenarios. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention proposes an automatic text semantic element extraction system based on a multimodal large model. This system solves the problem of element attachment object drift and structured cascading distortion caused by the inability to perceive the anchor point migration mechanism of local exception clauses when processing complex cross-format specification documents.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an automatic text semantic element extraction system based on a multimodal large model, comprising: The carrier parsing module is used to segment the specification document to generate a set of local semantic carriers. After extracting multi-dimensional structured attributes including text content, page position, structural category, hierarchical depth and local page clusters for the local semantic carriers, it constructs the local dominating domain. The trigger recognition module is used to calculate the difference between the local restricted semantic quantity and the independent action semantic quantity of the local semantic carrier. Based on the preset trigger threshold, the difference is judged to generate an exception trigger mark. The main region crossing state of the local semantic carrier is determined by combining the structure category and the local layout cluster to generate a local attachment indicator. The exception trigger mark and the local attachment indicator are fused to extract the migration trigger fragment set. The anchor point mapping module is used to construct a candidate set of anchor points based on the local dominance domain, calculate the total score of anchor point migration based on local spatial subordination, hierarchical subordinate and semantic contraction relationship, select the highest score to determine the real anchor point, and pair the migration trigger fragment with the real anchor point to generate an anchor point constraint mapping set. The constraint extraction module is used to deduplicate the anchor point constraint mapping set to extract the unique true anchor point. It aggregates the unique true anchor point, adjacent cooperative carriers within the local dominance domain to which the unique true anchor point belongs, and the bound migration trigger fragments into the anchor point constraint context. The anchor point constraint context is input into the multimodal large model to extract the process semantic result set.

[0006] Compared with existing technologies, it has the following advantages: This proposed system for automatic extraction of text semantic elements based on a multimodal large model overcomes the inherent limitations of traditional document element extraction, which relies excessively on linear adjacency relationships. It constructs a local dominance recovery mechanism based on multimodal attribute collaborative judgment, effectively solving the technical challenges of object drift and structured cascading distortion in complex normative documents. During the parsing phase, the system deconstructs unstructured text into local semantic carriers carrying multidimensional spatial and logical attributes. By quantifying the difference between constraint semantic quantities and action semantic quantities, it accurately detects exception constraint fragments with rewriting tendencies that are detached from the main page flow. When determining the true target, this solution abandons a single literal distance matching criterion, integrating local spatial subordination, hierarchical subordinate relationships, and semantic contraction relationships based on a natural language inference model to calculate the total anchor point migration score. The aforementioned operational logic transforms implicit typesetting cognition into strictly executable tension comparison instructions, ensuring from the underlying algorithm that every supplementary element at the visual edge can be accurately reconnected to its true constraint sub-branch.

[0007] After establishing a rigorous attachment mapping network, this invention substantially reduces the inherent divergence and illusion risks of large multimodal models when handling long-chain complex logic by constructing a constrained anchor point constraint for the final extraction of execution elements. The system forcibly aggregates real anchor points, adjacent collaborative carriers, and exception constraint fragments after binding and reconnection into tree-like graph-text multimodal input pairs in memory, defining extremely strict information boundaries for the reasoning of large models. Combined with a structured parsing and verification mechanism based on an abstract syntax tree in the decoding output stage, the system requires all exception elements to be nested as secondary fields of real trunk nodes. This closed-loop mechanism not only achieves unambiguous extraction of semantic elements in the document flow but also completely eliminates the risk of cross-page element mismatch in the underlying computer data architecture, resulting in a final output set of flow semantic results with extremely high data confidence and graph logic accuracy, enabling it to be directly integrated as a high-specification relational data pair into subsequent automated compliance review and business reasoning engines. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the system framework of the present invention.

[0009] Figure 2 This is a schematic diagram of the system flow of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] Please see Figures 1 to 2 This application provides an automatic text semantic element extraction system based on a multimodal large model, including a carrier parsing module, a trigger recognition module, an anchor mapping module, and a constraint extraction module; The carrier parsing module aims to construct local semantic carriers of a document and form local dominating domains. In the scenario of automatically processing long-chain process specification documents across different layouts, the original input is unstructured multimodal data. To solve the anchor point migration problem of local exception clauses in the layout, the system needs to break away from the traditional linear segmentation based on natural paragraphs and transform the original document into a fine-grained carrier set that can be directly located and includes spatial topological relationships. The specific execution flow of this module is as follows: Obtain the original multimodal specification document and perform page-level layout parsing; Specifically, the system first obtains the original multimodal specification document D. Document D contains a total of tp pages. The system then extracts each page of the document sequentially by page number np. ,in For each extracted page The system performs page-level layout parsing, identifies different functional layout areas on the page, and generates a layout area box bx for each identified area. The area box bx is defined by its position on the current page. The geometric parameters defined in the page coordinate system include the x-coordinate of the top-left corner of the area box (cx), the y-coordinate of the top-left corner of the area box (cy), the width of the area box (cw), and the height of the area box (ch). Through page layout analysis, the system distinguishes different visually functional page blocks such as the main text area, clause title area, table area, figure area, figure caption area, footnote area, parenthetical note area, and footer note area.

[0012] Specifically, the above process mainly utilizes a deep learning-based object detection network to segment the input multimodal standardized document images or layout files. The extracted region bounding box's top-left corner x-coordinate cx, top-left corner y-coordinate cy, width cw, and height ch together constitute the absolute geometric boundary of the visual block. The underlying mechanism of this processing lies in the spatial deconstruction and precise separation of visually mixed text and image layout content. This provides a clear analytical foundation for subsequent extraction of multimodal fine-grained features containing spatial coordinates, fundamentally isolating the main narrative area from local subordinate areas.

[0013] It should be noted that although layout analysis is a well-known technology in this field, this invention strictly defines the analysis categories to distinguish the visual carriers such as the caption area and the footnote area, aiming to provide a spatial dimension as a prerequisite for the subsequent identification of local exception clauses.

[0014] Within a region, text fragments are segmented to form a local semantic carrier set; Specifically, after obtaining the area bounding boxes (bx) for each page, the system further divides each area bounding box into the smallest text and page layout combination units according to natural semantic breakpoints and document logical structure boundaries. In this invention, this smallest combination unit is defined as a local semantic carrier. After segmentation, the system generates a set of local semantic carriers U covering the entire document D, whose mathematical expression is as follows:

[0015] In the formula, U represents the global and local semantic carrier data group formed after segmentation. This represents the k-th independent local semantic carrier element in the data set, and nu represents the total number of elements generated by the entire document segmentation. This formula discretizes continuous geometric regions into a set of independently addressable elements by traversing all page area boxes bx and performing fine-grained text segmentation. In essence, it further refines two-dimensional visual blocks to the smallest processing entities with independent logical coherence.

[0016] Specifically, the system uses natural language processing sentence segmentation algorithms combined with punctuation and document tags to segment long text within a region box. For example, if a region box bx is a table area, the segmented local semantic carrier... This corresponds to the content of a single cell in the table or a comment section within a merged cell; if the area box bx is the main text area, then it is a local semantic carrier. This corresponds to a single clause / segment. The underlying mechanism of this processing is to forcibly break down lengthy text chains while preserving the basic hierarchical relationships of the segmented fragments. This establishes the minimum granularity of the system's processing, ensuring that each carrier can serve as an independent benchmark unit for unambiguous attachment to subsequent process elements.

[0017] It should be noted that the local semantic carrier differs from the single sentence or word unit in traditional natural language processing. It is defined as the smallest unit of text and layout that can be processed as an independent semantic attachment object or independent semantic fragment without violating local semantic constraints. This carrier should be understood as the smallest input unit for subsequent multimodal large-scale model extraction operations.

[0018] Extract and record multidimensional structured attributes for each local semantic carrier; Specifically, the system addresses any local semantic carrier in set U. Extract and record six-dimensional structured attributes. The first dimension is the text content. Used as a recording medium The corresponding character sequence; the second dimension is the page layout. Used as a recording medium On its home page The precise coordinate frame information in the middle; the third dimension is the page index. Used as a recording medium The page number (np) belongs to; the fourth dimension is the structural category. The fifth dimension is the hierarchy depth, used to identify whether the carrier belongs to a specific category such as main text or figure caption. Used as a recording medium In the logical depth of the document's clause structure, if the carrier is not an explicit clause, it inherits the hierarchical information of its nearest superior clause; the sixth dimension is the local page cluster. Used as a recording medium The visual cluster identifier for the local layout.

[0019] Specifically, the above process extracts text content through rule parsing. With page index For multiple local semantic carriers segmented within the main text area, etc. The system does not directly inherit the macroscopic coordinates of the overall region bounding box bx. Instead, it recalculates the minimum bounding rectangle that encloses the current complete sentence segment based on the single-character-level coordinate boundaries provided by the underlying optical character recognition engine, using a polygon aggregation algorithm. This rectangle serves as the final layout position of the sentence segment. Structural categories The classification results are inherited from the aforementioned process. Simultaneously, the system combines the document directory tree mapping to determine the hierarchy depth. And the local page clusters are calculated using a spatial distance algorithm. This process expands a single text string into a multimodal data set that simultaneously includes visual spatial identity and cross-page continuity. It utilizes a precise independent coordinate system and multi-dimensional collaborative constraints to completely overcome the shortcomings of relying solely on linear text features in long documents, which can easily lead to semantic truncation. This enables large models to accurately perceive the true document location of each carrier based on three-dimensional structured features.

[0020] It should be noted that, in the actual implementation, the hierarchy depth... This can be achieved by using regular expressions to match chapter entry numbers and convert them into a depth-based numerical sequence. (Local page clusters) The position of each page within the same page can be calculated using a density-based spatial clustering algorithm. The Euclidean distance between the nodes is used to achieve this. When calculating the Euclidean distance, the system extracts the pixel value of the mode line height or the base font size of the current page as a dynamic normalization coefficient, thereby setting the density achievable radius of the clustering algorithm. These measures ensure that regardless of the resolution or scaling of the original document input, the visual compactness criterion remains relatively consistent with the text's own layout scale. The introduction of these two parameters and dynamic normalization provides an indispensable and robust quantitative basis for subsequent anchor point migration in breaking the linear order recognition space.

[0021] Constructing local dominance domains based on multidimensional attributes; The system uses the multidimensional attributes extracted in the aforementioned process as the basis for each local semantic carrier. Construct its own local dominance domain Constructing a local dominance domain Using a collaborative constraint strategy, if two carriers have the same local layout clusters... Or in structural category Objects belonging to the same table are included in each other's domains; based on hierarchy depth. Include logically subordinate carriers in the dominion domain; if the text content Explicit references include the surrounding elements of the target location within the dominion domain; finally, the page index is used. Restrict cross-page behavior, allowing the dominating field to extend only to the next page's continuation area with the same number or the same table.

[0022] Specifically, the system uses multidimensional attributes as input conditions and determines the local semantic carriers through logical relation mapping. The spatial and semantic affinity between them. When including logically subordinate carriers, the system implements a single-level depth blocking strategy, that is, only carriers with a depth of [insert depth here] are included. The direct child nodes (i.e., the direct child nodes of the next level determined by the hierarchy depth) are included in the dominance domain, without infinite recursive extension of descendant nodes, thus strictly limiting the data expansion of the local dominance domain. When processing explicit reference identifiers, the system first extracts the reference entity keywords, and then, within the already parsed structural categories... To obtain the unique identifier of the target carrier, regular expression matching is performed on the carrier set of the corresponding category. Then, the target carrier and carriers in its spatial neighborhood are included in the dominance domain. This process transforms the global linear retrieval space of the entire document into a directed relational network based on local topology, which effectively reduces the search scale for finding the actual target object, eliminates interference from irrelevant page numbers and irrelevant page clusters, and significantly reduces the probability of mismatches at the process node level from the data structure level.

[0023] It should be noted that the term "local dominance domain" is a non-public technical term used in this invention, defined as the domain that, without crossing obvious structural breaks, relates to the current semantic carrier. A set of semantic carriers that maintains local associations in page space and hierarchy, thus potentially becoming the object or being acted upon by it. In the low-level implementation of computer systems, local dominance domains... It is not a new storage area but an index list that records the unique identifier of the associated carrier. The aforementioned structural design, combined with a single-level deep blocking strategy, ensures that subsequent processes can be called efficiently without generating data redundancy and computational explosion.

[0024] Specifically, after processing by the carrier parsing module, the original document D has been fully reconstructed into a set of local semantic carriers U with rich multimodal attributes, and each carrier... All are equipped with clearly defined local action boundaries and local dominance domains. The aforementioned output will serve as the complete base data input for subsequent direct filtering of migration-triggered fragments from the set.

[0025] The trigger identification module aims to accurately identify special segments with exception constraints and page attachment offsets from the local semantic carrier set U. Its core objective is to avoid invalid calculations of the global main clauses. Through multi-dimensional feature fusion, it filters out the local exception clauses that truly cause distortion of the process structure, generating a migration trigger segment set M, thereby providing high-value candidate inputs for subsequent real anchor point mapping. The specific execution flow of this module is as follows: Perform semantic tension analysis on the local semantic carrier to generate exception triggering tags; Specifically, the system traverses each local semantic carrier in the set U of local semantic carriers. Extract its corresponding text content Exception semantic feature discrimination is performed. To avoid misjudgments caused by relying solely on keyword triggers, the system introduces an exception rewriting degree formula to evaluate the text content. The quantitative calculation is performed using the following formula:

[0026] In the formula, Represents local semantic carriers The score for the degree of rewriting of exceptions. This indicates a locally restricted semantic quantity. Represents independent action semantics. Locally restricted semantics. Used to measure text content The degree to which existing rules are contracted, excluded, or conditionally rewritten; the semantic quantity of independent actions. Used to measure text content Does it possess a complete subject-verb-object structure, thus constituting a new, independent processing action? This is determined by the calculated exception rewriting score. When the trigger value exceeds the preset threshold, it indicates that the carrier exhibits characteristics of strong restriction on rewriting and weak independent action, and the system identifies it as a local semantic carrier. Assign an exception trigger flag and will The value is set to a Boolean value of true, which is the number one; otherwise, it is set to the number zero.

[0027] Specifically, this process mainly utilizes semantic role labeling algorithms and dependency parsing trees from natural language processing. The system processes the text content... Input the pre-trained language model, extract the attention weights representing words rewritten within the range of exceptions or special cases, and normalize them to the real number interval of zero to one using an activation function as local restrictive semantic quantities. Simultaneously, the system traverses the dependency syntax tree, calculates the connectivity of the child nodes of the core predicate verb, and performs normalization processing to obtain independent action semantic quantities. The preset trigger threshold is set as an empirical constant greater than zero, with an optimal range of 0.2 to 0.4. This ensures that the system can detect the action as long as the restrictive semantics substantially outweigh the independent action semantics in terms of weight. The mechanism behind this process is to extract the substantial auxiliary constraint semantics through vectorized computation, effectively avoiding the large number of false positives generated by conventional block segmentation techniques that rely solely on dictionary matching. This ensures that the marked fragments do indeed have the potential to rewrite the attributes of other process nodes.

[0028] It should be noted that regarding the degree of exception rewriting... Traditional entity extraction systems typically focus on extracting complete action elements, while the calculation logic of this formula is exactly the opposite. Its purpose is to identify subordinate constraint fragments that lack independent actions but are highly dependent on other terms to be valid. This is an objective quantitative tool for computer systems to understand whether text has an exception constraint tendency.

[0029] Local attachment indicators are generated based on spatial structural property analysis; Based on the determination of text semantic features, the system further combines layout features to analyze local semantic carriers. Assess its marginalization properties. Systematically extract the structural category of the carrier. With partial page clusters A comprehensive assessment was conducted when the structural category Belonging to non-main narrative categories such as footnotes, figure captions, table notes, or footer notes, and identified by the system as a local page cluster of this medium. When it is shown that the carrier has not crossed the main page flow (i.e., the main region crossing status is determined to be not crossed), the system determines that the carrier has a tendency to cover locally and generates a local attachment indication for it. Local semantic carriers that satisfy the aforementioned conditions Local attachment indication A value set to true (a Boolean value) is one; otherwise, it is zero.

[0030] Specifically, this process utilizes a conditional branching algorithm to classify the structural categories that reflect visual categories. Clusters of local layouts that reflect spatial compactness This is used as a judgment criterion. When determining whether to cross the main page flow, the system will consider the local page clusters of the current carrier. The identifier is mathematically compared with the identifier of the main region cluster occupying the largest text connectivity area on the current page. If they do not match, the carrier is determined to belong to a local coverage area. This process simulates the typesetting conventions of long standard texts, performing computer-level identity verification on branch nodes that are formally subordinate explanations but actually affect local logic. It precisely strips away ordinary text clauses that, although containing exception terms, are actually part of the global general rules, allowing the system's computing power to accurately focus on those local subordinate regions that are easily misclassified in conventional linear extraction.

[0031] It should be noted that local attachment indication This is a technical parameter specifically designed for this invention, and it is related to the exception triggering flag. Together, they constitute the orthogonal feature dimensions for judging the risk of anchor point migration. Neither a single text feature nor a single layout feature can accurately describe the layout intent of a complex document. Through the joint verification of the two, those skilled in the art can achieve efficient locking of the target segment through logic gate circuits or Boolean intersection operations at the code level.

[0032] Fusion triggering and attachment feature extraction of migration triggering fragment set; Specifically, the system considers all local semantic carriers in set U. Perform a traversal check if and only if a certain carrier simultaneously satisfies the exception trigger flag. Equal to a number and a local attachment indicator When the value equals one, the system will use this local semantic carrier. Upgrade is identified as a migration trigger fragment The system will select all migration trigger fragments. The assembly process generates a set M of migration trigger fragments, whose mathematical expression is as follows:

[0033] In the formula, M represents the final generated migration trigger fragment data group. This represents the q-th identified exception clause object in the set, and nm represents the total number of extracted migration trigger fragments. Each migration trigger fragment... It fully inherits the original local semantic carrier All six-dimensional structured properties and their exclusive local dominating domains .

[0034] Specifically, the system performs strict Boolean logic AND operations to filter the massive set of local semantic carriers U through dimensionality reduction. This process involves forced intersection matching of cross-modal features to extract local constraint fragments that possess the essence of exception constraints but are located at the edge of the layout and are prone to attachment drift. This completes the extraction from general text to high-energy constraint nodes, greatly reducing the scale of data processing and ensuring the efficiency and accuracy of subsequent real anchor mapping algorithms.

[0035] It should be noted that a migration-triggered fragment is defined as a special semantic unit in a multimodal document structure that simultaneously possesses local exception constraints and experiences marginalization in its layout position. This is achieved through direct inheritance of the local dominating domain. The system does not need to recalculate the action boundary during computer memory allocation. The aforementioned data transfer mechanism ensures that subsequent processing can directly search for the real dominated object within the limited security range.

[0036] Specifically, after feature discrimination and logical filtering by the trigger recognition module, the system has successfully condensed the complex set of local semantic carriers U into a high-density set of migration trigger fragments M. In the subsequent anchor mapping module, the system will directly use set M as the core input object, utilizing each migration trigger fragment... Built-in multidimensional attributes and local dominance domain By calculating the anchor point migration score, the system identifies the actual process backbone node to which the anchor point is attached in the document, thereby reattaching the detached exception elements to the correct process tree branch.

[0037] The anchor mapping module is designed to trigger fragments for each migration in set M. The system identifies the core nodes in the document where exceptions truly function, addressing the fundamental question of which nodes a local exception clause applies to. By constructing a local dominance recovery mechanism, the system calculates the multidimensional correlation between migration trigger fragments and each candidate node, thereby repositioning detached exception elements to the correct process tree branch and ultimately outputting the anchor constraint mapping set MP. The specific execution flow of this module is as follows: Construct a candidate set of anchor points based on the local dominance domain; Specifically, the system iterates through each migration trigger fragment in the migration trigger fragment set M. The system extracts the local dominance domain inherent in the fragment itself. and in that local dominance domain Within the index range, select carriers capable of carrying the main process, and generate a migration trigger fragment accordingly. anchor point candidate set Its mathematical expression is as follows:

[0038] In the formula, Represents the candidate data group for anchor points. This represents the v-th candidate carrier element in the set that may be subject to the exception clause, and nc represents the number of candidate carriers retained after filtering. The system-defined filtering criteria require candidate carriers to meet at least one of the following characteristics: containing a core predicate verb lexical element, containing a named entity noun representing an organization or role, or belonging to a master data unit in a table structure, or belonging to the descriptive text of a diagram node.

[0039] Specifically, the system reads the local dominance domain. The system retrieves the text content and structural category of the corresponding carrier from the memory pointer. It uses natural language processing tools to extract the part-of-speech sequence of the text content, determines the presence of verbs, and obtains entity nouns through a pre-defined named entity recognition interface. Simultaneously, it excludes marginal carriers belonging to the same subordinate annotation category based on structural category. This process eliminates redundant data that is not qualified for attachment by pre-shrinking spatial and logical boundaries. This process precisely limits the enormous computational workload that would normally require searching the entire document to an extremely limited number of candidate carriers (nc), greatly improving the operating efficiency of the computer system and reducing the probability of attachment object drift from the perspective of physical search space isolation.

[0040] It should be noted that the anchor point candidate set It is not a collection of all the main nodes in the document, but is strictly limited to the local dominating domain. The dynamic subset. The aforementioned screening criteria based on underlying natural language features limit the candidate objects to core entities with substantial procedural significance, avoiding deadlocks in computer program logic caused by nested attachments between multiple exception clauses.

[0041] Calculate the multi-dimensional anchor point migration score to determine the attachment weight; Specifically, for the anchor point candidate set Each candidate vector The system analyzes its relationship with the corresponding migration trigger fragment. The system calculates the anchor point migration score (sc) based on the spatial and semantic deep relationships between the text elements. Instead of using a simple nearest-text-distance matching rule, the system uses three orthogonal criteria for joint determination and linear weighting. The calculation formula is as follows:

[0042] In the formula, the parameter sc represents the candidate vector. The total score for anchor point migration. The parameter sz represents the local subordination score, which is used to quantify the degree of co-occurrence between the trigger fragment and the candidate carrier in the page space; the parameter sh represents the structural subordinate score, which is used to quantify the specificity of the candidate carrier in the clause hierarchy tree; the parameter sr represents the semantic contraction score, which is used to quantify the actual constraint effect of the trigger fragment on the scope of applicability of the candidate carrier.

[0043] Specifically, when calculating the local subordinate score sz, if the local layout cluster identifiers of the two are consistent or belong to the same table row and column structure in terms of structural category, the system assigns sz a constant value of one (as the first constant value); otherwise, it assigns a constant value of zero. When calculating the structural subordinate score sh, the system extracts the hierarchical depth of the two; if the candidate carrier... If the hierarchical depth value of the transfer trigger fragment is greater than that of the transfer trigger fragment, the system assigns sh to a constant of one (as a second constant value); otherwise, it assigns it to a constant of zero. When calculating the semantic shrinkage score sr, the system uses the candidate carrier text as the premise input and the transfer trigger fragment text as the hypothesis input, both fed into a pre-trained natural language inference implication model. The probability distribution values ​​of the contradictory labels or implication labels output by the model are extracted as the sr score, which ranges from zero to one. Finally, the three scores are directly summed as floating-point numbers under the same dimensions. This process transforms the nonlinear reading derivation process of document experts into a rigorous feature vectorization calculation that can be executed by a computer. This process overcomes the serious deficiency of traditional information extraction algorithms that rely solely on literal distance, enabling the system to accurately pinpoint the true business constraint target of exception clauses based on the comprehensive tension of three dimensions: layout, directory hierarchy, and semantic implication.

[0044] It should be noted that the three scoring parameters mentioned above are not ordinary string similarity calculations, but rather a local dominance recovery quantification model specifically built around the anchor point migration mechanism that is prone to occur in local exception clauses. The clearly defined asymmetric natural language reasoning input structure with explicit premises and assumptions ensures that the system can accurately identify unidirectional conditional constraints, guaranteeing the objectivity and reproducibility of the scoring results from the algorithm's underlying layer.

[0045] Determine the actual anchor points and generate a set of anchor point constraint mappings; Specifically, after completing all candidate vectors After scoring, the system selects anchor point candidates. The candidate carrier with the highest total score (sc) in the anchor point migration is selected and identified as the migration trigger fragment. The true anchor point The system then triggers all migration fragments. Its corresponding real anchor point The pairs are matched to construct one-to-one mappings, thereby generating the anchor constraint mapping set MP for the entire document. Its mathematical expression is as follows:

[0046] In the formula, MP represents the final generated anchor constraint mapping data set, which contains a number of nm attachment relationship pairs.

[0047] Specifically, the system executes an array sorting algorithm to obtain the index of the highest score and then establishes the true anchor point. If multiple candidate carriers achieve the same highest total anchor migration score, the system triggers a fallback strategy, calculating the migration trigger fragments for each candidate carrier with the same highest total anchor migration score. The Euclidean distance between the center points of the pages is used to force the selection of the candidate carrier with the smallest absolute value of the Euclidean distance between the center points of the pages as the true anchor point. The system then uses hash tables or dictionary data structures to trigger migration fragments. The unique identifier is used as the key to store the real anchor. The unique identifier is used as the value for memory binding, thereby solidifying this corrected nonlinear dependency relationship in the computer's underlying storage. This process primarily involves reconnecting the marginalized local exception clauses to the actual core units of the process flow. This process outputs a set of fully computer-readable structured constraint mappings, completely avoiding the risk of simultaneous system downtime, ensuring that each exception element is firmly attached to its actual target object before final extraction, eliminating the hidden dangers of process node mismatch and cascading distortion.

[0048] It should be noted that the anchor constraint mapping set MP essentially constitutes a graph data structure with local semantic carriers of the document as nodes and actual constraint dependencies as directed edges. The generation of this data structure marks the completion of the semantic logic reconstruction of the document flow and will be directly invoked by subsequent steps as a rigid input rule.

[0049] Specifically, through multi-dimensional relationship calculations by the anchor point mapping module, the system has successfully restored and attached the marginalized migration trigger fragments to their actual local process branches, and output a high-density and structured anchor point constraint mapping set MP. In the subsequent constraint extraction module, the system will directly use the set MP as rigid constraint input, merge and package the exception clauses attached to the same real anchor point with the anchor point's own context, and send them into the multimodal large model to perform restricted text semantic element extraction, thereby ensuring that the final generated structured process semantic result no longer exhibits element drift.

[0050] The constraint extraction module aims to leverage the aforementioned mapping constraints to merge local exception clauses with their corresponding real process sub-units, and then invoke a multimodal large model to perform structured extraction within the boundaries of this merging context. Its core objective is to address the cascading distortion problem that occurs when exception elements are detached from their original layout attachment relationships. Ultimately, it transforms complex specification documents exhibiting anchor point migration into unambiguous, structured process semantic results for direct use by subsequent process graph modeling or intelligent review systems. The specific execution flow of this module is as follows: Anchor constraint context is constructed based on mapping and binding relationships; Specifically, the system performs a deduplication and aggregation operation on the anchor point constraint mapping set MP to extract all unique real anchor points, forming a unique anchor point set. For each unique true anchor point in this set. The system initiates the context wrapping process. The system collects the actual anchor point. The text content and multidimensional structured attributes are used to retrieve adjacent collaborative carriers belonging to the same process local structure in the local dominance domain, and the anchor point is extracted through the mapping set MP. All migration trigger fragments. The system assembles and aggregates the content of all the aforementioned related carriers to construct a dedicated anchor constraint context. .

[0051] Specifically, the system allocates an independent data structure space in the computer's underlying memory to store the real anchor points. The six-dimensional attributes serve as the core nodes of the tree structure, with the text content and layout positions of all migration trigger fragments belonging to it as its direct child nodes. The system extracts the unique true anchor point, adjacent collaborative carriers, and the corresponding local image slices of all migration trigger fragments belonging to it in the original document. Combined with the text content and layout position, and following a prompt engineering template with positional coding, it serializes these into a richly illustrated multimodal input pair. This processing is achieved through pre-coded hard-coding assembly, forcibly binding loose document fragments into a logical whole in the computer's low-level memory. This process sets strict information hard boundaries for subsequent large-scale model inference, effectively isolating attention interference from other irrelevant exception clauses across layouts, and ensuring that the extraction process is carried out in a strictly constrained and information-complete environment.

[0052] It should be noted that the anchor point constraint context Defined as a restricted semantic analysis entity formed by packaging the ontology content, local contextual relationships, and one or more exception clauses attached to a unique real anchor point during semantic element extraction. The aforementioned data aggregation and deduplication method avoids redundant calculations caused by multiple exception clauses pointing to the same node. It transforms the complex reasoning task, which originally required large models to search for relationships within large texts, into a simple local multimodal reading comprehension task through pre-dimensionality reduction, significantly reducing the model's computational cost and error rate.

[0053] Extracting multimodal semantic features from large models with performance limitations; The system will construct the anchor constraint context. In the pre-tuned multimodal large model, restricted text semantic feature extraction is performed. Its core extraction logic is expressed by the following synthesis formula:

[0054] In the formula, This indicates that the anchor point is unique. Extract a single structured result from the output. This represents the semantic feature extraction operation operator performed on the multimodal large model. The set part represents all operations in the mapping set MP that are determined to be attached to this anchor point. The migration triggers the convergence and combination of fragments. The calculation process of this formula first strictly locks the attachment relationship through set operations, and then initiates the forward propagation inference extraction of the multimodal large model. The multimodal large model is based on Parse and output standard six-tuple structured semantic elements, specifically including stage elements. , main elements , conditional elements , exception elements Successor relationship elements and anchor elements of evidence .

[0055] Specifically, when calling the large model application interface, the system uses mandatory data architecture output constraint instructions to introduce a structured output protocol or a constraint decoding mechanism based on context-free grammars to guide the model in generating results. In the decoding output stage, the system performs structured parsing and validation of the generated token stream based on an abstract syntax tree, mandating exception elements. Must be used as a real anchor point The nested output of secondary fields prevents the model from independently generating new global exception nodes detached from the anchor point. The mechanism behind this process leverages the powerful natural language generalization capabilities of the large model to parse complex long sentences, while simultaneously using a strict external data structure whitelist to suppress the inherent divergence illusion of the large model. This process ensures that extracted exception elements no longer drift to incorrect branches but are stably fixed under the actual object of action, overcoming the technical challenge of mismatched elements across page layouts in complex specification documents from the underlying architecture.

[0056] It should be noted that the above six structured elements are the customized extraction targets of this invention. Among them, The operation involves a deep neural network mapping process that combines spatial feature attention calculation with multimodal feature fusion, rather than simple keyword extraction. This includes evidence anchoring elements. It fully inherits the layout position extracted during the preprocessing. and page index This ensures that all intelligent extraction results have a traceable computational origin from the underlying data link.

[0057] The final semantic result set of the process is summarized and assembled. Specifically, the system collects all extracted single structured results. The data integrity is verified and null values ​​are removed, and finally, a complete document-level semantic result set R is generated. Its mathematical expression is as follows:

[0058] In the formula, R represents the structured set of all extracted elements from all process nodes in the entire document, and nt represents the total number of process nodes in the document, including real anchor points and ordinary trunk nodes. If there are ordinary trunk process units in the original document that are not attached to any migration trigger fragments, the system constructs ordinary contexts without abnormal attachments using the same dimensions and supplements them by calling the extraction interface to generate regular results, which are then uniformly incorporated into the result set R to ensure the continuity of the document's logical tree.

[0059] Specifically, the system performs a data serialization operation, converting all individual structured results into a single data sequence. The document's topology is sorted according to its original hierarchical depth and persistently stored in a standard data exchange format. The implementation logic of this process is to aggregate data from discrete extraction units into a globally connected graph. This process substantially transforms multimodal specification documents, which originally suffered from anchor point confusion and visual layout fragmentation, into high-quality relational data that computers can directly traverse and reason about, enabling seamless integration into subsequent business process engines or automated compliance review systems.

[0060] It should be noted that the final output set of process semantic results R is the overall target product of the implementation scheme of this invention. After the layer-by-layer noise reduction and forced topology alignment of the aforementioned four modules, this result set has significant architectural advantages over traditional single block extraction schemes in terms of data confidence and graph logic accuracy, and fully realizes the technical leap from manual reading and judgment to automated and accurate parsing.

[0061] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. An automatic text semantic element extraction system based on a multimodal large model, characterized in that, include: The carrier parsing module is used to segment the specification document to generate a set of local semantic carriers. After extracting multi-dimensional structured attributes including text content, page position, structural category, hierarchical depth and local page clusters for the local semantic carriers, it constructs the local dominating domain. The trigger recognition module is used to calculate the difference between the local restricted semantic quantity and the independent action semantic quantity of the local semantic carrier. Based on the preset trigger threshold, the difference is judged to generate an exception trigger mark. The main region crossing state of the local semantic carrier is determined by combining the structure category and the local layout cluster to generate a local attachment indicator. The exception trigger mark and the local attachment indicator are fused to extract the migration trigger fragment set. The anchor point mapping module is used to construct a candidate set of anchor points based on the local dominance domain, calculate the total score of anchor point migration based on local spatial subordination, hierarchical subordinate and semantic contraction relationship, select the highest score to determine the real anchor point, and pair the migration trigger fragment with the real anchor point to generate an anchor point constraint mapping set. The constraint extraction module is used to deduplicate the anchor point constraint mapping set to extract the unique true anchor point. It aggregates the unique true anchor point, adjacent cooperative carriers within the local dominance domain to which the unique true anchor point belongs, and the bound migration trigger fragments into the anchor point constraint context. The anchor point constraint context is input into the multimodal large model to extract the process semantic result set.

2. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, Extracting local layout clusters from multidimensional structured attributes of local semantic carriers, including: Extract the layout position of the local semantic carrier, and extract the pixel value of the mode line height or the base font size of the page where the local semantic carrier is located as a dynamic normalization coefficient. The density reachable radius of the density-based spatial clustering algorithm is set using dynamic normalization coefficients; Calculate the Euclidean distance between the layout positions of various local semantic carriers within the same page; Based on the density reachable radius and Euclidean distance, local layout clusters are generated using a density-based spatial clustering algorithm.

3. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, Constructing local dominance domains based on multidimensional structured attributes, including executing collaborative constraint strategies: Local semantic carriers that have the same local layout clusters or belong to the same table object in the structural category are included in each other's local dominance domains; When including logical subordinate carriers in the local dominance domain, a single-level depth blocking strategy is implemented, which only includes the direct child node carriers of the next level based on the level depth determination in the local dominance domain. Extract the reference entity keywords contained in the local semantic carriers, perform regular expression matching search in the set of local semantic carriers whose structural categories match the reference entity keywords to obtain the target carrier, and include the target carrier and the local semantic carriers in the spatial neighborhood of the target carrier into the local dominance domain.

4. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, Calculate the difference between local restriction semantics and independent action semantics, and determine the difference based on a preset trigger threshold to generate an exception trigger flag, including: Extract the text content contained in the local semantic carrier, input the text content into the pre-trained language model, and extract the attention weights of the words representing the scope of the rewritten words; Attention weights are normalized to the real number range of zero to one using an activation function to generate locally restricted semantic quantities; Construct a dependency parsing tree corresponding to the text content, traverse the dependency parsing tree to calculate the connectivity of the child nodes of the core predicate verb, and normalize the connectivity of the child nodes to generate independent action semantics. Calculate the difference between the local constraint semantic quantity and the independent action semantic quantity; If the difference is greater than the preset trigger threshold, an exception trigger flag with a true value is generated; if the difference is not greater than the preset trigger threshold, an exception trigger flag with a false value is generated. The preset trigger threshold is a constant greater than zero.

5. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, By combining structural categories and local layout clusters, the main region crossing status of local semantic carriers is determined to generate local attachment indicators. Exception trigger markers and local attachment indicators are then fused to extract a set of migration trigger fragments, including: When the structural category of a local semantic carrier belongs to a non-main narrative category, determine the main region cluster identifier that occupies the largest text connectivity area within the page where the local semantic carrier is located. Compare the local layout cluster identifier of the local semantic carrier with the trunk area cluster identifier; If the local layout cluster identifier is inconsistent with the trunk area cluster identifier, it is determined that the local semantic carrier has not crossed the trunk area, and a local attachment indicator with a true value is generated; otherwise, a local attachment indicator with a false value is generated. Determine whether both the exception trigger flag and the local attachment indicator are true. If so, extract the local semantic carrier as a migration trigger fragment. All extracted migration trigger fragments are aggregated to generate a migration trigger fragment set.

6. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, Based on the local dominance domain, a candidate set of anchor points is constructed, including: Local semantic carriers with the ability to carry the main process are selected from the local dominance domain as candidate carriers, and the candidate carriers are aggregated to generate a candidate set of anchor points; Among them, the candidate vectors satisfy at least one of the following characteristics: The candidate carrier contains a core predicate verb lexical element in its text content; The candidate carrier contains named entity nouns representing organizations or roles in its text content; The structural category of the candidate carrier belongs to the master data unit in the table structure; The structural category of the candidate carrier belongs to the illustration node description text.

7. The automatic text semantic element extraction system based on a multimodal large model according to claim 6, characterized in that, The total score for anchor point migration is calculated based on local spatial subordination, hierarchical hierarchy, and semantic contraction relationships, including: Determine whether the local layout cluster of the candidate carrier is consistent with the local layout cluster of the migration trigger fragment, or determine whether the structural category of the candidate carrier and the structural category of the migration trigger fragment belong to the same table row and column structure. If so, generate the first constant value as the local spatial subordination score; otherwise, generate the constant zero as the local spatial subordination score. Extract the hierarchical depth of the candidate carrier and the hierarchical depth of the migration trigger fragment. If the hierarchical depth of the candidate carrier is greater than the hierarchical depth of the migration trigger fragment, generate a second constant value as the hierarchical subordinate relationship score; otherwise, generate a constant zero as the hierarchical subordinate relationship score. The text content contained in the candidate carrier is used as the premise input, and the text content contained in the transfer trigger fragment is used as the hypothesis input. Both are fed into the pre-trained natural language inference implication model. The probability distribution value of the contradictory label or the probability distribution value of the implication label output by the natural language inference implication model are extracted as the semantic shrinkage relation score. The scores for local spatial subordination, hierarchical subordinate relationships, and semantic contraction are summed to obtain the total anchor point migration score.

8. The automatic text semantic element extraction system based on a multimodal large model according to claim 7, characterized in that, The highest score is selected to determine the true anchor point, including: If multiple candidate carriers obtain the same highest anchor point migration total score, then calculate the Euclidean distance between the page center point and the migration triggering segment for each candidate carrier that obtains the same highest anchor point migration total score. The candidate carrier with the smallest absolute value of the Euclidean distance from the center point of the page is selected as the real anchor point.

9. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, The unique true anchor point, adjacent cooperative carriers within the local dominance domain to which the unique true anchor point belongs, and the bound migration trigger fragments are aggregated into an anchor point constraint context, including: The multidimensional structured attributes of the unique real anchor point are used as the core node of the tree structure, and the text content and layout position of the bound migration trigger fragment are used as the direct child nodes of the core node. Extract the local image slices corresponding to the unique real anchor point, adjacent cooperative carriers, and bound migration trigger fragments in the specification document; The local image slices are combined with the unique real anchor point, adjacent collaborative carriers, and the text content and layout position contained in the bound migration trigger fragment. The slices are then serialized and assembled according to the prompt engineering template with position encoding marks to generate multimodal input pairs as anchor constraint contexts.

10. The automatic text semantic element extraction system based on a multimodal large model according to claim 1, characterized in that, Input the anchor point constraint context into the multimodal large model to extract a set of process semantic results, including: A structured output protocol or a constraint decoding mechanism based on context-free grammar is used to guide the multimodal large model to generate a token stream based on anchor constraint context; Perform structured parsing and validation of the token stream based on an abstract syntax tree, and extract exception elements from the token stream; Nested output of exception elements as secondary fields with unique true anchors to generate a set of process semantic results.