A structured text reconstruction method and system based on multi-modal feature fusion

CN122618639BActive Publication Date: 2026-10-09HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611060075.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-10-09
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

[0004]然而,上述现有技术在处理真实场景文档时存在多方面的底层技术缺陷

Benefits of technology

[0028]1. This invention improves reconstruction stability and fault tolerance when facing non-uniform physical degradation. By using a spatial smoothing filter to process the initial noise matrix, it can output a degradation matrix with spatial continuity and smoothness when dealing with continuously varying physical noise in the image, avoiding high-frequency oscillations in feature fusion weights near critical values. Simultaneously, when the degradation matrix indicates severe damage to a local region, it automatically invokes a visual language model, using global contextual priors to compensate for features in the damaged area. This mechanism can supplement the feature tensor through upper-level semantic reasoning even when underlying text features are lost, ensuring the continuity of information transmission in the computational chain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618639B_ABST
    Figure CN122618639B_ABST
Patent Text Reader

Abstract

The application relates to a structured text reconstruction method and system based on multi-modal feature fusion. The application generates an initial noise matrix by calculating local noise variance of a to-be-processed image, outputs a degradation matrix by using a spatial smoothing filter, generates an isolation mask by combining color distribution prior and the degradation matrix, controls asymmetric sparse convolution to decouple out text feature flow, extracts a two-dimensional topological anchor point of a text unit and calculates a Manhattan distance, maps to a spatial bias matrix to inject a multi-modal attention network, adjusts attention weight based on the degradation matrix, and calls a visual language model for feature compensation when local noise is over-limit, and finally decodes a fusion feature tensor to output a structured semantic sequence. Through physical degradation perception, spatial topological constraint and feature compensation mechanism, the application suppresses the cascade amplification of local extraction error along the calculation link, and improves the accuracy of reconstruction in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent document processing technology, and in particular to a structured text reconstruction method and system based on multimodal feature fusion. Background Technology

[0002] With the deepening of digital transformation, the intelligent parsing of complex structured documents (such as bidding documents and legal files) has become a research focus in the fields of natural language processing and computer vision. These documents not only contain dense textual information but also possess complex visual layout features, high semantic intertextuality, and strong logical structure. Accurately reconstructing these multimodal two-dimensional visual documents into machine-readable one-dimensional structured data sequences presents significant technical challenges.

[0003] Currently, the mainstream technologies in this field typically employ a single-modal cascaded architecture based on a serial processing pipeline. This architecture generally first uses an optical character recognition engine to reduce the dimensionality of a two-dimensional document image into a one-dimensional text character sequence. Then, an independent layout analysis module performs text region segmentation based on heuristic geometric rules. Finally, the text is input into a natural language processing module for semantic parsing and data extraction.

[0004] However, the aforementioned existing technologies suffer from several underlying technical deficiencies when processing real-world documents. First, in the image feature extraction stage, when documents contain pixel-level overlaps between stamps, handwritten signatures, and printed text, or non-uniform physical noise such as creases and stains, existing visual extraction models struggle to separate overlapping pixels while preserving the topological integrity of the underlying characters, leading to errors in the extraction of key text features or data loss. Second, regarding document topology reconstruction, existing table recognition algorithms overly rely on physical grid line feature detection, making them highly susceptible to failure when processing borderless tables or images with broken lines due to scanning degradation; furthermore, the separation of structure prediction and content extraction prevents the constraint of arithmetic consistency of multi-dimensional coupled data at the algorithm's underlying level. Additionally, in long text semantic parsing, the truncation of the context window size due to the self-attention mechanism of large language models prevents the model from handling cross-references across long texts, resulting in failures to resolve long-distance entity references and logical breaks.

[0005] In summary, how to eliminate system-level reconstruction failures caused by the propagation and amplification of local feature extraction errors along the computational link during the process of converting two-dimensional document images containing non-uniform physical noise and complex layout structures into one-dimensional structured semantic sequences is a technical problem that urgently needs to be solved in this field.

[0006] To address this, a structured text reconstruction method based on multimodal feature fusion is proposed. Summary of the Invention

[0007] In view of this, the purpose of this invention is to provide a structured text reconstruction method and system based on multimodal feature fusion. This invention generates an initial noise matrix by calculating the local noise variance of the image to be processed, and outputs a degradation matrix using a spatial smoothing filter. It then generates an isolation mask by combining color distribution priors with the degradation matrix, controlling asymmetric sparse convolution to decouple the text feature flow. Two-dimensional topological anchor points of text units are extracted and Manhattan distances are calculated, mapped to a spatial bias matrix, and injected into a multimodal attention network. Attention weights are adjusted based on the degradation matrix, and a visual language model is invoked for feature compensation when local noise exceeds limits. Finally, the fused feature tensor is decoded to output a structured semantic sequence. This invention, through physical degradation perception, spatial topological constraints, and feature compensation mechanisms, suppresses the cascading amplification of local extraction errors along the computational path, improving the accuracy of reconstruction in complex scenes.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: a structured text reconstruction method based on multimodal feature fusion, comprising:

[0009] Step 1: Acquire the document image to be processed and calculate the local noise variance; perform smoothing processing using a spatial smoothing filter with high and low thresholds, and output the degradation matrix;

[0010] Step 2: Combine the color distribution prior of the document image to be processed with the degradation matrix to generate a physical color gamut soft isolation mask; control the asymmetric sparse convolution through the physical color gamut soft isolation mask to decouple the document image to be processed into text feature stream and non-text feature stream;

[0011] Step 3: Extract the two-dimensional topological anchor points of each text unit in the text feature stream, calculate the spatial distance between each two-dimensional topological anchor point, and map it into a spatial bias matrix; inject the spatial bias matrix into the multimodal attention network, and adjust the fusion weights of the multimodal attention network based on the degradation matrix;

[0012] Step 4: If the degradation matrix shows that the noise in a local area is higher than the preset third confidence threshold, then reduce the weight of the text feature stream in the multimodal attention network and perform feature compensation on the local area. If the degradation matrix shows that the noise in a local area is not higher than the third confidence threshold, then maintain the current fusion weights. Fuse the non-text feature stream, the weighted text feature stream, and the optional compensation features to generate a fusion feature tensor. Decode based on the fusion feature tensor to output a structured semantic sequence.

[0013] In a preferred embodiment, step 1 utilizes a spatial smoothing filter with high and low thresholds for smoothing to output a degraded matrix. This includes: calculating the local noise variance of the document image to be processed to generate an initial noise matrix; the spatial smoothing filter iterates through the values ​​of each matrix element in the initial noise matrix, assigning a damaged state label to matrix elements greater than a preset first confidence threshold and a normal state label to matrix elements less than a preset second confidence threshold; for matrix elements between the first and second confidence thresholds, extracting the state labels of adjacent matrix elements in the spatial neighborhood of the matrix element and determining the state label of the matrix element according to the majority voting principle; and outputting the initial noise matrix with state labels as a degraded matrix.

[0014] In a preferred embodiment, step 2, combining the prior color distribution of the document image to be processed with the degradation matrix to generate a physical color gamut soft isolation mask, includes: calculating the characteristic distance between the document image to be processed and the stamp color reference distribution in the HSV color space to generate a basic color difference matrix; calculating the difference between a preset upper limit threshold for color distance and the basic color difference matrix to obtain a color similarity matrix; multiplying the color similarity matrix by a preset response gain coefficient, and performing point-by-point multiplication of the result with the elements of the corresponding coordinates of the degradation matrix, inputting it into a logistic activation function for numerical mapping, and outputting a continuous probability matrix with elements taking values ​​in the closed interval of zero to one as the physical color gamut soft isolation mask.

[0015] In a preferred embodiment, step 2 uses the physical color gamut soft isolation mask to control asymmetric sparse convolution, decoupling the document image to be processed into a text feature stream and a non-text feature stream. This includes: extracting target coordinate positions in the physical color gamut soft isolation mask with values ​​greater than the mask activation threshold to construct a non-zero coordinate hash table; performing a sparse convolution operation of the first kernel size dimension on pixel regions within the non-zero coordinate hash table to extract stamp texture features as a non-text feature stream; and performing a dense convolution operation of the second kernel size dimension on pixel regions not within the non-zero coordinate hash table to extract character stroke features as a text feature stream.

[0016] In a preferred embodiment, step 3 extracts two-dimensional topological anchor points for each text unit in the text feature stream, calculates the spatial distance between each two-dimensional topological anchor point, and maps it to a spatial bias matrix. This includes: dividing the text feature stream into independent text units and generating bounding rectangles; extracting the absolute physical coordinates of the intersection points of the diagonals of the bounding rectangles in the text feature stream as two-dimensional topological anchor points; calculating the absolute difference of the first coordinate in the horizontal direction and the absolute difference of the second coordinate in the vertical direction for any two two-dimensional topological anchor points, adding the first and second absolute coordinate differences to obtain the Manhattan distance between text units; and inputting the Manhattan distance into a preset negative penalty mapping function, wherein the negative penalty mapping function satisfies the following relation. ,in The Manhattan distance, It is a natural constant. This is the preset exponential smoothing factor. The preset negative basic control coefficient is used; negative bias values ​​with negative values ​​and whose absolute values ​​increase with the increase of physical spatial distance are obtained and constructed into a spatial bias matrix.

[0017] In a preferred embodiment, step 3 involves injecting the spatial bias matrix into the multimodal attention network and adjusting the fusion weights of the multimodal attention network based on the degradation matrix. This includes: calculating the product of the query vector matrix and the key vector matrix in the self-attention computation layer of the multimodal attention network to obtain an initial score matrix; performing matrix addition on the spatial bias matrix and the initial score matrix; performing global average pooling on the degradation matrix after binary numerical mapping to obtain a global degradation coefficient value; subtracting the global degradation coefficient value from a constant to obtain a text modality retention coefficient value; and performing pointwise multiplication on the text modality retention coefficient value and the tensor elements of the text feature stream to complete the adjustment of the fusion weights.

[0018] In a preferred embodiment, step 4, which involves reducing the weights of the text feature stream in the multimodal attention network and performing feature compensation on local regions to generate a fused feature tensor, includes: determining that a continuous pixel region is a damaged local region when the degradation value of a continuous pixel region in the degradation matrix is ​​greater than a third confidence threshold; extracting the context feature vectors of the text units surrounding the damaged local region and the global visual feature map of the document image to be processed, and inputting them into a visual language model to generate a one-dimensional semantic compensation tensor; performing channel dimension mapping on the one-dimensional semantic compensation tensor, and performing spatial broadcasting operation on the mapped one-dimensional semantic compensation tensor based on the two-dimensional physical coordinates of the damaged local region to generate a two-dimensional local compensation feature map; and inputting the two-dimensional local compensation feature map, the non-text feature stream, and the weighted text feature stream into a cross-attention mechanism for alignment and fusion to generate a fused feature tensor.

[0019] In a preferred embodiment, step 4, which decodes based on the fused feature tensor to output a structured semantic sequence, includes: inputting the fused feature tensor into a converter decoding network to output a basic text character sequence and paragraph layout tags; calculating the cosine similarity value of each sentence vector in the basic text character sequence using a locality-sensitive hashing retrieval algorithm; establishing directed connections between target sentence pairs with cosine similarity values ​​greater than a reference threshold to construct a heterogeneous document attribute graph from the basic text character sequence; performing graph convolution feature aggregation on the heterogeneous document attribute graph, and serializing and recombining the aggregated node features according to the hierarchical relationship of the paragraph layout tags to output a structured semantic sequence.

[0020] This invention provides a structured text reconstruction system based on multimodal feature fusion, and the method for structured text reconstruction based on multimodal feature fusion includes:

[0021] The degradation awareness module is used to acquire the document image to be processed and calculate the local noise variance to generate an initial noise matrix; the initial noise matrix is ​​smoothed using a spatial smoothing filter with high and low thresholds to output a degradation matrix.

[0022] The feature decoupling module is used to combine the color distribution prior of the document image to be processed with the degradation matrix to generate a physical color gamut soft isolation mask; and to control asymmetric sparse convolution through the physical color gamut soft isolation mask to decouple the document image to be processed into a text feature stream and a non-text feature stream.

[0023] The topology constraint module is used to extract the two-dimensional topological anchor points of each text unit in the text feature stream, calculate the spatial distance between each two-dimensional topological anchor point, and map it into a spatial bias matrix.

[0024] A multimodal attention fusion module is pre-installed with a multimodal attention network based on the Transformer architecture, which is used to inject the spatial bias matrix into the multimodal attention network and adjust the fusion weights of the multimodal attention network based on the degradation matrix.

[0025] The conditional compensation module is used to determine whether the noise in a local area shown by the degradation matrix is ​​higher than a preset third confidence threshold. If so, the weight of the text feature stream in the multimodal attention network is reduced and feature compensation is performed on the local area. If the noise in a local area shown by the degradation matrix is ​​not higher than the third confidence threshold, the current fusion weights are maintained.

[0026] The feature fusion and decoding module is used to fuse the non-text feature stream, the weighted text feature stream, and the compensation features output by the conditional compensation module to generate a fused feature tensor; and to decode the fused feature tensor to output a structured semantic sequence.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] 1. This invention improves reconstruction stability and fault tolerance when facing non-uniform physical degradation. By using a spatial smoothing filter to process the initial noise matrix, it can output a degradation matrix with spatial continuity and smoothness when dealing with continuously varying physical noise in the image, avoiding high-frequency oscillations in feature fusion weights near critical values. Simultaneously, when the degradation matrix indicates severe damage to a local region, it automatically invokes a visual language model, using global contextual priors to compensate for features in the damaged area. This mechanism can supplement the feature tensor through upper-level semantic reasoning even when underlying text features are lost, ensuring the continuity of information transmission in the computational chain.

[0029] 2. This invention improves the accuracy of multimodal image feature decoupling and the rationality of computational resource allocation. By combining prior color distribution to generate a physical color gamut soft isolation mask, and directly using this mask as the control condition for asymmetric sparse convolution, the separation of text feature streams and non-text feature streams is achieved at the underlying tensor dimension. This processing method not only achieves non-destructive isolation of overlapping interference elements (such as stamps and stains) while preserving the topological integrity of the underlying character strokes, but also effectively controls the memory usage and computational overhead during multimodal concurrent feature extraction by allocating convolutional calculations of different dimensions to different regions.

[0030] 3. This invention enhances the logical accuracy and layout consistency of structured sequence output under complex layouts. By extracting the two-dimensional topological anchor points of text units and calculating the Manhattan distance, it maps these points to a spatial bias matrix and directly injects it into the computational layer of a multimodal attention network. This mechanism transforms the absolute distance metric in physical two-dimensional space into a mandatory penalty term for the self-attention weights of the deep learning model. This allows the model to follow the objective laws of the physical layout of the visual page when performing semantic feature aggregation and alignment, suppressing the disordered splicing phenomenon between text blocks that are physically far apart but semantically similar, thus improving the rigor of the structured semantic sequence. Attached Figure Description

[0031] Figure 1 A flowchart of a structured text reconstruction method based on multimodal feature fusion provided in this embodiment of the invention;

[0032] Figure 2 This is a multimodal feature decoupling data flow graph provided in an embodiment of the present invention;

[0033] Figure 3 Spatial bias attention fusion and compensatory decoding flow graph provided in the embodiments of the present invention. Detailed Implementation

[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0035] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0036] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0037] Please see Figures 1 to 3 This invention provides a structured text reconstruction method based on multimodal feature fusion, the technical solution of which is as follows:

[0038] A structured text reconstruction method based on multimodal feature fusion includes:

[0039] Step 1: Obtain the image of the document to be processed and calculate the local noise variance, specifically satisfying the formula: ,in Let V be the local noise variance. This refers to the total number of pixels within each preset local grid segment of the document image to be processed. For the local grid, the first grayscale value of each pixel. The grayscale average of all pixels within the local grid is used; an initial noise matrix is ​​generated; the initial noise matrix is ​​smoothed using a spatial smoothing filter with both high and low thresholds, and a degraded matrix is ​​output.

[0040] Step 2: Combine the prior color distribution of the document image to be processed in the HSV color space with the degradation matrix to generate a physical gamut soft isolation mask; control the asymmetric sparse convolution through the physical gamut soft isolation mask to decouple the document image to be processed into text feature stream and non-text feature stream;

[0041] Step 3: Extract the two-dimensional topological anchor points of each text unit in the text feature stream, calculate the Manhattan distance between each two-dimensional topological anchor point, and map it into a spatial bias matrix; inject the spatial bias matrix into the multimodal attention network, and adjust the fusion weights of each modality feature in the multimodal attention network based on the degradation matrix;

[0042] Step 4: If the degradation matrix shows that the noise in a local area is higher than the preset third confidence threshold, then reduce the weight of the text feature stream in the multimodal attention network, and call the visual language model to perform feature compensation on the local area to generate a fused feature tensor; decode based on the fused feature tensor to output a structured semantic sequence.

[0043] Example 1:

[0044] This embodiment uses the structured reconstruction of paper legal files archived by the court as an application scenario. Such legal files typically contain dense legal provisions, multi-level indented layout, cross-references of clauses across pages, and are often stamped with red seals or official seals across the binding during the physical archiving process. Due to their age, they also have uneven deterioration noise such as creases and water stains.

[0045] As one embodiment of the present invention, refer to Figure 1 A flowchart of a structured text reconstruction method based on multimodal feature fusion, referencing... Figure 2 Multimodal feature decoupling data flow graph, refer to Figure 3 Spatial bias attention fusion and compensatory decoding flow graph.

[0046] Furthermore, a spatial smoothing filter with high and low thresholds is used to smooth the initial noise matrix and output a degraded matrix. This process includes: setting a first confidence threshold and a second confidence threshold, where the value of the first confidence threshold is greater than the value of the second confidence threshold; traversing each matrix element in the initial noise matrix, assigning a damaged state label to matrix elements with values ​​greater than the first confidence threshold, and assigning a normal state label to matrix elements with values ​​less than the second confidence threshold; for matrix elements with values ​​between the first and second confidence thresholds, extracting the state labels of adjacent matrix elements in the spatial neighborhood of the matrix element, and determining the state label of the matrix element according to the majority voting principle; and outputting the initial noise matrix with state labels as a degraded matrix.

[0047] Specifically, after acquiring the legal case file image, it is divided into multiple local grids of size 16×16 pixels. The statistical variance of the pixel grayscale within each local grid is calculated, satisfying the formula: ,in The statistical variance is... The total number of pixels within the local grid. For the local grid, the first grayscale value of each pixel. The variance is the average grayscale value of all pixels within the local grid. The variance values ​​are then arranged and concatenated according to the grid space to form the initial noise matrix. A first confidence threshold of 0.85 and a second confidence threshold of 0.35 are preset. The higher upper limit of 0.85 aims to prevent misjudging normal bold headings or dark shadows in the file as severely damaged, while the lower lower limit of 0.35 ensures that large areas of clean background can be quickly processed with low computing power. The initial noise matrix is ​​iterated. If the matrix element value is greater than 0.85, it indicates that the corresponding grid has severe stains or dark creases on the file, and its degradation state is marked as damaged. If the value is less than 0.35, it indicates that the grid is in a clean page background area and is marked as normal. For elements in the ambiguous zone with values ​​between 0.35 and 0.85, the degradation state of their eight neighboring elements in their spatial physical neighborhood is extracted. If five or more of the eight neighboring elements are marked as damaged, then according to the majority voting principle (this stringent voting condition of more than 50% means that the central mesh is physically topologically partially or fully surrounded by the damaged area, conforming to the continuous spread law of ink wash, which can effectively close the edge of the real stain and prevent the normal character boundary from being damaged due to excessive expansion), the central element is also forcibly marked as damaged; otherwise, it is marked as normal. Finally, all processed state sets are output as a degradation matrix.

[0048] This invention effectively eliminates isolated noise interference caused by uneven ink coverage or yellowing paper in legal case file images by introducing a hysteresis comparison mechanism with both high and low thresholds, combined with a majority voting rule based on spatial neighborhood. This mechanism ensures the spatial continuity of the output degradation matrix, preventing high-frequency switching and signal oscillations during subsequent modal feature weight allocation.

[0049] Furthermore, combining the prior color distribution of the document image to be processed in the HSV color space with the degradation matrix, a physical gamut soft isolation mask is generated, including: extracting the target hue channel value and target saturation channel value of each pixel in the document image to be processed in the HSV color space; calculating the first absolute difference between each target hue channel value and the preset stamp reference hue channel value, and the second absolute difference between each target saturation channel value and the preset stamp reference saturation channel value; using a preset weighting coefficient to perform a weighted summation of the first absolute difference and the second absolute difference to obtain a basic color difference matrix; calculating the difference between a preset color distance upper limit threshold and the basic color difference matrix to obtain a color similarity matrix; multiplying the color similarity matrix by a preset response gain coefficient, and performing point-by-point multiplication of the result with the elements of the corresponding coordinates of the degradation matrix, inputting it into a logistic activation function for numerical mapping, outputting a continuous probability matrix with element values ​​in the closed interval of zero to one, and using the continuous probability matrix as the physical gamut soft isolation mask.

[0050] Specifically, addressing the common issue in legal case files where red official seals obscure black printed text, a preset range of hue channel values ​​(0 to 10) and saturation channel values ​​(0.65) for the seal were obtained. This range strictly anchors to the typical spectral characteristics of standard red inkpads used in domestic official documents. The hue range of 0 to 10 effectively locks in the pure red to vermilion frequency band, while the saturation value of 0.65 fully accommodates color desaturation caused by fading of inkpads over time or attenuation of scanning light. After converting the case file image to the HSV color space, the target hue channel values ​​and target saturation channel values ​​were extracted pixel by pixel.

[0051] The first absolute difference between the target hue channel value and the preset stamp reference hue channel value at point (i,j) in the document image to be processed is calculated. Specifically, the calculation method is as follows: Take the normal absolute difference between the target hue channel value and the reference hue channel value, and the minimum of the two (the maximum range value of the hue channel minus this normal absolute difference) as the first absolute difference. Then, calculate the second absolute difference between the target saturation channel value and the preset stamp reference saturation channel value. Finally, use a preset weighting coefficient to perform a weighted sum of the first and second absolute differences to obtain the basic color difference matrix. The calculation formula is as follows:

[0052]

[0053] in, The values ​​of the elements of the basic color difference matrix at coordinate point (i,j); and These represent the actual target hue channel value and the actual target saturation channel value at that point, respectively. The preset base hue channel value for the seal; Set the preset stamp baseline saturation channel value; The first preset weighting coefficient is 0.6; The second preset weighting coefficient is set to 0.4. This weighting allocation emphasizes the dominant role of hue in color identification, while assigning appropriate auxiliary weight to saturation. This ensures that the red stamp is accurately located while maintaining sufficient tolerance for uneven ink application or reflective phenomena.

[0054] To generate a physical color gamut soft isolation mask, a preset color distance upper limit threshold is established (e.g., set to 1.0, which corresponds to the maximum theoretical color difference in the normalized color space; by subtracting this constant, the distance metric can be perfectly reversed into a similarity metric representing the probability that a pixel belongs to the stamp, achieving mathematical unification of the logical direction). The basic color difference matrix is ​​then subtracted from this upper limit threshold to transform the color difference data into a color similarity matrix, thus reversing the distance metric into a similarity metric. Furthermore, the color similarity matrix is ​​multiplied by a preset response gain coefficient (e.g., set to 5.0; this gain value is set to match the sensitive region of the subsequent logistic activation function; the significant amplification by multiplying by 5.0 stretches small color similarity differences to the saturation boundary of the activation function), thereby amplifying the numerical distinguishability between stamp pixels and background pixels. Subsequently, the amplified color similarity matrix is ​​multiplied point-by-point with the elements corresponding to the physical coordinates in the degraded matrix output from the pre-processor, generating a fused feature matrix. To ensure stable addressing of subsequent convolutional kernels, each element of the fused feature matrix is ​​substituted into the logistic activation function for nonlinear numerical mapping, thus smoothly mapping the original features to a closed interval between 0 and 1. At this point, the actual stamp pixels, after gain amplification, will approach a high activation state of 1.0, while the background pixels will completely approach 0. The calculation process satisfies the following formula:

[0055] in, This is a color similarity matrix; The preset upper limit threshold for color distance is 1.0; The basic color difference matrix; For fusion feature matrix; The preset response gain coefficient has a preferred value of 5.0; This represents the Hadamard product, i.e., the point-by-point matrix multiplication operation; This is the degradation matrix of the pre-output; The output is the continuous probability matrix of the physical color gamut soft isolation mask; This represents the logistic activation function.

[0056] Through this reverse similarity mapping mechanism, the real stamp pixels, due to their small initial color difference, appear as extremely positive numbers after being subtracted from and amplified by the upper limit threshold. After activation by the logistic function, their mask value will approach 1.0 (which is necessarily higher than the mask activation threshold of 0.75), and thus be accurately sent to the non-zero coordinate hash table for processing.

[0057] This invention achieves soft isolation of visually overlapping areas by quantizing the multidimensional distance between image pixels and standard stamp colors in the HSV color space and introducing a logistic function for smooth probability mapping. This method abandons the traditional hard clipping of binarization, and while separating the red stamp interference, it fully preserves the edge topological gradient information of the underlying black Chinese character strokes, significantly reducing the misrecognition rate of text in overlapping areas.

[0058] Generating a physical color gamut soft isolation mask further includes: calculating the pixel spatial gradient of the target saturation channel value in the horizontal and vertical directions, specifically satisfying the formula: , , ,in coordinate point The horizontal gradient component at that location, The gradient component is in the vertical direction. For the target saturation channel, The calculated pixel spatial gradient value is used to further determine that when the pixel spatial gradient value in a continuous pixel region is within a preset gradient range, the continuous pixel region is determined to be a stamp edge shading region. Further, in the process of outputting the physical gamut soft isolation mask using the logistic activation function, the pixel elements in the stamp edge shading region are numerically mapped using the first logistic mapping parameter, and the pixel elements in the non-shading region are numerically mapped using the second logistic mapping parameter, and the updated physical gamut soft isolation mask is output.

[0059] Specifically, after generating the basic color difference matrix, the gradient operator is called to calculate the pixel spatial gradient of the target saturation channel value in the horizontal X-axis and vertical Y-axis directions. When the pixel spatial gradient value in a certain continuous pixel region is detected to be within the preset gradient range of 0.15 to 0.45, the region is determined to be the smudged area at the boundary between the red stamp and the white paper background. When performing the logistic activation mapping, a strategy is employed: For pixel elements in the saturated region, the first logistic mapping parameter (e.g., a gain coefficient set to 2.5, which allows the logistic curve to maintain a relatively gentle slope in the transition zone, preventing the mask value from entering the saturation region too early, thus preserving the grayscale step details of the underlying text in the saturated region with a non-linear soft mapping, preventing the character stroke edges from being crudely truncated) is used to preserve the grayscale details of the underlying text with a lower compression intensity; for solid stamp pixels in the non-saturated region, the second logistic mapping parameter (e.g., a gain coefficient set to 8.0, which makes the activation function curve steeply near the input point, exhibiting a step function-like characteristic, thus ensuring that solid stamp pixels can be forcibly mapped to the high-confidence activation saturation region, achieving complete shielding and physical isolation of high-concentration pigment areas) is used to achieve complete pigment isolation with a high compression intensity, and finally outputting the updated physical gamut soft isolation mask.

[0060] This invention refines the processing accuracy of the isolation mask for edge transition zones by introducing a spatial gradient detection mechanism. This method effectively solves the technical problem commonly encountered in legal case files where the diffusion of red ink from seals obscures the key starting and ending points of characters. Through asymmetric mapping parameter adjustment, while completely isolating the core color of the seal, the topological features of the character edges are preserved to the maximum extent, avoiding character morphological degradation caused by hard cutting of the mask, and providing a higher fidelity visual input for subsequent text feature stream extraction.

[0061] Furthermore, by controlling asymmetric sparse convolution through the physical color gamut soft isolation mask, the document image to be processed is decoupled into a text feature stream and a non-text feature stream. This includes: extracting the target coordinate positions in the physical color gamut soft isolation mask whose values ​​are greater than the mask activation threshold to construct a non-zero coordinate hash table; performing a sparse convolution operation of the first kernel size dimension on pixel regions within the non-zero coordinate hash table to extract stamp texture features as a non-text feature stream; and performing a dense convolution operation of the second kernel size dimension on pixel regions not within the non-zero coordinate hash table to extract character stroke features as a text feature stream, while restricting the value of the second kernel size dimension to be less than the value of the first kernel size dimension.

[0062] Specifically, the preset mask activation threshold is 0.75. This threshold is chosen to match the nonlinear mapping characteristics of the logistic activation function in the forward computation. With the similarity feature input after gain amplification, 0.75 serves as the decision boundary, ensuring that only pixels entering the high-confidence saturation region of the activation function are identified as the seal area. This physically achieves accurate extraction of the core area of ​​the red seal in the legal case file, effectively filtering background noise or dark text edges at the activation edge. The physical color gamut soft isolation mask is scanned to accurately extract target coordinates with mask values ​​greater than 0.75. These coordinates physically correspond to the core area of ​​the red seal in the legal case file. These scattered coordinates are stored in a non-zero coordinate hash table. During feature extraction, the network employs a splitting mechanism. For the pixel areas recorded in the non-zero coordinate hash table, a first sparse convolution kernel with a kernel size of 5×5 is called to perform convolution calculations, specifically satisfying the formula: ,in The output feature map at coordinate points The characteristic values ​​at that location, The spatial weight parameters of the first sparse convolution kernel are... The local feature tensor is input to the convolution kernel. and The two-dimensional spatial bias index of the convolution kernel ranges from -2 to 2. Due to the complexity and global coherence of the seal texture, a larger convolution kernel can fully capture the seal's anti-counterfeiting texture and edge curvature, extracting seal texture features that constitute a non-textual feature stream. Conversely, for large areas of ordinary pixels not recorded in the hash table, a second dense convolution kernel with a kernel size of 3×3 is used to perform the operation, specifically satisfying the formula: ,in The output feature map at coordinate points The characteristic values ​​at that location, The spatial weight parameters of the second dense convolutional kernel are... The local feature tensor is input to the convolution kernel. and The two-dimensional spatial bias index is used for the convolution kernel, with values ​​ranging from -1 to 1. Since the main text of legal case files is mostly in small Song or Fangsong fonts, smaller convolution kernels can more sensitively capture subtle stroke features such as horizontal, vertical, and diagonal strokes, and the extracted stroke features constitute the text feature stream. This operation of assigning convolution kernels of unequal sizes to different regions achieves physical dimensionality reduction and decoupling of the feature stream.

[0063] This invention utilizes the probability distribution of the isolation mask as a hash index to achieve asymmetric splitting of tensor operations at the feature extraction layer. This mechanism performs sparse operations with a large receptive field only on the extremely small area covered by the stamp, while using fine-grained dense operations on the large text area. This successfully decouples the two types of visual features while significantly reducing unnecessary floating-point operations and lowering peak memory overhead.

[0064] Further, the two-dimensional topological anchor points of each text unit in the text feature stream are extracted, the Manhattan distance between each two-dimensional topological anchor point is calculated, and mapped to a spatial bias matrix. This includes: performing a connected component pixel aggregation algorithm on the text feature stream to divide it into independent text units and generate bounding rectangles for each text unit; extracting the absolute physical coordinates of the intersection points of the diagonals of the bounding rectangles of each text unit in the text feature stream as two-dimensional topological anchor points; calculating the absolute difference of the first coordinate in the horizontal direction and the absolute difference of the second coordinate in the vertical direction between any two two-dimensional topological anchor points, and adding the two to obtain the Manhattan distance between the text units; inputting the Manhattan distance into a preset negative penalty mapping function to obtain negative bias values ​​that are negative and whose absolute values ​​increase exponentially with the increase of physical spatial distance, and constructing a spatial bias matrix from all negative bias values.

[0065] Specifically, after acquiring the text feature stream, a connected component analysis algorithm is used to aggregate the pixels in the feature stream, extracting the minimum bounding rectangle bounding box of independent text units such as "plaintiff information block," "defendant information block," or "judgment result block." The intersection of the diagonals of this bounding rectangle is extracted to obtain its absolute physical coordinates in the two-dimensional coordinate system of the case file image. To eliminate the computational scale differences caused by images of different resolutions and to prevent numerical underflow during subsequent high-dimensional tensor exponential operations, the horizontal and vertical coordinate values ​​of the above absolute physical coordinates are normalized by division using the global pixel width and global pixel height of the document image to be processed. After dimensionless processing, the relative coordinates with values ​​strictly limited to the closed interval between 0 and 1 are obtained, and these relative coordinates are used as the two-dimensional topological anchor points of the text unit.

[0066] To quantify the spatial relationships between text blocks, two arbitrary text units are selected, and the spatial distance between any two two-dimensional topological anchor points is calculated. The normalized Manhattan distance is obtained by adding the absolute difference of their first relative coordinates along the horizontal X-axis to the absolute difference of their second relative coordinates along the vertical Y-axis. The calculation formula is as follows:

[0067]

[0068] in, The normalized Manhattan distance between the first and second text units; The relative coordinates of the two-dimensional topological anchor point of the first text unit; The relative coordinates of the two-dimensional topological anchor point of the second text unit.

[0069] To construct an effective soft attention mask, the calculated normalized Manhattan distance is used as the independent variable x and input into a predefined negative penalty mapping function. In order to achieve strong suppression of long-distance irrelevant information while ensuring gradient smoothness, the function... The scale-controlled exponential mapping logic is employed, which satisfies the mathematical relation:

[0070]

[0071] Where e is the natural constant, The preset exponential smoothing factor is set to 1.5 in this embodiment. As a negative baseline control coefficient, it is set to 10 in this embodiment. Through this mathematical logic, when the physical distance between text units is close, the bias value... It extends slowly toward the negative, maintaining a soft attentional connection; while the bias value increases with distance. The negative bias increases exponentially in a controlled manner until it reaches the negative saturation range (approximately -190 at the maximum distance). This magnitude, after entering the Softmax operation of the self-attention layer, effectively suppresses the weights of distant, unrelated text blocks to near zero, while strictly avoiding floating-point underflow and gradient vanishing problems caused by excessively large bias values ​​(e.g., exceeding -1000), thus ensuring the convergence stability of the model during the training phase. Finally, the negative bias values ​​between all pairs of text units within the dossier are constructed into a two-dimensional spatial bias matrix.

[0072] This invention transforms the layout information in the two-dimensional layout of legal case files into precise mathematical measurements by extracting the geometric center of the circumscribed rectangle and calculating the physical Manhattan distance. Combined with a negative penalty mapping function, it forcibly generates a constraint matrix that conforms to visual reading patterns, providing a reliable physical quantification basis for subsequently severing erroneous semantic associations between distant but unrelated text blocks.

[0073] Further, the spatial bias matrix is ​​injected into the multimodal attention network, and the fusion weights of each modality feature in the multimodal attention network are adjusted based on the degradation matrix. This includes: in the self-attention computation layer of the multimodal attention network, the spatial bias matrix is ​​used as a penalty term matrix, and matrix addition is performed on the product of the query vector matrix and the key vector matrix; after the degradation matrix is ​​binary numerically mapped, a global average pooling operation is performed to obtain a global degradation coefficient value; the text modality preservation coefficient value is obtained by subtracting the global degradation coefficient value from the constant 1. In order to prevent the coefficient value from being canceled by the layer normalization operation in the forward propagation of the multimodal attention network, the text feature stream is first processed by self-attention computation and layer normalization to output a feature tensor, and then the text modality preservation coefficient value is multiplied point by point with the elements of the output feature tensor to complete the adjustment of the fusion weights.

[0074] Specifically, when text feature streams and non-text feature streams enter the multimodal attention network for cross-modal interaction, the network performs a dot product operation between the query vector matrix and the key vector matrix to calculate the initial self-attention score, which specifically satisfies the formula: ,in The initial self-attention score matrix is... The query vector matrix, The key vector matrix, This is a matrix transpose operation. In this calculation step, the previously generated spatial bias matrix is ​​directly used as the penalty term matrix and element-wise added to the initial self-attention score matrix. Since the bias values ​​in the spatial bias matrix are mapped to negative numbers whose absolute values ​​increase exponentially with increasing physical spatial distance, after performing matrix addition, the initial self-attention scores between text units that are physically far apart will be forcibly pulled down to the deep negative range. Consequently, during the normalization activation function processing, the attention response activation values ​​corresponding to these deep negative values ​​are forcibly zeroed out by the mathematical mechanism. Simultaneously, a global average pooling operation is performed on the degraded matrix containing the degraded state of the entire legal case file. Specifically, before performing the pooling calculation, a numerical mapping operator is pre-constructed to map the normal state labels in the degraded matrix to the value 0 and the damaged state labels to the value 1, thereby converting the discrete label matrix into a binary numerical tensor. Subsequently, a global average pooling operation is performed on this binary numerical tensor, and the calculation formula is as follows:

[0075] in, The calculated text modality preservation coefficient values; This is the global degradation factor obtained through global average pooling operation; The text feature tensor after weight adjustment; This is the input tensor for the initial text feature stream. In this way, after receiving a signal that the overall quality of the case file has been severely degraded, the network automatically and uniformly adjusts the representation weights of the text modalities in the overall feature fusion process.

[0076] This invention effectively prevents word order disorder issues in complex multi-page or multi-column layouts by rigidly injecting a spatial bias matrix into the underlying formula for calculating self-attention scores. Simultaneously, by dynamically scaling the text feature tensor based on a global degradation coefficient, the network can adaptively adjust its dependence on underlying text visual features according to the objective physical degradation of the image, enhancing the robustness of the fusion mechanism.

[0077] Injecting the spatial bias matrix into the multimodal attention network and adjusting the fusion weights of each modality feature in the multimodal attention network based on the degradation matrix further includes: during the self-attention forward computation of the multimodal attention network, extracting the two-dimensional topological anchor coordinates of the first text unit corresponding to the current query vector and the two-dimensional topological anchor coordinates of the second text unit corresponding to the current key vector, calculating the two-dimensional physical space relative direction vector from the first text unit to the second text unit; and performing a dot product calculation between the two-dimensional physical space relative direction vector and the degradation gradient increase direction vector of the initial noise matrix.

[0078] A multimodal attention network is constructed, employing an encoder-decoder architecture. The encoder consists of six stacked standard Transformer encoding layers, each containing a multi-head self-attention sub-layer (8 attention heads, 512 hidden dimensions) and a feedforward sub-layer (2048 dimensions, GELU activation function). Each sub-layer is followed by residual connections and layer normalization. The text feature stream and the non-text feature stream aligned to the linear projection dimension are concatenated along the channel dimension and input into the encoder. The spatial bias matrix is ​​element-wise added to the query-key dot product score matrix in each self-attention sub-layer of the encoder.

[0079] Specifically, during the self-attention forward computation of the multimodal attention network, the two-dimensional directional derivative vector of each matrix element in the initial noise matrix, which contains continuous variance values, is calculated in real time. The direction of this vector represents the gradient increase direction of local image noise from mild to severe. Simultaneously, to achieve the mapping of high-dimensional semantic features to the underlying physical space, the original text entities corresponding to the current query vector and key vector at the input end are traced, and the two-dimensional topological anchor coordinates of the first text unit corresponding to the current query vector and the second text unit corresponding to the current key vector are extracted. Based on the above two-dimensional coordinates, the two-dimensional physical space relative direction vector from the first text unit to the second text unit is calculated. Subsequently, the two-dimensional physical space relative direction vector and the degradation gradient increase direction vector of the initial noise matrix are calculated as a two-dimensional vector dot product of the same dimension. If the dot product value is greater than 0, it indicates that the current attention search behavior is attempting to diverge towards a more noisy region. At this time, a division decay operation (such as dividing by a decay factor of 5.0) is immediately performed on the corresponding self-attention score matrix element to forcibly reduce the weight of this path.

[0080] In legal case file processing, when faced with localized water stains or ink smudges, this method forces the model to shift its attention from blurry areas back to clear contextual areas. By utilizing the continuous directional derivatives of the initial noise distribution as navigation logic, it overcomes the limitation of discrete masks in providing fine-grained guidance, achieving intelligent correction of the self-attention distribution. This ensures that the reconstruction logic is not skewed by local extreme noise, preventing the penetration of local errors into the global semantic reconstruction process at the algorithm execution level, and enhancing the reliability of resolution under poor image quality.

[0081] Furthermore, the weights of the text feature stream are reduced in the multimodal attention network, and a visual language model is invoked to perform feature compensation on local regions to generate a fused feature tensor. This includes: determining that a continuous pixel region is a damaged local region when the degradation value of a continuous pixel region in the degradation matrix is ​​greater than a third confidence threshold; extracting the context feature vector of the text unit surrounding the damaged local region and obtaining the global visual feature map of the document image to be processed, and inputting the two into the visual language model to generate a one-dimensional semantic compensation tensor; performing channel dimension mapping on the one-dimensional semantic compensation tensor, and performing spatial broadcasting operation on the mapped one-dimensional semantic compensation tensor based on the two-dimensional physical coordinates of the damaged local region to generate a two-dimensional local compensation feature map.

[0082] Finally, feature alignment and fusion are achieved through a cross-attention mechanism. Specifically, the concatenated tensor of the weighted text feature stream and the non-text feature stream is used as the query vector input; the two-dimensional local compensation feature map is linearly mapped to generate key vectors and value vectors as inputs. The query vector is used to calculate the spatial semantic relevance score in the key vector, and the value vector is weighted and aggregated accordingly. This accurately aligns and injects the high-dimensional semantic compensation information into the underlying visual feature tensor, generating a fused feature tensor.

[0083] The visual language model employs a multimodal pre-trained model based on a cross-attention architecture, where the visual encoder and text decoder are aligned via a cross-modal projection layer. Specifically, the damaged local region is cropped into fixed-size RGB image blocks, which are then concatenated with the character sequence of the surrounding parsed text units in the format "[CLS]surrounding context: [TEXT][SEP]damaged region image: [IMG][SEP]" to form a multimodal input prompt. The VLM's text decoder outputs a hidden state sequence, and the hidden vector corresponding to the last layer's [CLS] marker is taken as the one-dimensional semantic compensation tensor, whose dimension is consistent with the channel dimension of the text feature stream. During the training phase, the VLM maintains frozen pre-trained weights and only serves as a feature extractor to provide prior semantic compensation.

[0084] Specifically, the preset third confidence threshold is determined based on a balance analysis of empirical statistics on the reconstruction accuracy of the sample dataset and computational resource consumption, and is preferably 0.90. When the degradation value is below 0.90, it indicates that although there is local contamination in the image, the underlying convolutional neural network can still extract the remaining character stroke boundary features. At this time, reconstruction can be completed by relying on the conventional feature fusion of the multimodal attention network, thereby avoiding the high memory overhead and potential semantic illusion risks caused by frequent calls to large visual language models. When the degradation value reaches or exceeds 0.90, it indicates that the physical pixel area has been catastrophically damaged (such as complete coverage by dark ink or local physical tearing of paper), and the underlying visual topological features have been irreversibly lost. At this time, setting the high confidence threshold of 0.90 can accurately intercept such extreme failure areas and force the upper-layer visual language model to perform global semantic prior compensation. The setting of this threshold effectively controls the computation triggering frequency of complex large models while ensuring the recall rate of local extreme noise repair.

[0085] When a continuous 50×50 pixel region in a legal case file image is detected to have a degradation value exceeding 0.90 due to physical tearing, it is identified as a damaged local region. The global contextual feature vectors of successfully parsed text units surrounding this region are extracted and, along with the macroscopic layout feature map of the entire case file, are fed into the visual language model as input prompts. The visual language model outputs a one-dimensional semantic compensation tensor. To address the dimensionality mismatch between the one-dimensional semantic features and the underlying two-dimensional physical feature map, a linear projection layer maps this one-dimensional semantic compensation tensor to a channel dimension consistent with the text feature flow. Subsequently, based on the two-dimensional physical coordinate parameters of the damaged local region, a spatial broadcast operation is performed on the mapped tensor, expanding and upsampling it in the spatial dimension to generate a two-dimensional local compensation feature map aligned with the pixel size of the damaged region. Finally, through a cross-attention mechanism, this two-dimensional local compensation feature map is fused and aligned with the pre-weighted text feature flow and non-text feature flow.

[0086] This invention constructs a mathematical bridge for the reverse projection from high-dimensional abstract semantics to the underlying local physical space by introducing channel mapping and spatial broadcasting mechanisms. This approach overcomes the technical obstacle that prevents the direct fusion of the token sequence output by the visual language model and the feature map output by the convolutional neural network due to their different topological dimensions. By accurately "tiling" one-dimensional semantic information back into the damaged two-dimensional coordinate space, this invention achieves precise repair of broken logical chains while preserving spatial topological properties, ensuring the integrity of the reconstructed legal document structure.

[0087] Furthermore, decoding is performed based on the fused feature tensor to output a structured semantic sequence, including: inputting the fused feature tensor into the converter decoding network to output a basic text character sequence and paragraph layout labels.

[0088] The underlying text character sequence is semantically mapped using a pre-trained deep learning semantic encoder (such as the BERT-Base model based on a transformer architecture). Specifically, the character sequence of each sentence is input into the semantic encoder, which extracts global contextual features through its internal multi-layer self-attention mechanism and outputs a dense sentence vector of fixed dimensions (such as 768 dimensions).

[0089] After obtaining the sentence vectors, the cosine similarity value of each sentence vector in the basic text character sequence is calculated using the Local Sensitive Hashing (LSH) algorithm. Directed connections are established between target sentence pairs whose cosine similarity values ​​are greater than the reference threshold, and the basic text character sequence is constructed into a heterogeneous document attribute graph. Graph convolution feature aggregation is performed on the heterogeneous document attribute graph, and the aggregated node features are serialized and recombined according to the hierarchical relationship of the paragraph layout tags to output a structured semantic sequence.

[0090] The heterogeneous document attribute graph is subjected to graph convolutional feature aggregation, specifically including: initializing the edge weights of the graph attention mechanism for the established directed connections; each layer of the graph convolutional network aggregates the features of the first-order neighbor nodes of the target node and performs feature mapping using a weight matrix with a leaky linear rectified activation function, updating the node features, and repeating at least two layers of graph convolutional operations to complete deep feature fusion. Subsequently, the aggregated node features are used as tree nodes, with the document title or case number node as the root node, and a syntax tree is constructed according to the hierarchical relationship of the paragraph formatting tags. The features of each node in the syntax tree are serialized and expanded using a preorder traversal algorithm to output a structured semantic sequence.

[0091] The locality-sensitive hashing retrieval algorithm employs a cosine locality-sensitive hashing scheme based on random projection. Specifically, k unit Gaussian vectors with the same dimension as the sentence vector are randomly generated as hash projection vectors, and each hash table corresponds to a set of projection vectors. The number of hash tables L is set to 10, and the hash bit length k of each hash table is set to 16 to achieve a balance between recall and computational efficiency.

[0092] Before practical application of the system, the network needs to be jointly trained using a document image dataset labeled with real character sequences, layout classification tags, and cross-reference graph pairs. The total loss function is defined as a multi-task joint loss function, which includes three components: the first component is the sequence cross-entropy loss used to constrain character recognition accuracy; the second component is the class cross-entropy loss used to constrain paragraph layout tag classification accuracy; and the third component is the contrastive learning loss used to constrain edge connectivity in heterogeneous document attribute graphs. During the training phase, the gradient of the total loss function with respect to the multimodal attention network, converter decoding network, and graph convolutional network is calculated using the backpropagation algorithm, and the network weights are updated using the AdamW optimizer.

[0093] Specifically, a training dataset was constructed: 10,000 pages of high-resolution scanned document images covering types such as legal case files, bidding documents, and academic papers were collected. Each page image was labeled with three sets of tags by manual annotation: (1) character-level text sequence; (2) paragraph-level layout category tags (case number, title, body text, table, etc.); and (3) sentence-level cross-reference pairs. The dataset was divided into training set, validation set, and test set in an 8:1:1 ratio.

[0094] During the training phase, the weight ratios of each component in the multi-task loss function are set as follows: : : =1.0:0.5:0.3. The AdamW optimizer is used, with the initial learning rate set to... The weight decay coefficient is 0.01, the batch size is 16, and the training epochs are 50. A cosine annealing decay strategy is used for the learning rate, and training is terminated early if the validation set loss does not decrease for five consecutive epochs. The visual language model can maintain pre-trained weights frozen to stably provide high-dimensional semantic compensation priors. After iterative training until the total loss converges, the structured text reconstruction model is completed.

[0095] Specifically, a fused feature tensor containing multimodal information is input into the converter decoding network. The decoding network, through layer-by-layer reasoning, outputs a basic text character sequence containing the specific content of the legal case file, along with paragraph formatting labels corresponding to each sentence (such as "Case Number," "Plaintiff," "Claims," ​​and "Facts and Reasons"). Due to the large number of cross-references in the legal case file, a locality-sensitive hashing (LSH) retrieval algorithm is used to transform the basic text character sequence into high-dimensional sentence vectors, and the cosine similarity between each sentence vector is quickly calculated. A preset citation threshold of 0.75 is used (determined through grid search on the development set, aiming to maximize the F1 score, based on empirical analysis of semantic similarity statistical distribution). When the cosine similarity between two sentences (such as "See Appendix II of the Evidence List" and the actual appendix content paragraph) is greater than 0.75, a substantial citation association is determined, and a directed connection edge is established between the target sentence pair in the data dimension. In this way, the basic text character sequence, originally presenting a one-dimensional linear discrete state, is woven and reconstructed into a heterogeneous document attribute graph. Finally, a graph convolution feature aggregation operation is performed on the graph, allowing mutually referencing node features to be passed and fused along the directed edges. After the graph convolution calculation is completed, the node features are reorganized from top to bottom into a tree structure based on the hierarchical attributes of the previously output paragraph layout tags, ultimately outputting a structured semantic sequence containing complete nested hierarchies and logical references.

[0096] This invention overcomes the context distance limitation in converting long documents by employing locality-sensitive hashing retrieval and graph convolution aggregation mechanisms. This mechanism accurately captures and binds the cross-reference logic between clauses without increasing the computational complexity of the sequence. The resulting structured sequence is not only accurate in content but also highly restores the rigorous structure of the original document in terms of logical hierarchy.

[0097] Example 2:

[0098] This embodiment uses the automated review of bidding documents in the multinational industrial manufacturing sector as an application scenario. Such documents typically contain dense engineering parameters, complex financial quotation tables, multi-level list indexes, and numerous colored seals and company stamps. Because these documents are often electronically generated from multiple batches of scanning, they commonly suffer from paper folding, perspective distortion at the edges, and localized text blurring due to uneven moisture absorption. This embodiment will describe in detail how the present invention accurately reconstructs such heterogeneously degraded two-dimensional images into a structured semantic sequence.

[0099] As one embodiment of the present invention, refer to Figure 1 A flowchart of a structured text reconstruction method based on multimodal feature fusion, referencing... Figure 2 Multimodal feature decoupling data flow graph, refer to Figure 3 Spatial bias attention fusion and compensatory decoding flow graph.

[0100] After acquiring the image of the bidding document to be processed, the underlying physical environment perception program is first started. Specifically, the document image to be processed is divided into multiple local grids with the same pixel size; for each local grid, the statistical variance of the grayscale values ​​of all pixels within it is calculated. This variance reflects the degree of physical degradation in that local area, thereby generating an initial noise matrix corresponding to the spatial structure of the original image.

[0101] To eliminate isolated noise and smooth physically degraded signals, a spatial smoothing filter with a first confidence threshold and a second confidence threshold is invoked. Each element in the initial noise matrix is ​​traversed; if the value of the current element is greater than the first confidence threshold, the degradation state at that location is marked as damaged; if the value is less than the second confidence threshold, it is marked as normal. For intermediate elements between the two thresholds, the degradation states of their eight neighboring elements in their spatial neighborhood are extracted, and their final state is determined by majority voting. This process effectively filters out random noise during the scanning process, ultimately outputting a target degradation matrix that objectively reflects the local physical quality of the bidding documents.

[0102] In the feature decoupling stage, the color distribution prior of the HSV color space is used to handle stamp interference. The target hue channel value and target saturation channel value of each pixel in the document image to be processed are extracted in the HSV color space, and the first absolute difference and the second absolute difference between them and the preset red stamp reference value are calculated respectively. The above differences are linearly summed by preset weighting coefficients to generate the basic color difference matrix.

[0103] To address the text feature damage caused by red ink smudges at the edges of official seals commonly found in bidding documents, a refined edge protection logic is implemented. The pixel spatial gradients of the target saturation channel values ​​in the horizontal and vertical directions are calculated. When a pixel spatial gradient value within a continuous pixel region is detected to be within a preset gradient range, that region is determined to be a seal edge smudged area. Subsequently, the basic color difference matrix and the degradation matrix are multiplied point-by-point and substituted into the logistic activation function. During the mapping process, a first mapping parameter that preserves stroke grayscale details is applied to pixel elements in the seal edge smudged area, while a second mapping parameter with high isolation strength is applied to pixel elements in non-smudged areas. The final output is a physical color gamut soft isolation mask with element values ​​within a closed range of zero to one.

[0104] Subsequently, a physical color gamut soft isolation mask is used to control the addressing of the underlying convolutional operations. Coordinates with values ​​greater than the activation threshold are extracted from the mask to construct a non-zero coordinate hash table. For the stamp-covered area within the hash table, a first sparse convolutional kernel with a larger kernel size is used to extract non-textual feature streams containing stamp texture and anti-counterfeiting features. For regular areas not within the hash table, a second dense convolutional kernel with a smaller kernel size is used to extract textual feature streams containing fine character strokes. This asymmetric feature extraction strategy completely isolates stamp interference while protecting the feature integrity of the bidding document.

[0105] After entering the topology reconstruction stage, the system first corrects common scanning skew and binding deformation issues in bidding documents. A connected component analysis algorithm is used to extract the minimum bounding box of each text unit in the text feature flow. To handle paper bending caused by binding lines, the coordinates of the bottom center points of multiple vertically arranged text units are extracted. A one-dimensional quadratic polynomial function regression calculation is performed to obtain curve parameters reflecting perspective deformation, and based on this, affine transformation correction is applied to the two-dimensional topological anchor point coordinates of each text unit. Based on the corrected coordinates, the Manhattan distance between any two text units is calculated. To transform the physical distance into a mathematical constraint for the deep network, this Manhattan distance is input into a preset negative penalty mapping function. This function maps the zero point of distance to a numerical value of zero and maps the independent variable that increases with distance to a negative number whose absolute value increases exponentially, thus generating a negative bias value. The system constructs a two-dimensional spatial bias matrix from the negative bias values ​​between all text units.

[0106] In feature fusion within a multimodal attention network, a dual intervention is implemented. First, a spatial bias matrix is ​​injected as a penalty term into the self-attention computation layer, forcibly suppressing attention weights between physically distant, unrelated text blocks. Second, a degradation matrix is ​​used to guide the attention flow in real-time. The two-dimensional directional derivative vectors of each element in the degradation matrix are calculated to determine the direction of increase in noise gradients. In the forward attention computation, the spatial relative direction vector between the query vector and the key vector is calculated, and a dot product is performed with the noise gradient direction vector. If the dot product value is greater than zero, it indicates that the model is attempting to direct attention to more severely damaged noise areas, and a division decay operation is immediately performed to reduce the response of that path. Simultaneously, global average pooling of the degradation matrix dynamically adjusts the overall fusion ratio of the text modalities.

[0107] For situations where clauses in bidding documents are lost due to physical damage, the system activates a neural compensation mechanism based on spatial back projection. When the degradation matrix shows that the degradation value of a continuous region exceeds the third confidence threshold, the region is identified as a damaged local area. The system extracts the global contextual feature vectors of the undamaged text units surrounding this region (such as the prefix "product name" and the suffix "unit price"), and feeds them, along with the document's global visual feature map, as multimodal input prompts into the visual language model.

[0108] A visual language model is used to generate a one-dimensional semantic compensation tensor. To achieve mathematical alignment of heterogeneous features, the system first uses a linear projection layer to transform the channel dimension of the one-dimensional semantic compensation tensor, aligning its channel count with the text feature flow tensor. Next, the system obtains the bounding box parameters of the damaged local region in the original two-dimensional coordinate system and performs a spatial broadcast operation on the transformed tensor based on these parameters, thereby converting the discrete semantic embedding vectors into a two-dimensional local compensation feature map with H×W spatial resolution. Subsequently, the system invokes a cross-attention mechanism to perform tensor alignment calculations on the two-dimensional local compensation feature map, the non-text feature flow, and the weighted text feature flow in a unified two-dimensional topological space, merging them to generate a fused feature tensor.

[0109] Finally, the structured decoding and logical verification stages are performed. The fused feature tensor is input into the converter decoding network, outputting a basic text character sequence and corresponding paragraph formatting tags. To reconstruct the complex indexing and referencing relationships in the bidding documents (e.g., "Technical requirements are detailed in Chapter X"), a locality-sensitive hashing (LSH) algorithm is used to calculate the cosine similarity between sentence vectors. Directed connections are established between sentence pairs that meet a threshold condition, thus constructing a heterogeneous document attribute graph. Graph convolution feature aggregation is performed on this graph, enabling deep fusion of the features of interconnected clauses. Finally, according to the hierarchical relationship of the paragraph formatting tags, the aggregated node features are serialized and recombined, ultimately outputting a structured semantic sequence containing hierarchical nesting relationships, complete logical references, and immunity to physical noise interference.

[0110] This embodiment ensures that industrial bidding documents can still achieve high-precision structured reconstruction when faced with complex formats and harsh quality degradation environments through a closed-loop processing of the entire chain, from underlying physical perception, feature decoupling, spatial constraints to high-dimensional semantic compensation.

[0111] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A structured text reconstruction method based on multimodal feature fusion, characterized in that, include: Step 1: Acquire the document image to be processed and calculate the local noise variance; perform smoothing processing using a spatial smoothing filter with high and low thresholds, and output the degradation matrix; Step 2: Combine the color distribution prior of the document image to be processed with the degradation matrix to generate a physical gamut soft isolation mask; By controlling asymmetric sparse convolution through the physical color gamut soft isolation mask, the document image to be processed is decoupled into text feature stream and non-text feature stream; Step 3: Extract the two-dimensional topological anchor points of each text unit in the text feature stream, calculate the spatial distance between each two-dimensional topological anchor point, and map it into a spatial bias matrix; inject the spatial bias matrix into the multimodal attention network, and adjust the fusion weights of the multimodal attention network based on the degradation matrix; Step 4: If the degradation matrix shows that the noise in a local area is higher than the preset third confidence threshold, then reduce the weight of the text feature stream in the multimodal attention network and perform feature compensation on the local area; if the degradation matrix shows that the noise in a local area is not higher than the third confidence threshold, then maintain the current fusion weight. The non-text feature stream, the weighted text feature stream, and the optional compensation features are fused to generate a fused feature tensor. Decoding is performed based on fused feature tensors to output structured semantic sequences; Step 1 utilizes a spatial smoothing filter with both high and low thresholds for smoothing, outputting a degraded matrix. This includes: calculating the local noise variance of the document image to be processed, generating an initial noise matrix; the spatial smoothing filter iterates through the values ​​of each matrix element in the initial noise matrix, assigning a damaged state label to matrix elements greater than a preset first confidence threshold, and assigning a normal state label to matrix elements less than a preset second confidence threshold; for matrix elements between the first and second confidence thresholds, extracting the state labels of adjacent matrix elements within the spatial neighborhood of the matrix element, and determining the state label of the matrix element according to a majority voting principle; and outputting the initial noise matrix with state labels as a degraded matrix. Step 4 involves reducing the weights of the text feature stream in the multimodal attention network and performing feature compensation on local regions to generate a fused feature tensor. This includes: determining that consecutive pixel regions are damaged local regions when the degradation value of a continuous pixel region in the degradation matrix is ​​greater than the third confidence threshold; extracting the context feature vectors of the text units surrounding the damaged local regions and the global visual feature map of the document image to be processed, and inputting them into the visual language model to generate a one-dimensional semantic compensation tensor; performing channel dimension mapping on the one-dimensional semantic compensation tensor, and performing spatial broadcasting operation on the mapped one-dimensional semantic compensation tensor based on the two-dimensional physical coordinates of the damaged local regions to generate a two-dimensional local compensation feature map; and inputting the two-dimensional local compensation feature map, the non-text feature stream, and the weighted text feature stream into a cross-attention mechanism for alignment and fusion to generate a fused feature tensor.

2. The structured text reconstruction method based on multimodal feature fusion according to claim 1, characterized in that, Step 2 combines the prior color distribution of the document image to be processed with the degradation matrix to generate a physical color gamut soft isolation mask. This includes: calculating the characteristic distance between the document image to be processed and the stamp color reference distribution in the HSV color space to generate a basic color difference matrix; calculating the difference between a preset upper limit threshold for color distance and the basic color difference matrix to obtain a color similarity matrix; multiplying the color similarity matrix by a preset response gain coefficient, and performing point-by-point multiplication of the result with the elements at the corresponding coordinates of the degradation matrix, then inputting the result into the logistic activation function for numerical mapping, using the formula... The fusion feature matrix generated after the point-by-point multiplication operation Each element value is smoothly compressed and mapped to the [0,1] interval to obtain the continuous probability matrix of the physical color gamut soft isolation mask. ;in, Two-dimensional physical coordinates is a natural constant; the output element is a continuous probability matrix with values ​​ranging from zero to one, which serves as a soft isolation mask for the physical color gamut.

3. The structured text reconstruction method based on multimodal feature fusion according to claim 1, characterized in that, In step 2, the asymmetric sparse convolution is controlled by the physical color gamut soft isolation mask to decouple the document image to be processed into a text feature stream and a non-text feature stream. This includes: extracting the target coordinate positions in the physical color gamut soft isolation mask with values ​​greater than the mask activation threshold to construct a non-zero coordinate hash table; performing a sparse convolution operation of the first kernel size dimension on pixel regions within the non-zero coordinate hash table to extract the stamp texture features as the non-text feature stream; and performing a dense convolution operation of the second kernel size dimension on pixel regions not within the non-zero coordinate hash table to extract character stroke features as the text feature stream.

4. The structured text reconstruction method based on multimodal feature fusion according to claim 1, characterized in that, Step 3 extracts the two-dimensional topological anchor points of each text unit in the text feature stream, calculates the spatial distance between each two-dimensional topological anchor point, and maps it to a spatial bias matrix. This includes: dividing the text feature stream into independent text units and generating bounding rectangles; extracting the absolute physical coordinates of the intersection points of the diagonals of the bounding rectangles in the text feature stream as two-dimensional topological anchor points; calculating the absolute difference of the first coordinate in the horizontal direction and the absolute difference of the second coordinate in the vertical direction between any two two-dimensional topological anchor points, adding the first and second absolute coordinate differences to obtain the Manhattan distance between text units; and inputting the Manhattan distance into a preset negative penalty mapping function, wherein the negative penalty mapping function satisfies the following relation. ,in The Manhattan distance, It is a natural constant. This is the preset exponential smoothing factor. The preset negative basic control coefficient is used; negative bias values ​​with negative values ​​and whose absolute values ​​increase with the increase of physical spatial distance are obtained and constructed into a spatial bias matrix.

5. The structured text reconstruction method based on multimodal feature fusion according to claim 1, characterized in that, Step 3 involves injecting the spatial bias matrix into the multimodal attention network and adjusting the fusion weights of the multimodal attention network based on the degradation matrix. This includes: in the self-attention computation layer of the multimodal attention network, calculating the product of the query vector matrix and the key vector matrix to obtain the initial score matrix, specifically satisfying the formula: ,in The initial score matrix is... The query vector matrix, The key vector matrix, The process involves matrix transpose; matrix addition is performed between the spatial bias matrix and the initial score matrix; global average pooling is then performed on the degraded matrix after binary numerical mapping to obtain the global degraded coefficient value; the text modality preservation coefficient value is obtained by subtracting the global degraded coefficient value from a constant; and point-by-point multiplication is performed between the text modality preservation coefficient value and the tensor elements of the text feature stream to complete the adjustment of the fusion weights.

6. The structured text reconstruction method based on multimodal feature fusion according to claim 1, characterized in that, Step 4 involves decoding based on the fused feature tensor to output a structured semantic sequence. This includes: inputting the fused feature tensor into a converter decoding network to output a basic text character sequence and paragraph layout tags; calculating the cosine similarity of each sentence vector in the basic text character sequence using a locality-sensitive hashing (LSH) retrieval algorithm; establishing directed connections between target sentence pairs with cosine similarity values ​​greater than a reference threshold to construct a heterogeneous document attribute graph from the basic text character sequence; performing graph convolution feature aggregation on the heterogeneous document attribute graph; and serializing and recombining the aggregated node features according to the hierarchical relationship of the paragraph layout tags to output a structured semantic sequence.

7. A structured text reconstruction system based on multimodal feature fusion, characterized in that, Running a structured text reconstruction method based on multimodal feature fusion as described in any one of claims 1 to 6 includes: The degradation awareness module is used to acquire the document image to be processed and calculate the local noise variance to generate an initial noise matrix; the initial noise matrix is ​​smoothed using a spatial smoothing filter with high and low thresholds to output a degradation matrix. The feature decoupling module is used to combine the color distribution prior of the document image to be processed with the degradation matrix to generate a physical color gamut soft isolation mask; and to control asymmetric sparse convolution through the physical color gamut soft isolation mask to decouple the document image to be processed into a text feature stream and a non-text feature stream. The topology constraint module is used to extract the two-dimensional topological anchor points of each text unit in the text feature stream, calculate the spatial distance between each two-dimensional topological anchor point, and map it into a spatial bias matrix. A multimodal attention fusion module is pre-installed with a multimodal attention network based on the Transformer architecture, which is used to inject the spatial bias matrix into the multimodal attention network and adjust the fusion weights of the multimodal attention network based on the degradation matrix. The conditional compensation module is used to determine whether the noise in a local area shown by the degradation matrix is ​​higher than a preset third confidence threshold. If so, the weight of the text feature stream in the multimodal attention network is reduced and feature compensation is performed on the local area. If the noise in a local area shown by the degradation matrix is ​​not higher than the third confidence threshold, the current fusion weights are maintained. The feature fusion and decoding module is used to fuse the non-text feature stream, the weighted text feature stream, and the compensation features output by the conditional compensation module to generate a fused feature tensor; and to decode the fused feature tensor to output a structured semantic sequence.

Citation Information

Patent Citations

  • Text image restoration model and method based on structural attention and text perception

    CN116258652A

  • Text-guided multi-modal relationship extraction method and apparatus

    WO2025130069A1