Document layout analysis method based on mixing method
Through a hybrid method combining convolutional neural network, Transformer and graph neural network, the visual and text features of documents are extracted and fused, and the difficulty of the existing technology in detecting small-scale text areas and distinguishing different types of text areas in document layout analysis is solved, achieving more efficient and accurate document layout analysis.
Patent Information
- Application Number
- CN202510067725.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
Existing document layout analysis techniques have difficulties in detecting small-scale text areas and distinguishing different types of text areas. The Transformer-based model is computationally costly when processing long documents, and the GNN-based model has insufficient node feature definition.
Using a hybrid method, image visual features are extracted through convolutional neural networks, text features are extracted using Transformer, and connections between nodes are learned through graph neural networks. A dynamic range histogram attention mechanism based on channel shuffling and a Transformer model with multi-layer adapters are introduced to perform feature fusion and interaction.
It improves the accuracy and efficiency of document layout analysis, can better detect small-scale text areas and distinguish different types of text areas, reduces calculation costs, and improves the distinction and connection understanding of node characteristics.
Smart Images

Figure CN119992581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of document processing technology, and specifically to a document layout analysis method that uses a hybrid method to extract and combine document features. Background Art
[0002] The document layout analysis task refers to building the structure of each document page by understanding the images, texts, tables, and positional relationships in the document layout. Excellent document layout analysis helps improve the quality and consistency of downstream tasks, improve user experience and acceptance, and support professional translation. The document layout analysis task is essentially an object detection problem. Existing models with different architectures include Transformer-based encoders, graph deep network (GNN)-based models, and object detectors using convolutional neural networks (CNNs). Although these three architecture models have been successfully applied to the document layout analysis task, they each have their limitations.
[0003] CNN-based object detection detects the layout position of each object in a document by returning bounding boxes and categories. However, for all object detection or segmentation models, the predicted bounding boxes may overlap each other due to the blurred boundaries between instances. Since the training loss is not very sensitive to slightly offset predicted boxes, this limits the contribution of the model to reducing bounding box overlap during optimization. Therefore, it is difficult to achieve a high level of IoU (intersection over union, which represents the ratio of the intersection and union between the predicted bounding box and the true bounding box). In particular, it becomes difficult to accurately assign labels to text boxes that are located at the edge of the predicted region or shared by multiple predicted regions. This leads to unsatisfactory performance when IoU ≥ 0.9. In addition, this method cannot detect small-scale text regions that only span one or two text lines (such as titles, steps, and section titles) with high accuracy. Secondly, when two different types of text regions have similar visual textures, such as paragraphs and list items, paragraphs and section titles, and section titles and titles, these methods cannot robustly distinguish them.
[0004] Due to the limitation of input sequence length, the Transformer encoder series methods cannot process very long documents. The high computational cost introduced by self-attention and fully connected graphs also limits its applicability in industry.
[0005] GNN-based models have special advantages in modeling the spatial layout patterns of documents. Each text box in a document can be regarded as a spatially separated node in the graph; text boxes belonging to the same layout area can be regarded as connected by edges, so there are no ambiguous text boxes that are difficult to assign to a group. However, the current focus of GNN-based document layout analysis is on graph sampling and edge definition, and there is less work on the definition of node features. Therefore, how to use the image, text, layout and other features of the document itself to improve the distinction between nodes and enhance the connection between nodes is a key issue that needs to be solved to improve the accuracy of document layout analysis. Summary of the invention
[0006] The present invention aims to provide a document layout analysis method based on a hybrid method. To define node features, the present invention uses a convolutional neural network to extract visual features of an image, uses a Transformer to extract text features, and learns the connections between nodes through a graph neural network. To enhance the correlation between text pixels, the present invention proposes a dynamic range histogram attention mechanism based on channel shuffling. At the same time, in order to more efficiently extract text features using a pre-trained Transformer model, the present invention designs a Transformer model based on a multi-layer adapter. In the feature fusion stage, the present invention deeply fuses the extracted visual features and text features through a cross-attention mechanism, and uses a graph neural network to learn the information interaction between context nodes, thereby obtaining the final features of each text block. Subsequently, each text block is classified, and it is determined whether there is a connection between each pair of text blocks, and finally a prediction result is generated. For text blocks with connections, they are regarded as the same layout area, thereby achieving accurate analysis of document layout.
[0007] The technical solution adopted by the present invention is 1. A document layout analysis method based on a hybrid method, characterized in that the classification of each text block and the judgment of the connection between text block pairs include the following steps:
[0008] Step 1: Initialization;
[0009] Initialize, set batch size batch_size, number of samples per training round num_samples_per_epoch, maximum length of input sequence max_seq_length, maximum number of blocks per input sample max_block_num, maximum training round max_epochs, gradient clipping algorithm clip_gradient_algorithm and clipping threshold clip_gradient_value, optimizer type method, learning rate (lr), weight decay coefficient weight_decay, epsilon parameter eps of AdamW optimizer, number of learning rate warmup steps warmup_steps, verification interval val_interval;
[0010] Step 2: Construction of hybrid method document layout analysis algorithm;
[0011] The algorithm extracts features based on convolutional neural networks and Transformer models, uses graph neural networks to transfer information between text blocks, and finally classifies each text block and determines whether there is a connection between text blocks. First, read the batch data batch in the dataset. For the batch data, perform the following operations in sequence: Use convolutional neural networks to extract visual features of the image. Use the pre-trained Transformer model to extract text features, and further optimize the feature extraction effect through multi-layer adapters. The extracted visual features and text features are fused through the cross-attention mechanism to generate a cross-modal joint feature representation. Use graph neural networks to model the contextual information between text blocks, transfer information between nodes, and further enhance feature representation. Finally, classify the final features of each text block, predict its category, and determine whether there is a connection between each pair of text blocks.
[0012] Step 3: Visual feature extraction;
[0013] Step 3.1: Get the bounding box of the visual feature;
[0014] first_token_idxes: the first token index of each text block; B_batch_dim: constructs a batch index to extract the corresponding bounding box from bbox; feature_bbox: extracts the bounding box of each text block, the shape is [batch_size, num_first, 4]; block_num: the number of text blocks. Accurately obtain the bounding box information of each text block. Step 3.2: Extracting document image visual features
[0015] The ConvNext convolutional network model is used to extract multi-scale features of images. It passes the image to multiple stages in sequence, each of which includes a downsampling and several residual blocks, thereby gradually reducing the resolution of the feature map, increasing the number of channels, and finally outputting a multi-level feature map.
[0016] The number of channels of the input image, the default is 3. The number of residual blocks contained in each stage, depths, the default is [3,3,9,3], representing 4 stages. The number of channels in each stage, the default is [96,192,384,768].
[0017] In the four stages, downsampling layers and convolution blocks are used to process image features, as follows: 4 downsampling layers are defined. The first downsampling layer uses a Conv2d with a stride of 4 for large-stride convolution to quickly reduce the input size by 4 times; then a LayerNorm is used for normalization. The following three downsampling layers each include a LayerNorm and a Conv2d with a stride of 2 and a convolution kernel of 2. Before proceeding to the next stage, the number of channels will also increase. 4 convolution block sequences are defined. For each stage i, several blocks will be constructed, and the number is determined by depths[i]. Each block uses a depthwise separable convolution dwconv with a stride of 3 and a convolution kernel of 7 to extract features in the spatial dimension. The LayerNorm feature map is used for normalization to improve training stability. 1×1 convolution, GELU nonlinear layer and 1×1 convolution are used, among which 1x1 convolution is used for feature transformation in the channel dimension, and the nonlinear layer is introduced to improve the model's expressiveness. Finally, residual connections and random depth are used to prevent overfitting and improve training stability.
[0018] In the i-th stage, the i-th downsampling layer and the i-th convolution block sequence are used. First, the downsampling layer is used to perform channel transformation or resolution scaling on the current feature map. Then, multiple blocks in this stage are entered for deep convolution feature extraction. The input shape is usually an image tensor of [batch_size, in_chans, H, W], which refers to the number of batches, input channels, image feature length and width, respectively. The output is a tuple containing 4 feature maps, each of which corresponds to the output of a stage, and the shapes are: Stage 1: [batch_size, dims[0], H / 4, W / 4]; Stage 2: [batch_size, dims[1], H / 8, W / 8]; Stage 3: [batch_size, dims[2], H / 16, W / 16]; Stage 4: [batch_size, dims[3], H / 32, W / 32]
[0019] Step 3.3: Decode and fuse the visual features of the document image;
[0020] Use ConvNeXt as the backbone network and extract multi-scale features. Normalize the input image using the mean and standard deviation of ImageNet. Adjust the multi-scale features of ConvNeXt [96, 192, 384, 768] to a uniform channel number of 256 through 1×1 convolution. Use spatial attention to simulate human attention by focusing on specific parts of the input image. Spatial attention determines the location of attention in the feature map and then enhances these features. This process enhances the model's ability to recognize and respond to relevant spatial features. Multi-scale convolution blocks are then used to enhance the preservation of contextual relationships, and then the image feature maps of 4 scales are fused. Step 3.3.1: Channel attention block;
[0021] First, the adaptive maximum pooling P max (·) and adaptive average pooling P avg (·) is applied to the spatial dimension to extract the most important features of the entire feature map of each channel, where the adaptive maximum pooling P max (·) and adaptive average pooling P avg The output size of (·) is set to 1×1, which compresses the feature map of each channel into a single value to extract the global spatial information. Then, the correlation between channels is learned by two 1×1 convolutions, that is, the first Conv 1×1 The number of channels is reduced from C to C / rc (C is the number of channels, rc is the channel compression ratio set to 16), and the second Conv 1×1 The number of channels is restored from C / rc to C, and the nonlinear expression is learned through the ReLU activation function. Two features Y are obtained ca1 and Y ca2 . Add and fuse the two features to generate the channel weight W ca The activation function uses Sigmoid to limit the weight value to [0,1]. The generated channel weight W ca Multiply the input feature map X element-by-channel to adjust the feature strength of each channel.
[0022] Step 3.3.2: Spatial attention block;
[0023] First, along the channel direction, collect the maximum value Ch max (·) and the average value Ch avg (·) And the maximum and average features are concatenated in the channel dimension to focus on local features. A large 7×7 convolution kernel can capture more local context information. The concatenated feature is Y max_avg , the number of channels is 2. Then the Sigmoid function is applied to generate the spatial attention weight W spStep 3.3.3: Multi-scale convolutional block;
[0024] Perform depthwise convolutions at multiple scales and use a channel shuffle operation to shuffle channels between groups. More specifically, first use Conv 1×1 (·) expands the number of channels, followed by a batch normalization layer BN(·) and a ReLU6(·) activation layer. Multi-scale deep convolution DWC(·) [1×1, 3×3, 5×5] is used to capture multi-scale and multi-resolution context. And because deep convolution ignores the relationship between channels, a channel shuffle operation is used to merge the relationship between channels. Finally, through Conv 1×1 (·) The encoded multi-scale features F ms Restore and Enter X sp The same number of channels and adding residual connections.
[0025] Step 3.3.4: Multi-scale image feature map fusion;
[0026] The four input feature maps of different scales [X0,X1,X2,X3]∈X ms , through the global average pooling operation P in the spatial dimension avg (·) Obtain global feature representation, and then pass Conv 1×1 (·) Model the correlation between channels and generate channel descriptors through the sigmoid activation function. Concatenate the global features of the four scales in the channel dimension and obtain their respective weights W through the softmax function in the channel dimension. MSF It means that the generated weights are multiplied element-wise with the corresponding inputs and added.
[0027] Step 3.3.5: Dynamic range convolution block;
[0028] The feature map Y MSF The enhanced feature maps are divided into 4 groups by further mixing and sharing information through channel shuffling. The grouped feature maps are transposed to disrupt the channel order in each group, and the feature map Y is obtained. copy ;
[0029] For feature map Y MSF Through dynamic range convolution, the convolution is processed between similar pixels rather than between adjacent pixels. Specifically, the feature map Y is transformed along the channel dimension. MSF Divided into two partsY MSF1 and Y MSF2 For the first part of feature Y MSF1 The features are sorted in the W direction and the H direction so that the feature values increase gradually from the upper left corner to the lower right corner. The first feature branch after sorting It is reconnected with the second feature branch F2 to form a new feature map, and processed by depthwise separable convolution to obtain F out . The high-intensity and low-intensity pixels are organized into regular patterns at the diagonal of the matrix, thereby achieving dynamic range convolution, so that similar pixels are learned by the convolution kernel. Finally, the first part of the features is separated and restored in the H direction and W direction in turn to obtain F out .
[0030] For the feature map F' out The same process is also performed, so that the overall channel features can learn the features between similar pixels through dynamic range convolution. out and F' out Concatenate and pass through a Conv 1×1 (·) Restore the original channel feature dimension.
[0031] The output is divided into Q1, Q2, K1, K2, and V. First, V is reshaped into Then sort so that V is sorted from small to large in the HW dimension and save the index value d. Q1, Q2, K1, K2 are sorted according to index d. Given the number of bins B, the sorted shape is reshaped to and Perform self-attention on Q1, K1, V and Q2, K2, V. Finally, multiply the two self-attention results to get the final output.
[0032] Step 3.4: Obtain visual features of each text block;
[0033] Get the bounding box bbox of each text block according to first_token_idxes. The specific shape of bbox is [batch_size, num_tokens, 4], and each bounding box is represented by [x_min, y_min, x_max, y_max]. For the visual features of the entire document image, use ROIAlign according to the coordinate box position information of each text block, first scale the coordinates of the bounding box to the scale of the feature map, and then extract the document visual features of each text block at the corresponding position [batch_size, block_num, C, 1, 1], and finally reshape the feature shape to [batch_size, block_num, C] to represent the visual features of each text block.
[0034] Step 3.5: Extract document language text features;
[0035] Relying on existing document-related language text pre-training research, a multi-layer adapter fine-tuning method is used to extract language text features with a small training cost. The Bros model architecture pre-trained on the IIT-CDIP dataset through the geometry pre-training task in GeoLayoutLM is selected to extract text features in the document. Freeze the original pre-trained model parameters and add multi-layer adapters in the self-attention layer, the middle layer of the feedforward network, and the output layer.
[0036] Step 3.5.1: Extract document language text features;
[0037] Use a PDF parser or OCR engine to scan the document line by line or block by block to detect text fragments. At the same time, record the position information of each text fragment in the page. Finally, a data list of the form (text, coordinates) is obtained, and the text and coordinate information are converted into the following features: Text embedding Token Embedding, mapping each text segment or token after word segmentation into a vector; 1D Positional Embedding, using standard positional encoding to represent the position of each token in the sequence; Segment Order Embedding 1D Segment Order Embedding, attaching a Segment ID to each token to distinguish sentences; Segment BIE Tag Embedding 1D BIE Embedding; B, I, and E represent the beginning of a paragraph, the middle of a paragraph, and the end of a paragraph respectively; 2D Segment Box Embedding 2DBounding Box Embedding normalizes the coordinate information (x1, y1, x2, y2) to a certain range, and then maps it to a coordinate vector of a certain dimension.
[0038] The text features are learned by fine-tuning the adapter layer added to the frozen pre-trained model to obtain the characteristics of each text.
[0039] The frozen pre-trained model consists of 24 self-attention layers, each of which consists of a self-attention mechanism, an intermediate layer of a feedforward network, and an output layer. The feedforward layer consists of two layers of linear transformation and an activation function. Two layers of adapters are added after the self-attention mechanism, the intermediate layer of the feedforward network, and the output layer of each layer. The adapter module first introduces a downsampling layer, whose core goal is to effectively compress features while retaining key information. Through this design, the model can effectively extract key features in the image while reducing the computational burden, thereby improving operating efficiency. A ReLU activation function layer is added to the adapter module. Through the upsampling layer, the original feature shape is restored to retain sufficient information.
[0040] Freeze all parameters of the self-attention model architecture and only train the adapter parameters. This preserves the original network's parameter settings, avoids the tedious process of training the network from scratch, and does not require saving a new parameter weight for each task, significantly reducing storage requirements.
[0041] Step 4: Feature fusion and interaction;
[0042] Text features and corresponding visual features come from different modalities and need to be fused together. Direct concatenation or simple addition and multiplication may not fully capture cross-modal associations. Using interactive attention allows text and visual features to "pay attention" to each other, thereby extracting more discriminative cross-modal representations. The text vector and visual vector of the same text block are mapped to the same hidden layer dimension respectively.
[0043] In order to model the relationship between text blocks, a dynamic graph convolutional network Dynamic GCN is introduced. Each text block is regarded as a node in the graph, and the node feature is initially composed of the fused feature H in the previous step. i Representation. Since the positions of text blocks on the page are different, the neighbor relationship varies with the layout. Let each text block interact with its k neighbors and update the features of each text block to obtain the final features of each text block. This can improve each text block's understanding of the relationship between the context text blocks.
[0044] Step 5: Determine the classification of text blocks and the relationship between text block pairs;
[0045] Step 5.1: For text block classification:
[0046] In scenarios such as document understanding or form extraction, each text block often needs to be assigned a semantic label. i Input to a linear layer for classification:
[0047] Step 5.2: For the text block relation prediction task:
[0048] Predict whether there is a connection between two text blocks, that is, determine whether the text blocks are related in semantics or layout structure. In the case of low accuracy of relationship prediction tasks, rule-based methods can be used to strike a balance between time, space overhead and accuracy. Using a fully connected method to establish predefined edges between text blocks can capture potential relationships more comprehensively.
[0049] Specifically, by calculating the distance between pairs of text blocks, these edge features are combined with node features to generate edge features between each pair of text blocks. Then, a linear layer is used to determine whether there is an edge between pairs of text blocks. This process ensures that the relationship between text blocks can be accurately captured, thereby improving the overall effect of document layout analysis.
[0050] Step 6: Iteration termination condition;
[0051] During the training process, the model's loss function consists of two parts: annotation loss and link loss. The annotation loss is calculated using the cross entropy function and is used to measure the difference between the model's prediction results on the annotation task and the true label. The link loss is used to evaluate the relationship prediction effect between text blocks and is further divided into three parts: mask link loss, positive sample link loss, and variance loss. Finally, the annotation loss and link loss are weighted and accumulated to form a total loss, which is used to guide the optimization of the model.
[0052] During the training iteration process, two termination conditions are set to ensure the efficiency and stability of the training. The maximum number of training rounds is set to 400, and the training stops when the upper limit is reached. After each round of training, the performance of the model on the validation set is verified, and the current optimal model is saved to ensure the best effect of the final model. This training process ensures the convergence of the model and the training efficiency through reasonable design and termination conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Document layout analysis framework based on hybrid approach
[0054] Figure 2 Dynamic Range Histogram Attention Based on Channel Shuffling
[0055] Figure 3 Transformer model based on multi-layer adapter
[0056] Figure 4 Experimental results on entity recognition (named entity recognition) and entity linking (relation extraction). DETAILED DESCRIPTION
[0057] To achieve the above object, the present invention adopts the following technical solutions:
[0058] Step 1: Initialization
[0059] The parameters required for the initialization algorithm are as follows:
[0060] Batch_size: The batch size used in each training determines the amount of data input into the model each time and is set to 2.
[0061] num_samples_per_epoch: The number of samples used in each training round, set to 149.
[0062] max_seq_length: The maximum length of the input sequence. Sequences exceeding this length will be truncated. It is set to 512.
[0063] max_block_num: The maximum number of blocks included in each input sample, set to 150.
[0064] max_epochs: The maximum number of training rounds. The training process is performed for a maximum of 400 rounds.
[0065] clip_gradient_algorithm: The type of gradient clipping algorithm. Set it to norm to indicate clipping based on the gradient norm.
[0066] clip_gradient_value: The threshold for gradient clipping, set to 1.0, and gradients exceeding this value will be clipped.
[0067] method: The type of optimizer. If it is set to adamw, it means using the AdamW optimizer.
[0068] lr: learning rate, set to 2e-5.
[0069] weight_decay: Weight decay coefficient, used to prevent overfitting, set to 0.01.
[0070] eps: epsilon parameter in the AdamW optimizer, used for numerical stability, set to 1e-8.
[0071] warmup_steps: The number of steps for learning rate warmup, set to 200.
[0072] val_interval: The number of rounds of training after which verification is performed. If set to 1, verification is performed once after each round of training.
[0073] Step 2: Hybrid method document layout analysis algorithm construction
[0074] The algorithm extracts features based on convolutional neural networks and Transformer models, uses graph neural networks to transfer information between text blocks, and finally classifies each text block and determines whether there is a connection between text blocks. First, read the batch data in the dataset. For the batch data, perform the following operations in sequence: Use convolutional neural networks to extract visual features of the image. Use the pre-trained Transformer model to extract text features, and further optimize the feature extraction effect through multi-layer adapters. The extracted visual features and text features are fused through a cross-attention mechanism to generate a cross-modal joint feature representation. Use graph neural networks to model the contextual information between text blocks, transfer information between nodes, and further enhance feature representation. Finally, classify the final features of each text block, predict its category, and determine whether there is a connection between each pair of text blocks. The details are as follows:
[0075] Step 3: Visual feature extraction
[0076] Step 3.1: Get the bounding box of the visual feature
[0077] first_token_idxes: The first token index of each text block.
[0078] B_batch_dim: constructs a batch index for extracting the corresponding bounding box from bbox.
[0079] feature_bbox: Extract the bounding box of each text block, shape is [batch_size, num_first, 4].
[0080] block_num: The number of text blocks.
[0081] These can accurately obtain the bounding box information of each text block, providing a basis for subsequent visual feature extraction.
[0082] Step 3.2: Extracting document image visual features
[0083] The ConvNext convolutional network model is used to extract multi-scale features of images. It passes the image to multiple stages in sequence, each of which includes a downsampling and several residual blocks, thereby gradually reducing the resolution of the feature map, increasing the number of channels, and finally outputting a multi-level feature map.
[0084] The number of channels of the input image, the default is 3. The number of Blocks (residual blocks) contained in each stage, depths, the default is [3,3,9,3], representing 4 stages. The number of channels in each stage, the default is [96,192,384,768].
[0085] In the four stages, downsampling layers and convolution blocks are used to process image features, as follows: 4 downsampling layers are defined. The first downsampling layer uses a Conv2d with a stride of 4 for large-stride convolution to quickly reduce the input size by 4 times; then a LayerNorm is used for normalization. The following three downsampling layers each include a LayerNorm and a Conv2d with a stride of 2 and a convolution kernel of 2. Before proceeding to the next stage, the number of channels will also increase. 4 convolution block sequences are defined. For each stage i, several blocks are constructed (the number is determined by depths[i]). Each block uses a dwconv (depthwise separable convolution) with a stride of 3 and a convolution kernel of 7 to extract features in the spatial dimension. The LayerNorm feature map is used for normalization to improve training stability. 1×1 convolution, GELU nonlinear layer and 1×1 convolution are used, among which 1x1 convolution is used for feature transformation in the channel dimension, and the nonlinear layer is introduced to improve the model's expressiveness. Finally, residual connections and random depth are used to prevent overfitting and improve training stability.
[0086] In the i-th stage, the i-th downsampling layer and the i-th convolution block sequence are used. First, the downsampling layer is used to perform channel transformation or resolution scaling on the current feature map. Then, multiple blocks in this stage are entered for deep convolution feature extraction. The input shape is usually an image tensor of [batch_size, in_chans, H, W], which refers to the number of batches, input channels, image feature length and width, respectively. The output is a tuple containing 4 feature maps, each of which corresponds to the output of a stage, and the shapes are: Stage 1: [batch_size, dims[0], H / 4, W / 4]; Stage 2: [batch_size, dims[1], H / 8, W / 8]; Stage 3: [batch_size, dims[2], H / 16, W / 16]; Stage 4: [batch_size, dims[3], H / 32, W / 32]
[0087] Step 3.3: Decoding and fusing document image visual features Using ConvNeXt as the backbone network, extract multi-scale features. Normalize the input image using the mean and standard deviation of ImageNet. Adjust the multi-scale features of ConvNeXt [96, 192, 384, 768] to a uniform number of channels (256) through 1×1 convolution. Use channel attention blocks to assign different importance to each channel, thereby emphasizing more relevant features while suppressing less useful features. Use spatial attention to simulate human attention by focusing on specific parts of the input image. Spatial attention determines the location of attention in the feature map and then enhances these features. This process enhances the model's ability to recognize and respond to relevant spatial features. Multi-scale convolution blocks are then used to enhance the preservation of context. The image feature maps of the four scales are then fused.
[0088] Step 3.3.1: Channel Attention Block
[0089] We first use the adaptive maximum pooling P max (·) and adaptive average pooling P avg (·) is applied to the spatial dimension to extract the most important features of the entire feature map of each channel, where the adaptive maximum pooling P max (·) and adaptive average pooling P avg The output size of (·) is set to 1×1, which compresses the feature map of each channel into a single value to extract the global spatial information. Then, the correlation between channels is learned by two 1×1 convolutions, that is, the first Conv 1×1 The number of channels is reduced from C to C / rc (C is the number of channels, rc is the channel compression ratio set to 16), and the second Conv 1×1 The number of channels is restored from C / rc to C, and the nonlinear expression is learned through the ReLU activation function. In this way, two features Y are obtained ca1 and Y ca2 .
[0090] Y ca1 =Conv 1×1 (ReLU(Conv 1×1 (P max (X))) (1)
[0091] Y ca2 =Conv 1×1 (ReLU(Conv 1×1 (R avg (X))) (2)
[0092] Then the two features Y ca1 and Y ca2 Add and fuse to generate channel weight W caThe activation function uses Sigmoid to limit the weight value to [0,1]. The generated channel weight W ca Multiply the input feature map X element-by-channel to adjust the feature strength of each channel.
[0093] W ca =Sigmoid(Y ca1 +Y ca2 ) (3)
[0094] X ca =W ca ·X#(4)
[0095] Step 3.3.2: Spatial Attention Block
[0096] First, along the channel direction, collect the maximum value Ch max (·) and the average value Ch avg (·) And concatenate the maximum and average features in the channel dimension to focus on local features:
[0097] Y max_avg =Concat[Ch max (X ca ),Ch avg (X ca )] (5)
[0098] A large 7×7 convolution kernel can capture more local context information. The concatenated feature Y max_avg The number of channels is 2. Then the Sigmoid function is applied to generate the spatial attention weight W sp .
[0099] W sp =Sigmoid(Conv 7×7 (Y max_avg )) (6)
[0100] X sp =W sp ·X ca (7)
[0101] Step 3.3.3: Multi-scale convolutional blocks
[0102] Perform depthwise convolutions at multiple scales and use a channel shuffle operation to shuffle channels between groups. More specifically, first use Conv 1×1 (·) expands the number of channels, followed by a batch normalization layer BN(·) and a ReLU6(·) activation layer:
[0103] F u =ReLU6(BN(Conv 1×1 (Xsp )) (8)
[0104] Multi-scale deep convolution DWC(·)[1×1, 3×3, 5×5] is used to capture multi-scale and multi-resolution context. In addition, the deep convolution ignores the relationship between channels and uses the channel shuffle operation to merge the relationship between channels:
[0105] F u1×1 =ReLU6(BN(DWC 1×1 (F u )) (9)
[0106] F u3×3 =ReLU6(BN(DWC 3×3 (F u )) (10)
[0107] F u5×5 =ReLU6(BN(DWC 5×5 (F u )) (11)
[0108] F ms =ChannelShuffle(F u1×1 +F u3×3 +F u5×5 ) (12)
[0109] Finally, through Conv 1×1 (·) The encoded multi-scale features F ms Restore and Enter X sp Same number of channels, and adding residual connections:
[0110] X ms =BN(Conv 1×1 (F ms )+X sp ) (13)
[0111] Step 3.3.4: Multi-scale image feature map fusion
[0112] The four input feature maps of different scales [X0,X1,X2,X3]∈X ms , through the global average pooling operation P in the spatial dimension avg (·) Obtain global feature representation, and then pass Conv 1×1 (·) Model the correlation between channels and generate channel descriptors through the sigmoid activation function:
[0113]
[0114] The global features of the four scales are concatenated in the channel dimension, and their respective weights W are obtained through the softmax function in the channel dimension. MSF express:
[0115]
[0116] Multiply the generated weights by the corresponding input element-wise and add them together:
[0117] Y MSF =W MSF [:,1,:]·X1+Y MSF [:,2,:]·X2+W MSF [:,3,:]·X3+Y MSF [:,4,:]·X4 (16)
[0118] Step 3.3.5: Dynamic Range Convolution Block
[0119] The feature map Y MSF The enhanced feature maps are divided into 4 groups by further mixing and sharing information through channel shuffling. The grouped feature maps are transposed to disrupt the channel order in each group, and the feature map Y is obtained. copy ;
[0120] For feature map Y MSF Through dynamic range convolution, the convolution is processed between similar pixels rather than between adjacent pixels. Specifically, the feature map Y is transformed along the channel dimension. MSF Divided into two partsY MSF1 and Y MSF2 For the first part of feature Y MSF1 The features are sorted in the W direction and the H direction in turn so that the feature values increase gradually from the upper left corner to the lower right corner.
[0121] F sort =SortH(SortW(Y MSF )) (17)
[0122] The first feature branch after sorting It is reconnected with the second feature branch F2 to form a new feature map, and processed by depthwise separable convolution to obtain F out .
[0123] F out =DWConv 3×3 (Conv 1×1 (Convat[F sort ,F2])) (18)
[0124] In this way, high-intensity and low-intensity pixels are organized into regular patterns at the diagonal of the matrix, thereby achieving dynamic range convolution, so that similar pixels can be learned by the convolution kernel. Finally, the first part of the features is separated and restored in the H direction and W direction in turn to obtain F out .
[0125] For the feature map F' out The same process is also performed, so that the overall channel features can learn the features between similar pixels through dynamic range convolution. out and F' out Concatenate and pass through a Conv 1×1 (·) Restore the original channel feature dimension.
[0126] Y DRC =Conv 1×1 (Concat[F out ,F' out ]) (19)
[0127] The output is divided into Q1, Q2, K1, K2, and V. First, V is reshaped into Then sort so that V is sorted from small to large in the HW dimension and save the index value d. Q1, Q2, K1, K2 are sorted according to index d. Given the number of bins B, the sorted shape is reshaped to and Implement self-attention on Q1, K1, V and Q2, K2, V.
[0128]
[0129] Finally, the two self-attention results are multiplied to get the final output.
[0130] Y DRC =Y DRC1 ·Y DRC2 (twenty two)
[0131] Step 3.4: Get the visual features of each text block
[0132] Get the bounding box bbox of each text block according to first_token_idxes. The specific shape of bbox is [batch_size, num_tokens, 4], and each bounding box is represented by [x_min, y_min, x_max, y_max]. For the visual features of the entire document image, use ROIAlign to scale the coordinates of the bounding box to the scale of the feature map according to the coordinate box position information of each text block, and then extract the document visual features of each text block at the corresponding position [batch_size, block_num, C, 1, 1]. Finally, reshape the feature shape to [batch_size, block_num, C] to represent the visual features of each text block.
[0133] Step 3.5: Extract document language text features
[0134] Rely on existing document-related language text pre-training research. In order to make better use of the pre-trained model, a multi-layer adapter fine-tuning method is used to extract language text features with a smaller training cost. The Bros model architecture pre-trained on the IIT-CDIP dataset through the geometry pre-training task in GeoLayoutLM is selected to extract text features in the document. Choose to freeze the original pre-trained model parameters and add multi-layer adapters in the self-attention layer, the middle layer of the feedforward network, and the output layer.
[0135] Step 3.5.1: Extract document language text features
[0136] Use a PDF parser or OCR engine to scan the document line by line or block by block to detect text fragments. At the same time, record the position information of each text fragment in the page (such as x-coordinate, y-coordinate, width, height). Finally, a data list of the form (text, coordinate) is obtained, for example: [("Invoice", (x1, y1, x2, y2)), ("Total", (x3, y3, x4, y4)), ...], and convert these text and coordinate information into the following features: Token Embedding, which maps each text segment or token after word segmentation into a vector; 1D Positional Embedding, which uses standard positional encoding to represent the position of each token in the sequence; 1DSegment Order Embedding, which attaches a Segment ID (0 or 1) to each token to distinguish sentences; 1D BIE Embedding, where B, I, and E represent Beginning, Inside, and End, respectively; 2D Bounding Box Embedding (two-dimensional segment frame embedding) normalizes the coordinate information (x1, y1, x2, y2) to a certain range (such as 0 to 1000), and then maps it to a coordinate vector of a certain dimension.
[0137] The text features are learned by fine-tuning the adapter layer added to the frozen pre-trained model to obtain the characteristics of each text.
[0138] The frozen pre-trained model consists of 24 self-attention layers, each of which consists of a self-attention mechanism, an intermediate layer of a feedforward network, and an output layer. The feedforward layer consists of two layers of linear transformation and an activation function.
[0139]
[0140] Intermediate = ReLU (Attention out W I +b I ) (twenty four)
[0141] FFN out =Intermediate·W FFN +b FFN (25)
[0142] Two layers of adapters are added after the self-attention mechanism of each layer, the intermediate layer of the feedforward network, and the output layer. The adapter module first introduces a downsampling layer, whose core goal is to effectively compress features while retaining key information. Through this design, the model can effectively extract key features in the image while reducing the computational burden, thereby improving operating efficiency. In addition, in order to enhance the network's ability to capture nonlinear features, a ReLU activation function layer is added to the adapter module. Finally, the upsampling layer is used to restore the original feature shape and retain enough information.
[0143] f(x)=W up (ReLU(W down ·x+b down ))+b up (26)
[0144] Freeze all parameters of the self-attention model architecture and only train the adapter parameters. This preserves the original network's parameter settings, avoids the tedious process of training the network from scratch, and does not require saving a new parameter weight for each task, significantly reducing storage requirements.
[0145] Step 4: Feature fusion and interaction
[0146] Text features and corresponding visual features come from different modalities and need to be fused together. Direct concatenation or simple addition and multiplication may not fully capture cross-modal associations. Using interactive attention allows text and visual features to "pay attention" to each other, thereby extracting more discriminative cross-modal representations. The text vector and visual vector of the same text block are mapped to the same hidden layer dimension respectively.
[0147] In order to model the relationship between text blocks, we introduce a dynamic graph convolutional network (Dynamic GCN). Each text block is regarded as a node in the graph, and the node feature is initially composed of the fused feature H i Representation. Since the positions of text blocks on the page are different, the neighbor relationship varies with the layout. Let each text block interact with its k neighbors and update the features of each text block to obtain the final features of each text block. This can improve each text block's understanding of the relationship between the context text blocks.
[0148] Step 5: Determine the classification of text blocks and the relationship between text block pairs
[0149] Step 5.1: For text block classification:
[0150] In scenarios such as document understanding or form extraction, each text block often needs to be assigned a semantic label. i Input to a linear layer for classification:
[0151] Y i =softmax(W cls ·H i +b cls )#(27)
[0152] Step 5.2: For the text block relation prediction task:
[0153] Predict whether there is a connection between two text blocks, that is, determine whether the text blocks are related in semantics or layout structure. When the accuracy of the relationship prediction task is low, a rule-based approach can be used to strike a balance between time, space overhead and accuracy. However, as the accuracy of the relationship prediction task increases, the rule-based approach may cause some edges to be mistakenly abandoned in the initial stage, thus affecting the overall performance. To this end, a fully connected approach is used to establish predefined edges between text blocks, which can capture potential relationships more comprehensively.
[0154] Specifically, by calculating the distance between pairs of text blocks, these edge features are combined with node features to generate edge features between each pair of text blocks. Then, a linear layer is used to determine whether there is an edge between pairs of text blocks. This process ensures that the relationship between text blocks can be accurately captured, thereby improving the overall effect of document layout analysis.
[0155] Step 6: Iteration termination condition
[0156] During the training process, the loss function of the model consists of two parts: labeling loss and linking loss. Labeling loss is calculated by the cross entropy function and is used to measure the difference between the model's prediction results on the labeling task and the true label. Linking loss is used to evaluate the relationship prediction effect between text blocks, which is further divided into three parts: masked link loss, positive link loss, and variance loss. Masked link loss ignores the impact of invalid links by applying a mask to the pairwise link loss and normalizes the loss of valid links; positive link loss focuses on the link prediction of positive samples, filters positive samples through masks and calculates their losses; variance loss constrains the consistency of positive sample prediction probabilities and reduces the uncertainty of predictions by calculating the mean and variance of positive sample probabilities. Finally, labeling loss and link loss are weighted and accumulated to form a total loss to guide model optimization.
[0157] During the training iteration process, two termination conditions are set to ensure the efficiency and stability of the training. The maximum number of training rounds is set to 400, and the training stops when the upper limit is reached. After each round of training, the performance of the model on the validation set is verified, and the current optimal model is saved to ensure the best effect of the final model. This training process ensures the convergence of the model and the training efficiency through reasonable loss design and termination conditions.
[0158] Experiment and analysis
[0159] 1) Experimental conditions
[0160] The hardware environment of the present invention is 13th Gen Intel(R) Core(TM) i9-13900K, 128GB memory, and GPU is RTX4090; the software platform is Ubuntu 22.04, and the programming language is Python.
[0161] 2) Experimental data
[0162] The experiments of this invention mainly use the FUNSD dataset.
[0163] The FUNSD (Form Understanding in Noisy Scanned Documents) dataset is a public dataset focused on document understanding. It contains complex and diverse real-world form documents and aims to evaluate the form understanding capabilities of models in noisy environments. The dataset consists of 199 scanned form images, of which 149 are used for training and 50 for testing. The document types cover a variety of handwritten and printed forms, including application forms, questionnaires, and receipts. The annotation information includes entity-level annotations, where each text block is annotated as a specific category, such as "question", "answer", "title", and "other"; and relation-level annotations, which define the association between entities, such as the link between questions and answers. As it is a scanned document, the FUNSD dataset has problems such as noise, distortion, and uneven lighting. The text is also arranged irregularly, including complex tables, nested structures, and nonlinear layouts. In addition, some documents contain handwritten text, which increases the difficulty of recognition. These characteristics make the FUNSD dataset a valuable resource for studying tasks such as document visual understanding, information extraction, and relational reasoning.
[0164] 3) Performance comparison
[0165] The present invention is compared with the existing document layout analysis method. The experimental results are shown in the attached figure. SER is named entity recognition (entity recognition), and RE is entity linking (relation extraction). The method proposed by the present invention improves the F1 value of named entity recognition and entity linking.
[0166] In summary, the document layout analysis method proposed in the present invention comprehensively considers the use of convolutional neural networks to extract local features, and uses dynamic range multi-scale attention based on channel shuffling to strengthen the connection between similar pixel features, so that when using graph neural networks to interact with text blocks, the information between each other can be better distinguished and the connection between each other can be found. The use of Transformer based on multiple adapters can learn the knowledge of pre-training and new data set training with fewer parameters. Experimental results show that the present invention provides a powerful solution for document layout analysis and understanding.
[0167] The present invention proposes a document layout method based on a hybrid method, which can infer the category of each text box and the relationship between each pair of text boxes from an input document.
[0168] The document layout analysis based on the hybrid method proposed by the present invention is described in detail below in conjunction with a specific implementation. For ease of explanation, the algorithm simulates a document, specifically including a document image Image and {{text: "text1", box: [x11, y11, x12, y12]}, {text: "text2", box: [x21, y21, x22, y22]}, {text: "text3", box: [x31, y31, x32, y32]}, {text: "text4", box: [x41, y41, x42, y42]}} extracted by the PDF parsing engine / OCR engine.
[0169] Step 1: Initialization
[0170] Initialize the parameters required by the algorithm, such as max_seq_length is 512, max_block_num is 150, etc.
[0171] Step 2: Hybrid method document layout analysis algorithm construction
[0172] Step 2.1: Using ConvNext as the visual backbone, pass the document image Image through ConvNext-tiny to obtain four scale feature maps.
[0173] Step 2.2: The feature maps of each scale are adjusted to the same number of channels. The channel attention block is used to strengthen the important features in the channel and suppress the minor features. The spatial position information of the text block is learned through the spatial attention block. The multi-scale convolution block is used to learn the information of text blocks of different scales.
[0174] Step 2.3: Fuse the multi-scale image features of four different scales to obtain an image feature representation.
[0175] Step 2.3: The obtained image feature representation is passed through a dynamic range convolution block to learn the relationship between similar pixels rather than adjacent pixels, so as to enhance the long-range dependency of the feature map and obtain the final document image feature representation.
[0176] Step 2.4: Use ROIAlign to extract the visual features corresponding to text1 according to the bounding box coordinates box:[x11,y11,x12,y12]. The visual features of text2-4 are also obtained accordingly.
[0177] Step 2.5: Based on the text and the corresponding coordinates, such as text1 and box:[x11,y11,x12,y12], convert these text and coordinate information into text embedding, one-dimensional position embedding, segment order embedding, segment BIE label embedding and two-dimensional segment box embedding.
[0178] Step 2.6: Pass these embeddings through a frozen transformer backbone (the adapter is not frozen) to get the embeddings for each word.
[0179] Step 3: Calculate the correlation between the two modal features through the interactive attention module and generate a fused cross-modal feature representation. The fused features are exchanged with the surrounding 20 neighbors through a dynamic graph convolutional network and the features of each text block are updated.
[0180] Step 4: Determine the classification of text blocks and the relationship between text block pairs
[0181] Step 4.1: Classify each text block feature, whether text1 belongs to title, question, answer or other.
[0182] Step 4.2: For each pair of text blocks, concatenate the features of the two text blocks and the distance between the two text blocks, and use this feature to determine whether there is a relationship between the two text blocks. For example, if text1 and text2 are questions and answers, there is a relationship connection between them, and if text3 is a title, there is no relationship connection between it and other texts.
[0183] See also Figure 1, the image of the upper left document is used as the input of the visual module, and ConvNext-tiny is used as the visual backbone. The input image is passed through the visual backbone to obtain four feature maps. The four feature map channels are unified, and the channel importance, spatial importance and multi-scale information are learned through the channel attention, spatial attention and multi-scale feature extraction modules. Then the four feature maps are fused and the similar pixel information is learned by dynamic range convolution. Finally, the features of each text block are extracted through ROIAlign to obtain N1-N3 in the figure. The text information and coordinate information extracted from the upper left document are converted into text embedding, one-dimensional position embedding, segment order embedding, segment BIE label embedding and two-dimensional segment box embedding. The embedding Ft1-Ft5 of each word is obtained through a language model. Then, the visual and text features are fused through an attention mechanism, and the neighbor node information is learned by the graph neural network. Finally, the downstream task layer makes predictions.
[0184] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation schemes of the present invention. They are not used to limit the scope of protection of the present invention. All equivalent implementation schemes or changes that do not deviate from the spirit of the invention should be included in the scope of protection of the present invention.
Claims
1. A document layout analysis method based on a hybrid approach, characterized in that: The classification of each text block and the determination of the connection between pairs of text blocks include the following steps: Step 1: Initialization; Initialize, set batch size batch_size, number of samples per training round num_samples_per_epoch, maximum length of input sequence max_seq_length, maximum number of blocks per input sample max_block_num, maximum training round max_epochs, gradient clipping algorithm clip_gradient_algorithm and clipping threshold clip_gradient_value, optimizer type method, learning rate lr, weight decay coefficient weight_decay, epsilon parameter eps of AdamW optimizer, number of learning rate warmup steps warmup_steps, verification interval val_interval; Step 2: Construction of hybrid method document layout analysis algorithm; Based on the convolutional neural network and Transformer model, features are extracted and the information between text blocks is transmitted using the graph neural network. Finally, each text block is classified and whether there is a connection between the text blocks is determined. Step 3: Visual feature extraction; Step 3.1: Get the bounding box of the visual feature; first_token_idxes: the first token index of each text block; B_batch_dim: constructs a batch index to extract the corresponding bounding box from bbox; feature_bbox: extracts the bounding box of each text block, the shape is [batch_size, num_first, 4]; block_num: the number of text blocks; accurately obtain the bounding box information of each text block; Step 3.2: Extracting document image visual features; The ConvNext convolutional network model is used to extract multi-scale features of images. The image is passed to multiple stages in sequence. Each stage includes a downsampling and several residual blocks, thereby gradually reducing the resolution of the feature map, increasing the number of channels, and finally outputting a multi-level feature map. The number of channels of the input image is 3 by default; the number of residual blocks contained in each stage, depths, is [3,3,9,3] by default, representing 4 stages; the number of channels in each stage is [96,192,384,768] by default; Step 3.3: Decode and fuse the visual features of the document image; Using ConvNeXt as the backbone network, after extracting multi-scale features; normalizing the input image using the mean and standard deviation of ImageNet; adjusting the multi-scale features of ConvNeXt [96, 192, 384, 768] to a uniform channel number of 256 through 1×1 convolution; using spatial attention to simulate human attention by focusing on specific parts of the input image; spatial attention determines the location of attention in the feature map and then enhances these features; this process enhances the model's ability to recognize and respond to relevant spatial features; then using multi-scale convolution blocks to enhance the preservation of contextual relationships, and then fusion of image feature maps of 4 scales; Step 3.4: Obtain visual features of each text block; Get the bounding box bbox of each text block according to first_token_idxes. The specific shape of bbox is [batch_size, num_tokens, 4], and each bounding box is represented by [x_min, y_min, x_max, y_max]. For the visual features of the entire document image, use ROIAlign according to the coordinate box position information of each text block, first scale the coordinates of the bounding box to the scale of the feature map, and then extract the document visual features of each text block at the corresponding position [batch_size, block_num, C, 1, 1]. Finally, reshape the feature shape to [batch_size, block_num, C] to represent the visual features of each text block. Step 3.5: Extract document language text features; Relying on the existing document-related language text pre-training research, a multi-layer adapter fine-tuning method is used to extract language text features with a small training cost; the Bros model architecture pre-trained on the IIT-CDIP dataset through the geometry pre-training task in GeoLayoutLM is selected to extract text features in the document; the original pre-training model parameters are frozen, and multi-layer adapters are added to the self-attention layer, the middle layer of the feedforward network, and the output layer; Step 4: Feature fusion and interaction; Text features and corresponding visual features come from different modalities and need to be fused together. Direct concatenation or simple addition and multiplication may not be able to fully capture cross-modal associations. Using interactive attention can make text and visual features "pay attention" to each other, thereby extracting more discriminative cross-modal representations. The text vector and visual vector of the same text block are mapped to the same hidden layer dimension respectively. In order to model the relationship between text blocks, a dynamic graph convolutional network DynamicGCN is introduced; each text block is regarded as a node in the graph, and the node feature is initially composed of the fused feature H i Representation; Due to the different positions of text blocks on the page, the neighbor relationship varies with the layout; Let each text block interact with the surrounding k neighbors, and update the features of each text block to obtain the final features of each text block; This can improve the understanding of the relationship between each text block and the context text block; Step 5: Determine the classification of text blocks and the relationship between text block pairs; Step 5.1: For text block classification: In scenarios such as document understanding or form extraction, each text block often needs to be assigned a semantic label. i Input to a linear layer for classification: Step 5.2: For the text block relation prediction task: Predict whether there is a connection between two text blocks, that is, determine whether the text blocks are related in semantics or layout structure. When the accuracy of the relationship prediction task is low, a rule-based method can be used to strike a balance between time, space overhead and accuracy. The full connection method is used to establish predefined edges between text blocks, which can capture potential relationships more comprehensively. Specifically, by calculating edge features such as the distance between pairs of text blocks, these edge features are combined with node features to generate edge features between each pair of text blocks; then, a linear layer is used to determine whether there is an edge between pairs of text blocks; this process ensures that the relationship between text blocks can be accurately captured, thereby improving the overall effect of document layout analysis; Step 6: Iteration termination condition; During the training process, the model's loss function consists of two parts: annotation loss and link loss. The annotation loss is calculated through the cross entropy function to measure the difference between the model's prediction results on the annotation task and the true label. The link loss is used to evaluate the relationship prediction effect between text blocks, which is further divided into three parts: mask link loss, positive sample link loss, and variance loss. Finally, the annotation loss and link loss are weighted and accumulated to form a total loss to guide the optimization of the model. During the training iteration process, two termination conditions are set to ensure the efficiency and stability of the training; the maximum number of training rounds is set to 400, and the training stops when the upper limit is reached; after each round of training, the performance of the model on the validation set is verified, and the current optimal model is saved to ensure that the final model has the best effect; this training process ensures the convergence of the model and the training efficiency through reasonable design and termination conditions.
2. A document layout analysis method based on a hybrid method according to claim 1, characterized in that: In step 2, first, read the batch data batch in the dataset; for the batch data, perform the following operations in sequence: use a convolutional neural network to extract the visual features of the image; use a pre-trained Transformer model to extract text features, and further optimize the feature extraction effect through a multi-layer adapter; fuse the extracted visual features and text features through a cross-attention mechanism to generate a cross-modal joint feature representation; use a graph neural network to model the contextual information between text blocks, transfer information between nodes, and further enhance the feature representation; finally, classify the final features of each text block, predict its category, and determine whether there is a connection between each pair of text blocks.
3. The document layout analysis method based on a hybrid method according to claim 1, characterized in that: In step 3.2, downsampling layers and convolution blocks are used to process image features in four stages, as follows: four downsampling layers are defined. The first downsampling layer uses a Conv2d with a stride of 4 for large-stride convolution to quickly reduce the input size by 4 times; then a LayerNorm is used for normalization; the next three downsampling layers each include a LayerNorm and a Conv2d with a stride of 2 and a convolution kernel of 2. Before proceeding to the next stage, the number of channels will also increase; four convolution block sequences are defined, and for each stage i, several Blocks will be constructed, the number of which is determined by depths[i]; each Block uses a depth-wise separable convolution dwconv with a stride of 3 and a kernel of 7 to extract features in the spatial dimension, and uses LayerNorm feature maps for normalization to improve training stability. 1×1 convolution, GELU nonlinear layer and 1×1 convolution are used, where 1x1 convolution is used for feature transformation in the channel dimension, and a nonlinear layer is introduced to improve the model's expressiveness; finally, residual connections and random depth are used to prevent overfitting and improve training stability; In the i-th stage, the i-th downsampling layer and the i-th convolution block sequence are used; the downsampling layer is first used to perform channel transformation or resolution scaling on the current feature map; then multiple blocks of this stage are entered for deep convolution feature extraction; the input shape is usually an image tensor of [batch_size, in_chans, H, W], which refers to the batch size, input channel, image feature length and width respectively; the output is a tuple containing 4 feature maps, each feature map corresponds to the output of a stage, and the shapes are: stage 1: [batch_size, dims[0], H / 4, W / 4]; stage 2: [batch_size, dims[1], H / 8, W / 8]; stage 3: [batch_size, dims[2], H / 16, W / 16]; stage 4: [batch_size, dims[3], H / 32, W / 32].
4. The document layout analysis method based on a hybrid method according to claim 1, characterized in that: The implementation steps of step 3.3 are as follows: Step 3.3.1: Channel attention block; First, the adaptive maximum pooling P max (·) and adaptive average pooling P avg (·) is applied to the spatial dimension to extract the most important features of the entire feature map of each channel, where the adaptive maximum pooling P max (·) and adaptive average pooling P avg The output size of (·) is set to 1×1, that is, the feature map of each channel is compressed into a single value to extract the global spatial information; then, the correlation between channels is learned through two 1×1 convolutions, that is, the first Conv 1×1 The number of channels is reduced from C to C / rc (C is the number of channels, rc is the channel compression ratio set to 16), and the second Conv 1×1 The number of channels is restored from C / rc to C, and the nonlinear expression is learned through the ReLU activation function in the middle; two features Y are obtained ca1 and Y ca2 ; Add and fuse the two features to generate the channel weight W ca ; The activation function uses Sigmoid to limit the weight value to [0, 1]; the generated channel weight W ca Multiply the input feature map X element by channel to adjust the feature strength of each channel; Step 3.3.2: Spatial attention block; First, along the channel direction, collect the maximum value Ch max (·) and the average value Ch avg (·) And the maximum and average features are concatenated in the channel dimension to focus on local features; a large 7×7 convolution kernel can capture larger local context information; the concatenated feature is Y max_avg , the number of channels is 2; then the Sigmoid function is applied to generate the spatial attention weight W sp ; Step 3.3.3: Multi-scale convolutional block; Perform depthwise convolution at multiple scales and use a channel shuffle operation to shuffle channels between groups; more specifically, first use Conv 1×1 (·) Expand the number of channels, followed by a batch normalization layer BN(·) and a ReLU6(·) activation layer; use multi-scale deep convolution DWC(·) [1×1, 3×3, 5×5] to capture multi-scale and multi-resolution context; and because deep convolution ignores the relationship between channels, a channel shuffle operation is used to merge the relationship between channels; finally, Conv 1×1 (·) The encoded multi-scale features F ms Restore and Enter X sp The same number of channels and adding residual connections; Step 3.3.4: Multi-scale image feature map fusion; The four input feature maps of different scales [X0, X1, X2, X3]∈X m , through the global average pooling operation P in the spatial dimension avg (·) Obtain global feature representation, and then pass Conv 1×1 (·) Model the correlation between channels and generate channel descriptors through the sigmoid activation function; concatenate the global features of the four scales in the channel dimension and obtain their respective weights W through the softmax function in the channel dimension. MSF Indicates that the generated weight is multiplied element-wise with the corresponding input and added; Step 3.3.5: Dynamic range convolution block; The feature map Y MSF The enhanced feature maps are divided into 4 groups by further mixing and sharing information through channel shuffling. The grouped feature maps are transposed to disrupt the channel order in each group and obtain feature map Y copy ; For feature map Y MSF Through dynamic range convolution, the convolution is processed between similar pixels rather than between adjacent pixels; specifically, the feature map Y is transformed along the channel dimension MSF Divided into two partsY MSF1 and Y MSF2 ; For the first part feature Y MSF1 The features are sorted in the W direction and the H direction in turn, so that the feature values gradually increase from the upper left corner to the lower right corner; the first feature branch after sorting It is reconnected with the second feature branch F2 to form a new feature map, and processed by depthwise separable convolution to obtain F out ; Organize high-intensity and low-intensity pixels into regular patterns at the diagonal of the matrix to achieve dynamic range convolution, so that similar pixels are learned by the convolution kernel; finally, separate the first part of the features and restore them in the H direction and W direction in turn to obtain F out ; For the feature map F′ out The same process is also performed, so that the overall channel features can learn the features between similar pixels through dynamic range convolution; F out and F′ out Concatenate and pass through a Conv 1×1 (·) Restore the original channel feature dimension; The output is divided into Q1, Q2, K1, K2, and V. First, V is reshaped into Then sort so that V is sorted from small to large in the HW dimension and the index value d is saved; Q1, Q2, K1, K2 are sorted according to the index d; given the number of bins B, the sorted shape is reshaped into and Perform self-attention on Q1, K1, V and Q2, K2, V; finally, multiply the two self-attention results to get the final output.
5. The document layout analysis method based on a hybrid method according to claim 1, characterized in that: Extract document language text features: Use PDF parser or OCR engine to scan the document line by line or block by block to detect text fragments; at the same time, record the position information of each text fragment in the page; finally get a data list in the form of (text, coordinates), convert the text and coordinate information into the following features: text embedding TokenEmbedding, map each text segment or token after word segmentation into a vector; one-dimensional position embedding 1DPositionalEmbedding, use standard position encoding to represent the position of each token in the sequence; Segment order embedding 1DSegmentOrderEmbedding adds a SegmentID to each token to distinguish sentences; segment BIE tag embedding 1DBIEEmbedding; B, I, E represent the beginning of a paragraph, the middle of a paragraph, and the end of a paragraph respectively; two-dimensional segment box embedding 2DBoundingBoxEmbedding normalizes the coordinate information (x1, y1, x2, y2) to a certain range, and then maps it to a coordinate vector of a certain dimension; Fine-tune the adapter layer added to the frozen pre-trained model to learn text features and obtain the features of each text. The frozen pre-trained model consists of 24 self-attention layers. Each layer consists of a self-attention mechanism, an intermediate layer of a feedforward network, and an output layer. The feedforward layer consists of two layers of linear transformation and an activation function. Two layers of adapters are added after the self-attention mechanism, the intermediate layer of the feedforward network, and the output layer of each layer. The adapter module first introduces a downsampling layer, the core goal of which is to effectively compress features while retaining key information. Through this design, the model can effectively extract key features in the image while reducing the computational burden, thereby improving operating efficiency. A ReLU activation function layer is added to the adapter module; through the upsampling layer, the original feature shape is restored to retain sufficient information; Freeze all parameters of the self-attention model architecture and only train the adapter parameters; this preserves the original network parameter settings, avoids the tedious process of training the network from scratch, and does not require saving a new parameter weight for each task, significantly reducing storage requirements.
Citation Information
Cited By
File reading sequence correction and optimization system based on graph neural network
CN121257476A
A document reading order correction and optimization system based on a graph neural network
CN121257476B
Table recognition method and device based on thinking chain, electronic equipment and storage medium
CN121259854A
Output length prediction method and system based on large model activation and sampling parameters
CN121524628A
A Method and System for Predicting Output Length Based on Large Model Activation and Sampling Parameters
CN121524628B