Layout analysis method, layout analysis device, electronic equipment and computer medium

By extracting and integrating visual features and natural language features of the layout, the problem of low detection accuracy of picture and text in the existing technology of layout analysis is solved, and more efficient layout information detection and separation is achieved.

CN119992580APending Publication Date: 2025-05-13BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311490258.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the detection accuracy of layout analysis on pictures and text is low, making it difficult to effectively identify and separate text and picture information in the layout.

Method used

By extracting the visual features and natural language features of the layout and fusing them, the text information in the layout block is processed using the pre-trained model to obtain natural language features and combine them with visual features to finally obtain the layout analysis results.

Benefits of technology

It realizes simultaneous detection of text content and pictures, improves the feature representation ability, and thus improves the detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992580A_ABST
    Figure CN119992580A_ABST
Patent Text Reader

Abstract

The invention discloses a layout analysis method, a layout analysis device, electronic equipment and a computer readable storage medium. The layout analysis method comprises the following steps: extracting visual features of a layout; performing text analysis on the layout to obtain text information in the layout; dividing the layout into a plurality of layout blocks; processing the text information in the layout blocks through a pre-training model to obtain natural language features; fusing the visual features and the natural language features to obtain fused features; and processing the fusion features to obtain a layout analysis result. According to the layout analysis method, when the natural language features are extracted, the layout is divided into the layout blocks, the text information in the layout blocks is processed through the pre-training model to obtain the natural language features, meanwhile, the visual features are extracted, the visual features and the natural language features are fused and processed, and the layout analysis result is obtained. The text content and the picture are detected at the same time, the characterization capability of the features is improved, and therefore the detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of layout analysis technology, and more specifically, to a layout analysis method, a layout analysis device, an electronic device and a computer-readable storage medium. Background Art

[0002] Layout analysis is the process of automatically analyzing, identifying and understanding the images, texts, table information and positional relationships within a layout. Through layout analysis, the readable text can be separated from the title, page number, pictures, etc. In related technologies, the detection accuracy of layout analysis for pictures and texts is low. Summary of the invention

[0003] Embodiments of the present invention provide a layout analysis method, a layout analysis device, an electronic device, and a computer-readable storage medium.

[0004] An embodiment of the present invention provides a layout analysis method, which includes: extracting visual features of a layout; performing text analysis on the layout to obtain text information in the layout; dividing the layout into multiple layout blocks; processing the text information in the layout blocks through a pre-trained model to obtain natural language features; fusing the visual features and the natural language features to obtain fused features; and processing the fused features to obtain a layout analysis result.

[0005] In this way, the layout is divided into multiple layout blocks when extracting natural language features, and the text information in the layout blocks is processed through a pre-trained model to obtain natural language features, and visual features are extracted at the same time, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and images, and improving the feature representation ability, thereby improving the detection accuracy.

[0006] In some embodiments, the processing of the text information in the layout block by a pre-trained model to obtain natural language features includes: processing the text information in the layout block by the pre-trained model to map the text information into a natural language vector; linearly processing the natural language vector to obtain the natural language vector as a natural language feature.

[0007] In this way, by processing the text information in the layout through the pre-trained model, the text information can be mapped into a natural language vector, and then through linear processing of the natural language vector, the natural language vector can be obtained as a natural language feature, thereby realizing the extraction of natural language features.

[0008] In some embodiments, extracting the visual features of the layout includes: dividing the layout into a plurality of layout blocks; inputting the layout blocks into a linear projection layer for processing to obtain visual vectors; and position encoding the visual vectors to obtain the visual features.

[0009] In this way, the layout is divided into multiple layout blocks, and then the layout blocks are input into the linear projection layer for processing, so that the visual vector can be obtained, and the visual vector is position-encoded to obtain the visual features of the set dimension. After block embedding and position encoding, the visual features in the layout can be extracted.

[0010] In some embodiments, the visual features include visual information features and position information features, and the fusing of the visual features and the natural language features to obtain fused features includes: fusing the visual information features and the natural language features to obtain an initial fused feature; adding the position information features to the initial fused feature to obtain the fused feature.

[0011] In this way, the visual information features in the visual features are first fused with the natural language features to obtain the initial fused features, and then the position information is added to the initial fused features to obtain the fused features, so as to realize the fusion of natural language features and visual features of different dimensions.

[0012] In some embodiments, the processing of the fused features to obtain a layout analysis result includes: inputting the fused features into a sequence model of an attention mechanism for processing to obtain multiple output vectors of the same scale; processing the output vectors to obtain the layout analysis result.

[0013] In this way, by inputting the fused features into a sequence model of an attention mechanism such as Transformer, multiple output vectors of the same scale can be obtained. By processing the output vectors, the results of the layout analysis can be obtained to complete the analysis of the layout.

[0014] In some embodiments, the processing of the output vector to obtain the layout analysis result includes: sampling the output vector; upsampling the sampled output vector to obtain a feature pyramid, wherein the feature pyramid includes multiple feature vectors of different scales; and processing the feature pyramid to obtain the layout analysis result.

[0015] In this way, by sampling the output vector and upsampling the sampled output vector to obtain a feature pyramid, and then processing the feature pyramid, the structure of the layout analysis can be obtained to complete the analysis of the layout.

[0016] In some embodiments, the processing of the feature pyramid to obtain the layout analysis result includes: using a target detection head of Yolov5 to perform target detection on the feature pyramid, and outputting the detection result as the layout analysis result.

[0017] In this way, the target detection head of Yolov5 is used to detect the target on the feature pyramid, and three convolutional layers are used to process the feature pyramid to obtain the layout analysis result, which has higher detection accuracy and faster speed.

[0018] An embodiment of the present invention provides a layout analysis device, which includes: an extraction module, a parsing module, a division module, a first processing module, a fusion module and a second processing module, wherein the extraction module is used to extract visual features of the layout; the parsing module is used to perform text parsing on the layout to obtain text information in the layout; the division module is used to divide the layout into multiple layout blocks; the first processing module is used to process the text information in the layout block through a pre-trained model to obtain natural language features; the fusion module is used to fuse the visual features and the natural language features to obtain fused features; the second processing module is used to process the fused features to obtain a layout analysis result.

[0019] In this way, by dividing the layout into multiple layout blocks when extracting natural language features, and processing the text information in the layout blocks through a pre-trained model to obtain natural language features, the accuracy of extracting natural language features is improved; and at the same time, visual features are extracted, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and images and improving the detection accuracy.

[0020] An embodiment of the present invention provides an electronic device, which includes one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the layout analysis method of any of the above embodiments are implemented.

[0021] In this way, by dividing the layout into multiple layout blocks when extracting natural language features, and processing the text information in the layout blocks through a pre-trained model to obtain natural language features, the accuracy of extracting natural language features is improved; and at the same time, visual features are extracted, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and images and improving the detection accuracy.

[0022] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the layout analysis method of any of the above embodiments are implemented.

[0023] In this way, by dividing the layout into multiple layout blocks when extracting natural language features, and processing the text information in the layout blocks through a pre-trained model to obtain natural language features, the accuracy of extracting natural language features is improved; and at the same time, visual features are extracted, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and images and improving the detection accuracy.

[0024] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0026] Figure 1 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0027] Figure 2 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0028] Figure 3 is a schematic diagram of a layout analysis device according to an embodiment of the present invention;

[0029] Figure 4 is a schematic diagram of a layout analysis result according to an embodiment of the present invention;

[0030] Figure 5 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0031] Figure 6 is a schematic diagram of a first processing module according to an embodiment of the present invention;

[0032] Figure 7 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0033] Figure 8 is a schematic diagram of an extraction module according to an embodiment of the present invention;

[0034] Fig. 9 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0035] Fig.10 is a schematic diagram of a fusion module according to an embodiment of the present invention;

[0036] Fig.11 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0037] Fig.12 is a schematic diagram of a second processing module according to an embodiment of the present invention;

[0038] Fig.13 is a schematic diagram of an encoder according to an embodiment of the present invention;

[0039] Fig.14 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0040] Fig.15 is a schematic diagram of a fourth processing submodule according to an embodiment of the present invention;

[0041] Fig.16 It is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;

[0042] Fig.17 It is a flowchart of a method for training a layout analysis model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The embodiments of the present invention are described in detail below, and the embodiments of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0044] Layout analysis is the process of automatically analyzing, identifying and understanding the images, texts, table information and positional relationships within a layout. Through layout analysis, the readable text can be separated from the title, page number, pictures, etc. In related technologies, the detection accuracy of layout analysis for pictures and texts is low.

[0045] See also Figure 1 and Figure 2 The embodiment of the present invention provides a layout analysis method, the layout analysis method comprising:

[0046] 01: Extract visual features of layout 200;

[0047] 02: Perform text analysis on page 200 to obtain text information in page 200;

[0048] 03: Divide the layout 200 into a plurality of layout blocks 201;

[0049] 04: Processing the text information in the layout block 201 through the pre-training model 300 to obtain natural language features;

[0050] 05: Fusion of visual features and natural language features to obtain fusion features;

[0051] 06: Process the fused features to obtain layout analysis results.

[0052] Specifically, see Figure 3 The layout analysis method of the embodiment of the present invention can be implemented by the layout analysis device 100 of the embodiment of the present invention. The layout analysis device 100 includes: an extraction module 10, a parsing module 20, a division module 30, a first processing module 40, a fusion module 50 and a second processing module 60, wherein step 01 can be implemented by the extraction module 10, step 02 can be implemented by the parsing module 20, step 02 can be implemented by the division module 30, step 02 can be implemented by the first processing module 40, step 02 can be implemented by the fusion module 50, and step 02 can be implemented by the second processing module 60. That is to say, The extraction module 10 can be used to extract the visual features of the layout 200; the analysis module 20 can be used to perform text analysis on the layout 200 to obtain the text information in the layout 200; the division module 30 can be used to divide the layout 200 into multiple layout blocks 201; the first processing module 40 can be used to process the text information in the layout block 201 through the pre-trained model 300 to obtain natural language features; the fusion module 50 can be used to fuse the visual features and the natural language features to obtain fused features; the second processing module 60 can be used to process the fused features to obtain the layout analysis results.

[0053] In addition, natural language usually refers to a language that evolves naturally with culture. In order to enable machines to communicate with humans using natural language, NLP (Natural Language Processing) technology has emerged and continues to develop. NLP can be broadly defined as the automatic analysis, processing and operation of natural languages ​​such as voice and text through software. Natural language features can be represented by NLP features. In order to extract NLP features from layout 200, the format of layout 200 must first be converted into a corresponding PDF. Text parsing of layout 200 in PDF format can obtain text information in layout 200. In addition, text box information in layout 200 can also be obtained. In related technologies, the entire layout 200 is extracted during the NLP feature extraction process, so the extraction effect is poor. The embodiment of the present invention first divides the layout 200 into multiple layout blocks 201 (patches). The size of each layout block 201 can be fixed and consistent. For example, the size of the layout block 201 can be 16×16. The text information contained in each layout block 201 is gathered in one place. Natural language features are extracted from the divided layout blocks 201, and visual features are extracted at the same time. The visual features and natural language features are fused, and the fused features obtained after fusion are processed to obtain the layout analysis results. Please refer to Figure 4 The layout analysis results include text content (Text), title (Section-header), image (Picture) and page footer (Page-footer), etc.

[0054] In this way, the layout 200 is divided into multiple layout blocks 201 when extracting natural language features, and the text information in the layout blocks is processed through the pre-trained model 300 to obtain natural language features, and visual features are extracted at the same time, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and pictures, and improving the feature representation ability, thereby improving the detection accuracy.

[0055] See also Figure 5 In some embodiments, step 04 (processing the text information in the layout block 201 by the pre-trained model 300 to obtain natural language features) includes:

[0056] 041: Processing the text information in the layout block 201 through the pre-trained model 300 to map the text information into a natural language vector;

[0057] 042: Linearly process the natural language vector to obtain the natural language vector as a natural language feature.

[0058] Specifically, see Figure 6 The first processing module 40 includes a first processing sub-module 41 and a linear processing sub-module 42, wherein step 041 can be implemented by the first processing sub-module 41, and step 042 can be implemented by the linear processing sub-module 42. In other words, the first processing sub-module 41 can be used to process the text information in the layout block 201 through the pre-trained model 300 to map the text information into a natural language vector; the linear processing sub-module 42 can be used to linearly process the natural language vector to obtain the natural language vector as a natural language feature.

[0059] Among them, the pre-trained model 300 can be a model such as a Bert (Bidirectional Encoder Representations from Transformers) model, and the Bert model can map text information such as words and sentences into natural language vectors. The set dimension can be 2500×768 dimensions. If the number of text information contained in each layout block 201 is M, and each text information can generate a 64-dimensional natural language vector through the Bert model, then the text information of each layout block 201 can generate an M×64-dimensional natural language vector after being processed by the Bert model, and then the natural language vector is supplemented to a L×64-dimensional natural language vector with a set long dimension. All natural language vectors are reshaped to obtain a 2500×L×64-dimensional natural language vector, and the reshaped natural language vector is linearly processed to obtain a 2500×768-dimensional natural language vector as a natural language feature.

[0060] In this way, by processing the text information in the layout 200 through the pre-training model 300, the text information can be mapped into a natural language vector, and then through linear processing of the natural language vector, the natural language vector can be obtained as a natural language feature, thereby realizing the extraction of natural language features with better extraction effect.

[0061] See also Figure 7 In some embodiments, step 01 (extracting visual features of layout 200) includes:

[0062] 011: Divide the layout 200 into a plurality of layout blocks 201;

[0063] 012: Input the layout block 201 into the linear projection layer for processing to obtain a visual vector;

[0064] 013: Positionally encode the visual vector to obtain visual features.

[0065] Specifically, see Figure 8The extraction module 10 includes a division submodule 11, a second processing submodule 12 and a position encoding submodule 13, wherein step 011 can be implemented by the division submodule 11, step 012 can be implemented by the second processing submodule 12, and step 013 can be implemented by the position encoding submodule 13, that is, the division submodule 11 can be used to divide the layout 200 into multiple layout blocks 201 of a set size; the second processing submodule 12 can be used to input the layout blocks 201 into the linear projection layer for processing to obtain a visual vector; the position encoding submodule 13 can be used to perform position encoding on the visual vector to obtain a visual feature.

[0066] In addition, the size of the layout 200 can be 800×800, and the batch size (Batch_size) is set to 1. The layout 200 is divided into multiple layout blocks 201 (patch), and the size of the layout blocks 201 divided when extracting visual features can be consistent with the size of the layout blocks 201 divided when extracting natural language features. The size of the layout block 201 can be 16×16, so each layout 200 can generate 800×800 / (16×16)=2500 layout blocks 201, and the length of the sequence that can be generated is 2500. The dimension of each layout block 201 is 768. Since the dimension of the linear projection layer (Linear) is also 768×N, the dimension of the visual vector obtained after the layout block 201 is input into the linear projection layer for processing is 2500×768, that is, there are a total of 2500 words (tokens), and the dimension of each token is 768. Since special characters (cls) need to be added to the visual vector, the final dimension of the visual vector is 2501×768. The above process is the process of patch embedding. After patch embedding, a visual problem is transformed into a sequence-to-sequence (seq2seq) problem. Then the visual vector is position encoded. Position encoding can be understood as a table with N rows. The size of N is the same as the length of the input sequence (2501). Each row represents a vector. The dimension of the vector is the same as the dimension of the input sequence embedding (768). Since the way to add position encoding information is sum (sum) instead of concat (connection), the dimension of the visual feature obtained after adding position encoding is also 2501×768.

[0067] In this way, the layout 200 is divided into a plurality of layout blocks 201, and then the layout blocks 201 are input into the linear projection layer for processing, so that a visual vector can be obtained, and the visual vector is position-encoded to obtain a visual feature of a set dimension. After block embedding and position coding, the visual feature in the layout 200 can be extracted.

[0068] See also Fig. 9 In some embodiments, the visual features include visual information features and position information features, and step 05 (fusing the visual features and the natural language features to obtain fused features) includes:

[0069] 051: Fuse visual information features and natural language features to obtain initial fusion features;

[0070] 052: Add the location information feature to the initial fusion feature to obtain the fusion feature.

[0071] Specifically, see Fig.10 The fusion module 50 includes a first fusion submodule 51 and a second fusion submodule 52, wherein step 051 can be implemented by the first fusion submodule 51, and step 052 can be implemented by the second fusion submodule 52, that is, the first fusion submodule 51 can be used to fuse visual information features and natural language features to obtain initial fusion features; the second fusion submodule 52 can be used to add position information features to the initial fusion features to obtain fusion features.

[0072] In addition, since the dimension of visual features is 2501×768 and the dimension of natural language features is 2500×768, the dimensions of visual features and natural language features are different and cannot be fused. nlp ) and the visual information features with a dimension of 2500×768 in the visual features (feature vision ) are fused to obtain the initial fusion feature (feature merge ):

[0073] feature merge =featrue vision +feature nlp

[0074] The dimension of the initial fusion feature is 2500×768, and the location information feature is added to the initial fusion feature to obtain the fusion feature, and the dimension of the fusion feature is 2501×768.

[0075] In this way, the visual information features in the visual features are first fused with the natural language features to obtain the initial fused features, and then the position information is added to the initial fused features to obtain the fused features, so as to realize the fusion of natural language features and visual features of different dimensions.

[0076] See also Fig.11 In some embodiments, step 06 (processing the fused features to obtain layout analysis results) includes:

[0077] 061: Input the fused features into the sequence model of the attention mechanism for processing to obtain multiple output vectors of the same scale;

[0078] 062: Process the output vector to obtain the layout analysis result.

[0079] Specifically, see Fig.12 The second processing module 60 includes a third processing sub-module 61 and a fourth processing sub-module 62, wherein step 061 can be implemented by the third processing sub-module 61, and step 062 can be implemented by the fourth processing sub-module 62, that is, the third processing sub-module 61 can be used to input the fusion features into the sequence model of the attention mechanism for processing to obtain multiple output vectors of the same scale; the fourth processing sub-module 62 can be used to process the output vector to obtain the layout analysis result.

[0080] In addition, the second processing module 60 can be used as a backbone structure. The sequence model of the attention mechanism can be a Transformer model. The Transformer model includes: Fig.13 The encoder shown in FIG. 1 is composed of multiple encoder layers stacked together, each of which includes two sub-layer connection structures. The first sub-layer connection structure includes a multi-head self-attention sub-layer (Multi-Head-Attention) and a normalization and residual connection layer (Add&Norm), and the second sub-layer connection structure includes a feed-forward fully connected sub-layer (Feed Forward) and a normalization and residual connection layer (Add&Norm). The fused features are input into the encoder of the Transformer model for processing, and the dimensions of the output vectors are all 2501×768. By processing the output vectors, the results of the layout analysis can be obtained to complete the analysis of layout 200.

[0081] In this way, by inputting the fused features into a sequence model of an attention mechanism such as Transformer, multiple output vectors of the same scale can be obtained. By processing the output vectors, the results of the layout analysis can be obtained to complete the analysis of layout 200.

[0082] See also Fig.14 In some embodiments, step 062 (processing the output vector to obtain layout analysis results) includes:

[0083] 0621: Sample the output vector;

[0084] 0622: Upsampling the sampled output vector to obtain a feature pyramid, which includes feature vectors of multiple scales;

[0085] 0623: Process the feature pyramid to obtain layout analysis results.

[0086] Specifically, see Fig.15 The fourth processing submodule 62 includes a sampling unit 621, an upsampling unit 622 and a processing unit 623, wherein step 0621 can be implemented by the sampling unit 621, step 0622 can be implemented by the upsampling unit 622, and step 0623 can be implemented by the processing unit 623, that is, the sampling unit 621 can be used to sample the output vector; the upsampling unit 622 can be used to upsample the sampled output vector to obtain a feature pyramid, the feature pyramid includes multiple feature vectors of different scales; the processing unit 623 can be used to process the feature pyramid to obtain a layout analysis result.

[0087] In addition, the fourth processing submodule 62 may be a neck structure in the Yolov5 model. Sample the output vectors of the same scale output by different layers of the Transformer encoder, for example, extract the output vector feature of the Nth layer of the encoder encoder1 , the output vector feature of the N-2th layer encoder2 , the output vector feature of the N-4th layer encoder3 , and upsample the sampled output vector in different processing methods to obtain a feature pyramid. Since the dimension of the output vector is 2501×768, in order to facilitate subsequent processing, only the part with a dimension of 2500×768 is taken and reshaped to obtain feature vectors of different scales. The feature pyramid includes feature vectors of different scales Scale1, Scale2, Scale3:

[0088] Scale1 = fun1(feature encoder1 )

[0089] Scale2 = fun2(feature encoder2 )

[0090] Scale3 = fun3 (feature encoder3 )

[0091] Among them, the dimension of Scale1 is 768×100×100, the dimension of Scale2 is 768×50×50, and the dimension of Scale3 is 768×25×25. The layout analysis result can be obtained by processing the feature pyramid containing multi-scale feature vectors.

[0092] In this way, by sampling the output vector and upsampling the sampled output vector to obtain a feature pyramid, and then processing the feature pyramid, a layout analysis structure can be obtained to complete the analysis of the layout 200.

[0093] See also Fig.16 In some embodiments, step 0623 (processing the feature pyramid to obtain layout analysis results) includes:

[0094] 06231: Use Yolov5’s object detection head to perform object detection on the feature pyramid and output the detection results as layout analysis results.

[0095] Specifically, the processing unit 623 includes a detection subunit, and step 06231 can be implemented by the detection subunit, that is, the detection subunit can be used to use the target detection head of Yolov5 to perform target detection on the feature pyramid and output the detection result as the layout analysis result.

[0096] In the related art, the detection head structure of the Cascade R-CNN model is usually used for layout analysis. The embodiment of the present invention uses the target detection head of Yolov5 to detect the target on the feature pyramid. The target detection head of Yolov5 uses three detectors (Detect) to process the feature pyramid, which has higher detection accuracy and faster speed. Each Detect is a 1×1 convolution layer. The multi-scale feature vectors in the feature pyramid are subjected to the convolution layer to reduce the dimension of the number of channels and scale the feature vectors. Then, the feature vectors of different scales are fused to obtain richer feature information, thereby improving the detection performance.

[0097] In this way, the target detection head of Yolov5 is used to detect the target on the feature pyramid, and three convolutional layers are used to process the feature pyramid to obtain the layout analysis result, which has higher detection accuracy and faster speed.

[0098] See also Fig.17 The embodiment of the present invention provides a training method for a layout analysis model. Through the training method, a model to be trained can be trained according to a data set to obtain a layout analysis model. The data set includes multiple training layouts. The layout analysis model can implement the steps of the layout analysis method of any of the above embodiments. The training method includes:

[0099] 071: Extract visual features of training layout;

[0100] 072: Perform text parsing on the training page to obtain text information in the training page;

[0101] 073: Divide the training layout into multiple training layout blocks;

[0102] 074: Process the text information in the training page block by setting a pre-training model to obtain natural language features;

[0103] 075: Fusion of visual features and natural language features to obtain fusion features;

[0104] 076: Process the fused features to obtain the training layout analysis results;

[0105] 077: Train the model to be trained according to the training layout analysis result to obtain a layout analysis model.

[0106] Specifically, the explanation of the layout analysis method in the above embodiment is applicable to the training method of the layout analysis model in this embodiment, and will not be repeated here.

[0107] In this way, the layout analysis model trained by the training method improves the accuracy of extracting natural language features by dividing the layout 200 into multiple layout blocks 201 when extracting natural language features, and processes the text information in the layout block 201 through the pre-trained model 300 to obtain natural language features; and simultaneously extracts visual features, merges and processes visual features and natural language features to obtain layout analysis results, thereby realizing the simultaneous detection of text content and images and improving detection accuracy.

[0108] An embodiment of the present invention provides an electronic device, which includes one or more processors and a memory. The memory stores a computer program. When the computer program is executed by the processor, the steps of the layout analysis method of any of the above embodiments are implemented.

[0109] In this way, by dividing the layout 200 into multiple layout blocks 201 when extracting natural language features, and processing the text information in the layout block 201 through the pre-trained model 300 to obtain natural language features, the accuracy of extracting natural language features is improved; and at the same time, visual features are extracted, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and pictures, and improving the detection accuracy.

[0110] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the layout analysis method of any of the above embodiments are implemented.

[0111] In this way, by dividing the layout 200 into multiple layout blocks 201 when extracting natural language features, and processing the text information in the layout block 201 through the pre-trained model 300 to obtain natural language features, the accuracy of extracting natural language features is improved; and at the same time, visual features are extracted, and the visual features and natural language features are fused and processed to obtain the layout analysis results, thereby realizing the simultaneous detection of text content and pictures, and improving the detection accuracy.

[0112] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0113] In addition, the term "connection" should be understood in a broad sense, for example, it can include fixed connection, detachable connection, or integral connection; it can include direct connection, indirect connection through an intermediate medium, and internal communication between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0114] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0115] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present invention belong.

[0116] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A layout analysis method, characterized in that: The layout analysis method comprises: Extract visual features of the layout; Performing text analysis on the page layout to obtain text information in the page layout; Dividing the layout into a plurality of layout blocks; Processing the text information in the layout block through a pre-trained model to obtain natural language features; fusing the visual feature and the natural language feature to obtain a fused feature; The fusion features are processed to obtain layout analysis results.

2. The layout analysis method according to claim 1, characterized in that: The processing of the text information in the layout block by the pre-trained model to obtain natural language features includes: Processing the text information in the layout block by using the pre-trained model to map the text information into a natural language vector; The natural language vector is linearly processed to obtain a natural language vector as a natural language feature.

3. The layout analysis method according to claim 1, characterized in that: The extracting of visual features of the layout includes: Dividing the layout into a plurality of layout blocks; Inputting the layout block into a linear projection layer for processing to obtain a visual vector; Position encoding is performed on the visual vector to obtain the visual feature.

4. The layout analysis method according to claim 3, characterized in that: The visual features include visual information features and position information features, and the fusing the visual features and the natural language features to obtain fused features includes: fusing the visual information feature and the natural language feature to obtain an initial fused feature; The position information feature is added to the initial fused feature to obtain the fused feature.

5. The layout analysis method according to claim 1, characterized in that: The processing of the fusion features to obtain a layout analysis result includes: Inputting the fused features into a sequence model of an attention mechanism for processing to obtain multiple output vectors of the same scale; The output vector is processed to obtain the layout analysis result.

6. The layout analysis method according to claim 5, characterized in that: The processing of the output vector to obtain the layout analysis result includes: Sampling the output vector; Upsampling the sampled output vector to obtain a feature pyramid, wherein the feature pyramid includes feature vectors of multiple different scales; The feature pyramid is processed to obtain the layout analysis result.

7. The layout analysis method according to claim 6, characterized in that: The processing of the feature pyramid to obtain the layout analysis result includes: The target detection head of Yolov5 is used to perform target detection on the feature pyramid, and the detection result is output as the layout analysis result.

8. A layout analysis device, characterized in that: The layout analysis device comprises: An extraction module, the extraction module is used to extract visual features of the layout; A parsing module, the parsing module is used to perform text parsing on the layout to obtain text information in the layout; A division module, the division module is used to divide the layout into a plurality of layout blocks; A first processing module, the first processing module is used to process the text information in the layout block through a pre-trained model to obtain natural language features; A fusion module, the fusion module is used to fuse the visual feature and the natural language feature to obtain a fusion feature; The second processing module is used to process the fusion features to obtain layout analysis results.

9. An electronic device, characterized in that: The electronic device includes one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the layout analysis method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the layout analysis method according to any one of claims 1 to 7 are implemented.