Document layout analysis method and device, equipment, storage medium and program product

By using object detection and image classification models in document layout analysis to determine the bounding boxes of document elements and calculate similarity, the problems of large data requirements and high model complexity are solved, thus improving the efficiency and accuracy of document layout analysis.

CN120877306APending Publication Date: 2025-10-31AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510987540.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies for document layout analysis suffer from problems such as large data requirements, difficulty in data acquisition, and high model complexity, resulting in insufficient analysis efficiency and accuracy.

Method used

By inputting the document image to be analyzed into the object detection model, the bounding boxes of the elements are obtained, and cropping is performed based on the bounding boxes. The first and second relationship features between the element image and the reference element image are determined using an image classification model, and the similarity is calculated to determine the element category.

Benefits of technology

It improves the efficiency and accuracy of document layout analysis by determining the similarity between element images and reference element images of each category through first and second relation features, thereby accurately classifying elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877306A_ABST
    Figure CN120877306A_ABST
Patent Text Reader

Abstract

The invention discloses a document layout analysis method and device, equipment, a storage medium and a program product. Inputting a to-be-analyzed document image into the target detection model to obtain a bounding box of each element; cutting each element based on the bounding box to obtain a plurality of element images; for each element image, determining a first relation feature and a second relation feature between the element image and a plurality of categories of reference element images based on an image classification model; determining the similarity between the element image and the reference element image of each category based on the first relation feature and the second relation feature; and determining the category of the element image according to the similarity. According to the document layout analysis method provided by the embodiment of the invention, the similarity between the element image and the reference element image of each category is determined through the first relation feature and the second relation feature, so that the category of the element is determined, and the efficiency and accuracy of document layout analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a document layout analysis method, apparatus, device, storage medium and program product. Background Technology

[0002] In recent years, with the rapid development of machine learning and artificial intelligence technologies, layout analysis technology has gradually become the mainstream method for automated document processing. Especially in areas such as deep learning and multimodal fusion, these technologies have not only improved the efficiency and accuracy of automated document processing but also provided new solutions for understanding and analyzing complex documents. However, existing methods still suffer from drawbacks such as large data requirements, difficulty in data acquisition, and high model complexity. Summary of the Invention

[0003] This invention provides a document layout analysis method, apparatus, device, storage medium, and program product, which can improve the efficiency and accuracy of document layout analysis.

[0004] In a first aspect, embodiments of the present invention provide a document layout analysis method, including:

[0005] Input the document image to be analyzed into the object detection model to obtain the bounding boxes of each element;

[0006] Based on the bounding box, each element is cropped to obtain multiple element images;

[0007] For each element image, a first relation feature and a second relation feature are determined between the element image and reference element images of multiple categories based on an image classification model;

[0008] The similarity between the element image and reference element images of each category is determined based on the first relation feature and the second relation feature.

[0009] The category of the element image is determined based on the similarity.

[0010] Secondly, embodiments of the present invention also provide a document layout analysis device, comprising:

[0011] The element detection module is used to input the document image to be analyzed into the object detection model to obtain the bounding boxes of each element;

[0012] The element cropping module is used to crop each element based on the bounding box to obtain multiple element images;

[0013] The relationship feature determination module is used to determine, for each element image, a first relationship feature and a second relationship feature between the element image and reference element images of multiple categories based on an image classification model;

[0014] The similarity determination module is used to determine the similarity between the element image and reference element images of various categories based on the first relation feature and the second relation feature;

[0015] An element category determination module is used to determine the category of the element image based on the similarity.

[0016] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the document layout analysis method described in the embodiments of the present invention.

[0020] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions, which are used to cause a processor to execute the document layout analysis method described in the embodiments of the present invention.

[0021] Fifthly, embodiments of the present invention provide a computer program product, including a computer program that, when executed by a processor, implements the document layout analysis method as described in the embodiments of the present invention.

[0022] This invention discloses a document layout analysis method, apparatus, device, storage medium, and program product. The method involves inputting a document image to be analyzed into a target detection model to obtain bounding boxes for each element; cropping each element based on the bounding boxes to obtain multiple element images; for each element image, determining first and second relationship features between the element image and multiple categories of reference element images based on an image classification model; determining the similarity between the element image and each category of reference element images based on the first and second relationship features; and determining the category of the element image based on the similarity. The document layout analysis method provided by this invention determines the element category by using first and second relationship features to determine the similarity between an element image and each category of reference element images, thereby improving the efficiency and accuracy of document layout analysis. Attached Figure Description

[0023] Figure 1 This is a flowchart of a document layout analysis method according to Embodiment 1 of the present invention;

[0024] Figure 2This is a schematic diagram of the structure of a document layout analysis device according to Embodiment 2 of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of an electronic device according to Embodiment 3 of the present invention. Detailed Implementation

[0026] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0027] Example 1

[0028] Figure 1 This is a flowchart of a document layout analysis method provided in Embodiment 1 of the present invention. This embodiment is applicable to the analysis of document layout. The method can be executed by a document layout analysis device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the specific steps include the following:

[0029] S110: Input the document image to be analyzed into the object detection model to obtain the bounding boxes of each element.

[0030] The document image can be obtained by image acquisition of a document, such as a contract or resume. The object detection model can be a pre-trained YOLO model used to detect various element regions in the document, and the elements can include titles, images, tables, and body text.

[0031] Specifically, the method for inputting the document image to be analyzed into the object detection model to obtain the bounding boxes of each element can be as follows: perform at least one of the following preprocessing steps on the document image to be analyzed: tilt correction, noise reduction, binarization, and image size scaling; input the preprocessed document image into the object detection model to obtain the bounding boxes of each element.

[0032] In this embodiment, to improve image quality, enhance feature representation, and reduce noise interference, preprocessing of the document image to be analyzed is necessary. During document image acquisition, image tilting can occur due to shooting angle or equipment instability, affecting detection results. Therefore, Hough transform is used to detect straight lines in the image, and rotation angles are calculated to correct the tilt, ensuring that targets in the document image maintain standard positions and orientations, thus improving model stability. Since salt-and-pepper noise often occurs during acquisition and annotation, this embodiment requires denoising of the element images, which can be achieved using median filtering, mean filtering, and Gaussian filtering. Preferably, this embodiment uses median filtering for denoising. Median filtering effectively removes isolated noise points by calculating the median pixel value within the window, while maintaining the integrity of edge features. Compared to mean filtering and Gaussian filtering, it has higher robustness and accuracy. To improve the contrast between text and background, binarization of the element images is required. In this embodiment, an adaptive thresholding algorithm (such as the Otsu thresholding method) can be used for binarization. The Otsu method automatically determines the threshold by maximizing the inter-class variance, which can stably segment the foreground and background under different lighting conditions, making the target outline clear and helping to improve the detection accuracy. To adapt to the input requirements of the neural network, this embodiment can use bilinear interpolation to scale the image to a set size (e.g., 224*224). The bilinear interpolation method effectively reduces scaling distortion by using a four-neighbor linear weighted average, balancing visual effect and computational efficiency.

[0033] In this embodiment, the preprocessed document image is input into the target detection model to obtain the bounding boxes of each element in the document image.

[0034] S120: Based on the bounding box, each element is cropped to obtain multiple element images.

[0035] In this embodiment, after the document image is processed by the object detection model, each element region in the image is enclosed by a bounding box. Each element region enclosed by the bounding box is cropped out from the document image, thereby obtaining multiple element images.

[0036] S130, For each element image, determine the first and second relation features between the element image and reference element images of multiple categories based on the image classification model.

[0037] The reference element image can be a pre-labeled element image, such as a reference element image with the category "title", a reference element image with the category "image", a reference element image with the category "table", and a reference element image with the category "body text". The first relation feature and the second relation feature are used to characterize the relevant and unrelated information between the element image and the reference element image.

[0038] In this embodiment, for each element image, the element image is paired with reference element images of each category to form an image pair, which are then input into the image classification model to determine the first and second relationship features between the element image and the reference element images of multiple categories.

[0039] The image classification model includes a multi-level embedding module, a cross-correlation module, and a collaborative attention module. Specifically, for each element image, the process of determining the first and second relationship features between the element image and multiple categories of reference element images based on the image classification model can be as follows: input the element image and the reference element images into the multi-level embedding module, and output the first and second multi-level features; input the first and second multi-level features into the cross-correlation module, and output the cross-correlation information; input the cross-correlation information, the first multi-level features, and the second multi-level features into the collaborative attention module, and output the first and second relationship features.

[0040] The first multi-level feature is the multi-level feature corresponding to the element image, and the second multi-level feature is the multi-level feature corresponding to the reference element image. The multi-level embedding module can use a ResNet18 network. The ResNet18 network includes an initial convolutional layer and eight basic blocks, each block outputting a tensor of different dimensions, which can be represented by the notation Z. t ∈R h×w×ct Where t∈{0,1,2,……,8} represents the layer number, h and w represent the height and width of the image, respectively, and c t This represents the number of channels in the tensor output at layer t. In this implementation, L outputs are selected from 9 tensors to form the L-level features.

[0041] In this embodiment, the multi-level embedding module is used to extract multi-level features from the element image and the reference element image respectively. Assume the element image is represented as I. q and reference element image I s Then the first multi-level feature corresponding to the element image can be represented as: The second multi-level feature corresponding to the element image can be represented as:

[0042] In this embodiment, the cross-correlation module is divided into two stages: intra-layer feature extraction and inter-layer feature extraction, used to extract the cross-correlation information between the element image and the reference element image. Specifically, the cross-correlation module includes: a fusion unit, an intra-layer feature extraction unit, and an inter-layer feature extraction unit. The process of inputting the first multi-level features and the second multi-level features into the cross-correlation module and outputting the cross-correlation information can be as follows: the fusion unit fuses the first multi-level features and the second multi-level features to obtain initial related information; the intra-layer feature extraction unit processes the initial related information to obtain intra-layer feature information; and the inter-layer feature extraction unit processes the intra-layer feature information to obtain cross-correlation information.

[0043] The intra-layer feature extraction unit and inter-layer feature extraction unit both include a layer normalization layer, a multi-head attention layer, and a multilayer perceptron. The method for fusing the first and second multi-level features based on the fusion unit can be as follows: First, each level feature in the first and second multi-level features is mapped one-to-one, and the corresponding two features are fused to obtain initial relevant information. That is, the initial relevant information is also a multi-level feature, which can be represented as:

[0044] Specifically, the process of fusing the first and second multi-level features based on the fusion unit can be as follows: First, perform a reshape() operation on both the first and second multi-level features. Then, perform a dot product on the two reshaped features to obtain initial relevant information. The reshape() operation changes the feature dimensions from h×w×c. t Change to hw×c t The formula for calculating the initial relevant information of the l-th layer can be expressed as: corr l The dimension is hw×hw. All the initial relevant information from each level is stacked layer by layer to form corr=∈R. hw×hw×L .

[0045] Specifically, the process of obtaining intra-layer feature information by processing initial relevant information based on the intra-layer feature extraction unit can be as follows: The initial relevant information is accumulated with the position embedding vector (pre-set) and then input into the layer normalization layer for processing. The processed result is then input into the multi-head attention layer, and finally into the multilayer perceptron to obtain the intra-layer feature information. The formula can be expressed as: corr' = MLP(Intra(LN(corr+E)) pos ))), where E pos Let M represent the position embedding vector, LN represent layer normalization, and Intra() represent the processing of the multi-head attention layer. The calculation process of Intra(M) can be represented as: head i =Attention(MW) iQ MW i K MW i V ), MultiHead(M)=Concat(head1,...head h W 0 Among them, W i Q W i K W i V W 0 For parameters in a multi-head attention layer, Attention(Q, K, V) = softmax(QK). T )V, MLP stands for Multilayer Perceptron.

[0046] Specifically, the process of processing intra-layer feature information based on inter-layer feature extraction units to obtain mutual correlation information can be as follows: Intra-layer feature information is accumulated with a pre-set position embedding vector and then input into a layer normalization layer for processing. The processed result is then input into a multi-head attention layer, and finally into a multilayer perceptron to obtain the mutual correlation information. The formula can be expressed as: corr map =MLP(Inter(LN(corr'+E) pos ))), where E pos Let M represent the position embedding vector, LN represent layer normalization, and Inter() represent the processing of the multi-head attention layer. The calculation process of Inter(M) can be represented as: head i =Attention(MW) i Q MW i K MW i V ), MultiHead(M)=Concat(head1,...head h W 0 Among them, W i Q W i K W i V W 0 For parameters in a multi-head attention layer, Attention(Q, K, V) = softmax(QK). T )V, MLP stands for Multilayer Perceptron.

[0047] In this embodiment, the cross-correlation module extracts intra-layer and inter-layer features from the initial relevant information, thereby removing some useless information to ensure the high accuracy of the cross-correlation information.

[0048] Among them, the related information corr map It contains information related to the consistency between the element image and the reference element image. map The dimension is hw×hw. The collaborative attention module is used to process the mutual related information to obtain the first relation feature and the second relation feature.

[0049] The collaborative attention module includes a mean unit, a product unit, and a pooling layer. Specifically, the method for inputting mutual information, the first multi-level feature, and the second multi-level feature into the collaborative attention module and outputting the first and second relationship features can be as follows: The mean unit performs a mean operation on the mutual information along both directions and then performs a dimensionality transformation to obtain the first and second collaborative features; the product unit performs a product operation on the first collaborative feature and the first multi-level feature, and on the second collaborative feature and the second multi-level feature, respectively, and then concatenates them along the channel dimension to obtain the first and second concatenated features; the pooling layer performs average pooling on the first and second concatenated features to obtain the first and second relationship features.

[0050] The process of performing averaging operations on the cross-related information along both directions based on the mean unit and then transforming the dimension can be as follows: First, perform averaging operations on the cross-related information along both the horizontal and vertical directions to obtain dimensions R respectively. 1×hw and R hw×1 We have two tensors, and then we perform a dimensionality transformation on these two tensors, transforming them into two features of dimension h×w, which are the first collaborative feature A. q Second cooperative feature A s .

[0051] The process of multiplying the first collaborative feature with the first multi-level feature using the product unit and then concatenating them along the channel dimension can be as follows: The first collaborative feature is multiplied by each layer of the first multi-level feature, and then the multi-level features are concatenated along the channel dimension. The calculation formula can be expressed as: The process of multiplying the second collaborative feature with the second multi-level feature using the product unit and then concatenating them along the channel dimension can be as follows: The second collaborative feature is multiplied by each layer of the second multi-level feature, and then the multi-level features are concatenated along the channel dimension. The calculation formula can be expressed as: E' q E' sThese represent the first and second concatenation features, respectively. Concat indicates concatenation along the channel dimension.

[0052] The average pooling process applied to the first and second concatenated features based on the pooling layer can be expressed as: q = pool(E') q ), s = pool(E' s ), where q and s represent the first relation feature and the second relation feature, respectively. pool() represents average pooling.

[0053] S140, determine the similarity between the element image and the reference element images of each category based on the first relation feature and the second relation feature.

[0054] The similarity can be cosine similarity. The calculation formula can be expressed as: Where γ is a scalar.

[0055] S150, determine the category of the element image based on similarity.

[0056] Specifically, the category of the reference element image that has the highest similarity to the element image is determined as the category of the element.

[0057] In this embodiment, the training method of the image classification model is as follows: acquire training set images and divide the training set images into a reference image set and a query image set; determine a first loss function based on the first multi-level features corresponding to the query image set; determine a second loss function based on the first and second relationship features corresponding to the reference image set and the query image set; and train the image classification model based on the first and second loss functions.

[0058] The reference image set includes N×K images, where N represents the number of categories and K represents the number of images in each category. Specifically, the method for determining the first loss function based on the first multi-level features corresponding to the query image set can be as follows: First, perform mean pooling on the first multi-level images to obtain feature E. q Finally based on E q Determine the first loss function. The formula for calculating the first loss function is: Where w and b are the weights and biases of the fully connected layer, and c is the channel.

[0059] In this embodiment, since the reference image set includes N×K images, N×K pairs of first relation features and second relation features can be obtained for each query image. The average of the K first relation features q in each category is then calculated to obtain... The average of the K second relation features s in each category is obtained. The formula for calculating the second loss function can be expressed as:

[0060] The process of training the image classification model based on the first loss function and the second loss function can be as follows: The first loss function and the second loss function are linearly superimposed, and the image classification model is then back-tuned based on the superimposed loss function. The formula for calculating the linear superposition of the first loss function and the second loss function is expressed as follows: α is a hyperparameter.

[0061] The technical solution of this embodiment involves inputting the document image to be analyzed into a target detection model to obtain bounding boxes for each element; cropping each element based on the bounding boxes to obtain multiple element images; for each element image, determining first and second relationship features between the element image and reference element images of multiple categories based on an image classification model; determining the similarity between the element image and each category of reference element images based on the first and second relationship features; and determining the category of the element image based on the similarity. The document layout analysis method provided by this embodiment of the invention determines the category of elements by using first and second relationship features to determine the similarity between element images and each category of reference element images, thereby improving the efficiency and accuracy of document layout analysis.

[0062] Example 2

[0063] Figure 2 This is a schematic diagram of the structure of a document layout analysis device provided in Embodiment 2 of the present invention, as shown below. Figure 2 As shown, the device includes:

[0064] The element detection module 210 is used to input the document image to be analyzed into the target detection model to obtain the bounding boxes of each element;

[0065] The element cropping module 220 is used to crop each element based on the bounding box to obtain multiple element images;

[0066] The relation feature determination module 230 is used to determine, for each element image, a first relation feature and a second relation feature between the element image and reference element images of multiple categories based on an image classification model;

[0067] Similarity determination module 240 is used to determine the similarity between the element image and reference element images of various categories based on the first relation feature and the second relation feature;

[0068] The element category determination module 250 is used to determine the category of the element image based on the similarity.

[0069] Optionally, the element detection module 210 is also used for:

[0070] Perform at least one of the following preprocessing steps on the document image to be analyzed: tilt correction, denoising, binarization, and image resizing;

[0071] The preprocessed document image is input into the object detection model to obtain the bounding boxes of each element.

[0072] Optionally, the image classification model includes: a multi-level embedding module, a cross-correlation module, and a collaborative attention module; for each element image, the relation feature determination module 230 is further used for:

[0073] The element image and the reference element image are input into the multi-level embedding module, and a first multi-level feature and a second multi-level feature are output; wherein, the first multi-level feature is the multi-level feature corresponding to the element image, and the second multi-level feature is the multi-level feature corresponding to the reference element image;

[0074] The first multi-level feature and the second multi-level feature are input into the cross-correlation module, and the cross-correlation information is output.

[0075] The mutual information, the first multi-level feature, and the second multi-level feature are input into the collaborative attention module, and the first relationship feature and the second relationship feature are output.

[0076] Optionally, the cross-correlation module includes: a fusion unit, an intra-layer feature extraction unit, and an inter-layer feature extraction unit; the relationship feature determination module 230 is further used for:

[0077] The first multi-level features and the second multi-level features are fused based on the fusion unit to obtain initial relevant information;

[0078] The initial relevant information is processed based on the intra-layer feature extraction unit to obtain intra-layer feature information;

[0079] The inter-layer feature extraction unit processes the intra-layer feature information to obtain mutual related information; wherein, both the intra-layer feature extraction unit and the inter-layer feature extraction unit include a layer normalization layer, a multi-head attention layer and a multilayer perceptron.

[0080] Optionally, the collaborative attention module includes: a mean unit, a product unit, and a pooling layer; the relation feature determination module 230 is further used for:

[0081] Based on the mean unit, the mutual related information is mean-valued in two directions and then transformed in dimension to obtain the first collaborative feature and the second collaborative feature.

[0082] Based on the product unit, the first collaborative feature and the first multi-level feature, and the second collaborative feature and the second multi-level feature are respectively multiplied and then concatenated along the channel dimension to obtain the first concatenated feature and the second concatenated feature.

[0083] The first splicing feature and the second splicing feature are subjected to average pooling based on the pooling layer to obtain the first relation feature and the second relation feature.

[0084] Optionally, it also includes: the image classification model training module, used for:

[0085] Acquire training set images and divide the training set images into a reference image set and a query image set;

[0086] A first loss function is determined based on the first multi-level features corresponding to the query image set;

[0087] The second loss function is determined based on the first and second relation features corresponding to the reference image set and the query image set;

[0088] The image classification model is trained based on the first loss function and the second loss function.

[0089] The above-described apparatus can execute the methods provided in all the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in this embodiment can be found in the methods provided in all the foregoing embodiments of the present invention.

[0090] Example 3

[0091] Figure 3 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components, connections and relationships between components, and their functions shown herein are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0092] like Figure 3As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0093] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0094] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as document layout analysis methods.

[0095] In some embodiments, the document layout analysis method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the document layout analysis method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the document layout analysis method by any other suitable means (e.g., by means of firmware).

[0096] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0097] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0098] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0099] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0100] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0101] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0102] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the document layout analysis method provided in any embodiment of this application.

[0103] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0104] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0105] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A document layout analysis method, characterized in that, include: Input the document image to be analyzed into the object detection model to obtain the bounding boxes of each element; Based on the bounding box, each element is cropped to obtain multiple element images; For each element image, a first relation feature and a second relation feature are determined between the element image and reference element images of multiple categories based on an image classification model; The similarity between the element image and reference element images of each category is determined based on the first relation feature and the second relation feature. The category of the element image is determined based on the similarity.

2. The method according to claim 1, characterized in that, The document image to be analyzed is input into the object detection model to obtain the bounding boxes of each element, including: Perform at least one of the following preprocessing steps on the document image to be analyzed: tilt correction, denoising, binarization, and image resizing; The preprocessed document image is input into the object detection model to obtain the bounding boxes of each element.

3. The method according to claim 1, characterized in that, The image classification model includes: a multi-level embedding module, a cross-correlation module, and a collaborative attention module; for each element image, based on the image classification model, a first relation feature and a second relation feature are determined between the element image and reference element images of multiple categories, including: The element image and the reference element image are input into the multi-level embedding module, and a first multi-level feature and a second multi-level feature are output; wherein, the first multi-level feature is the multi-level feature corresponding to the element image, and the second multi-level feature is the multi-level feature corresponding to the reference element image; The first multi-level feature and the second multi-level feature are input into the cross-correlation module, and the cross-correlation information is output. The mutual information, the first multi-level feature, and the second multi-level feature are input into the collaborative attention module, and the first relationship feature and the second relationship feature are output.

4. The method according to claim 3, characterized in that, The cross-correlation module includes: a fusion unit, an intra-layer feature extraction unit, and an inter-layer feature extraction unit; The first multi-level features and the second multi-level features are input into the cross-correlation module, and the cross-correlation information is output, including: The first multi-level features and the second multi-level features are fused based on the fusion unit to obtain initial relevant information; The initial relevant information is processed based on the intra-layer feature extraction unit to obtain intra-layer feature information; The inter-layer feature extraction unit processes the intra-layer feature information to obtain mutual related information; wherein, both the intra-layer feature extraction unit and the inter-layer feature extraction unit include a layer normalization layer, a multi-head attention layer and a multilayer perceptron.

5. The method according to claim 3, characterized in that, The collaborative attention module includes: a mean unit, a product unit, and a pooling layer; The mutual information, the first multi-level feature, and the second multi-level feature are input into the collaborative attention module, and the first relation feature and the second relation feature are output, including: Based on the mean unit, the mutual related information is mean-valued in two directions and then transformed in dimension to obtain the first collaborative feature and the second collaborative feature. Based on the product unit, the first collaborative feature and the first multi-level feature, and the second collaborative feature and the second multi-level feature are respectively multiplied and then concatenated along the channel dimension to obtain the first concatenated feature and the second concatenated feature. The first splicing feature and the second splicing feature are subjected to average pooling based on the pooling layer to obtain the first relation feature and the second relation feature.

6. The method according to claim 3, characterized in that, The training method for the image classification model is as follows: Acquire training set images and divide the training set images into a reference image set and a query image set; A first loss function is determined based on the first multi-level features corresponding to the query image set; The second loss function is determined based on the first and second relation features corresponding to the reference image set and the query image set; The image classification model is trained based on the first loss function and the second loss function.

7. A document layout analysis device, characterized in that, include: The element detection module is used to input the document image to be analyzed into the object detection model to obtain the bounding boxes of each element; The element cropping module is used to crop each element based on the bounding box to obtain multiple element images; The relationship feature determination module is used to determine, for each element image, a first relationship feature and a second relationship feature between the element image and reference element images of multiple categories based on an image classification model; The similarity determination module is used to determine the similarity between the element image and reference element images of various categories based on the first relation feature and the second relation feature; An element category determination module is used to determine the category of the element image based on the similarity.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the document layout analysis method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the document layout analysis method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the document layout analysis method as described in any one of claims 1-6.