A layout analysis method, device, equipment and storage medium

Through the combined model of dual-stream network and graph neural network, the layout analysis problem of cross-page correlation of question diagrams is solved, efficient and accurate matching and recognition of question diagrams is achieved, and the effect of layout analysis is improved.

CN119091459BActive Publication Date: 2025-07-11SHENZHEN CHUANGZHI MINIMALIST TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411590795.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-07-11
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

The existing technical solutions are difficult to effectively realize the cross-page relationship between the question chart, resulting in insufficient function of cross-page relationship in layout analysis.

Method used

The layout analysis model based on dual-stream network and graph neural network is adopted, and the multi-convolution neural network structure and the cross-page correlation network structure of the question graph are used to match and identify the question and question graph. The graph structure data processing capabilities of the graph neural network are used, and feature fusion and relationship modeling are combined with the image segmentation model and attention mechanism.

Benefits of technology

It realizes efficient correlation between the question map across pages, improves the accuracy and efficiency of layout analysis, and ensures accurate matching between the question map and the question.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091459B_ABST
    Figure CN119091459B_ABST
Patent Text Reader

Abstract

The present invention discloses a layout analysis method, device, equipment and medium. By preprocessing the image to be processed, an initial image after preprocessing is obtained, and the initial image is input into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result of matching between the title and the figure associated with the title. Among them, the graph neural network includes a plurality of convolutional neural network structures and a figure-cross-page association network structure, and each convolutional neural network structure has different receptive fields from large to small in turn. Compared with the prior art, the present application inputs the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result of matching between the title and the figure associated with the title, which can effectively analyze the layout of the test paper, and through the figure-cross-page association network, perform cross-page association on the figure associated with the title and the title, so as to better achieve the layout effect of figure-cross-page association, and improve the accuracy and efficiency of layout analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a layout analysis method, apparatus, device, and medium. Background Art

[0002] With the exponential increase in the production and storage requirements of a large number of electronic documents, higher requirements are put forward for the automatic retrieval and layout analysis of documents.

[0003] The existing technical solutions tend to use NLP technology to analyze the associations between paragraphs in the document layout, but lack the realization of scenarios where the title diagram spans multiple pages. For example, in the model VGT, through the grid Transformer model, with the help of network linear Projection, the document content data of the image is converted into a serialized representation that can be input to the Transformer encoder, and a model for pre-training with 2D Token-level and segment-level semantic understanding is proposed to solve the association relationship between paragraphs, but it is not applicable to the function of the title diagram spanning multiple pages and being associated. Therefore, how to effectively analyze the layout of the test paper to achieve the layout effect of the title diagram spanning multiple pages and being associated has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] Based on this, it is necessary to provide a layout analysis method, apparatus, device, and medium for the above technical problems, which can effectively analyze the layout of the test paper, so as to better achieve the layout effect of the title diagram spanning multiple pages and being associated.

[0005] In the first aspect of the embodiments of the present application, a layout analysis method is provided. The layout analysis method includes:

[0006] Preprocess the image to be processed to obtain an initial image after preprocessing;

[0007] Input the initial image into a layout analysis model based on a dual-stream network and a graph neural network to obtain a layout recognition result that matches between the title and the title diagram. Among them, the graph neural network includes multiple convolutional neural network structures and a title diagram cross-page association network structure, and each convolutional neural network structure has different receptive fields from large to small in sequence.

[0008] In the second aspect of the embodiments of the present application, a layout analysis apparatus is provided. The layout analysis apparatus includes:

[0009] A processing module, configured to preprocess the image to be processed to obtain an initial image after preprocessing;

[0010] A module is obtained, which is used to input the initial image into a layout analysis model based on a dual-stream network and a graph neural network to obtain a layout recognition result that matches the title and the title image, wherein the graph neural network includes multiple convolutional neural network structures and a title-image cross-page association network structure, and each of the convolutional neural network structures has different receptive fields from large to small.

[0011] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the layout analysis method as described in the first aspect when executing the computer program.

[0012] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the layout analysis method as described in the first aspect is implemented.

[0013] In summary, the present invention provides a layout analysis method, device, equipment and medium, which preprocesses the image to be processed to obtain a preprocessed initial image, and inputs the initial image into a layout analysis model based on a dual-stream network and a graph neural network to obtain a layout recognition result that matches the question and the question image, wherein the graph neural network includes multiple convolutional neural network structures and a question-image cross-page association network structure, and each convolutional neural network structure has different receptive fields from large to small. Compared with the prior art, the present application can effectively analyze the test paper layout by inputting the initial image into a layout analysis model based on a dual-stream network and a graph neural network to obtain a layout recognition result that matches the question and the question image, and through the question-image cross-page association network, the question image and the question are cross-page associated, thereby better achieving the layout layout effect of the question-image cross-page association, and improving the accuracy and efficiency of the layout analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0015] Figure 1 It is a schematic diagram of an application environment of a layout analysis method provided by an embodiment of the present invention;

[0016] Figure 2 It is a flow chart of a layout analysis method provided by one embodiment of the present invention;

[0017] Figure 3 It is a schematic structural diagram of a layout analysis device provided by an embodiment of the present invention;

[0018] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Specific embodiments

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] It should be understood that when used in the specification and appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0021] It should also be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0022] As used in the specification and appended claims of the present invention, the term "if" can be interpreted as "when...", "once" or "in response to determining" according to the context. Similarly, the phrase "if it is determined" or "if [the described condition or event] is compared" can be interpreted as meaning "once it is determined" or "in response to determining" or "once [the described condition or event] is compared" or "in response to comparing [the described condition or event]" according to the context.

[0023] In addition, in the description of the specification and appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0024] Reference to "one embodiment" or "some embodiments" etc. described in the specification of the present invention means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear at different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0025] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0026] In order to illustrate the technical solution of the present invention, the following specific embodiments are used for illustration.

[0027] A layout analysis method provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 to solve common defects in images, thereby improving the accuracy of layout analysis to complete the image of the cross-page association of the test questions and their corresponding figures. Among them, the scanning device communicates with the server through the network. The scanning device is used to obtain the image to be processed and send the image to be processed to the server. The server can use the image to be processed as a training sample image in a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result of the matching between the questions and their corresponding figures. Among them, the scanning device can be but is not limited to various cameras, sensors, laser scanners, scanning measuring instruments, etc. The server can be implemented by an industrial computer, an independent server, or a server cluster composed of multiple servers.

[0028] See Figure 2 , which is a schematic flowchart of a layout analysis method provided by an embodiment of the present invention. The above layout analysis method can be applied to the server in Figure 1 such as Figure 2 shown, and the layout analysis method can be implemented through the following steps.

[0029] S201: Preprocess the image to be processed to obtain an initial image after preprocessing.

[0030] In the embodiments of the present disclosure, the image to be processed can be any image that needs to be subjected to layout analysis, such as an image of a test paper written by a student. The image to be processed can be a photo obtained by photographing teaching materials (such as books, test papers, and workbooks, etc.) through the camera of a smart phone, the camera of a tablet computer, the lens of a digital camera, etc., and can also include an image obtained by scanning teaching materials (such as books, test papers, and workbooks, etc.) through an instrument with a scanning function such as a printer. It should be noted that although the embodiments of the present disclosure take the image to be processed as an example of a test paper image, it should not be regarded as a limitation of the present disclosure.

[0031] As an optional embodiment, the preprocessing of the image to be processed includes at least one random combination: random rotation, random scaling, random cropping, random flipping, random contrast / brightness enhancement, random RGB-gray-RGB color space conversion, random addition of different Gaussian or salt-and-pepper noises, image normalization, and Gaussian bilateral filtering.

[0032] Specifically, random rotation refers to rotating the image within a certain range, which can enhance the diversity of the image, help the model better learn features in various directions, and the rotation angle can be randomly selected from the specified range; random scaling is an operation of enlarging or reducing the image. By changing the size of the image, the model can be more robust to objects of different sizes, and usually scaling is performed within the set ratio range; random cropping is randomly selecting a region from the image for cropping, which can change the perspective of the image and also improve the model's ability to process different scenes and partial information; random flipping refers to randomly flipping the image horizontally or vertically, which can help the model learn and recognize different image layouts, especially in the case of symmetric objects; randomly enhancing the contrast / brightness means randomly adjusting the contrast or brightness of the image, so that the visual effect of the image changes, which can make the model more robust to situations such as light changes; randomly converting the RGB - grayscale - RGB color space can randomly convert the image into a grayscale image or keep it as a color image. This kind of change can not only introduce color diversity, but also help the model's adaptability in low - light or high - contrast situations; randomly adding different Gaussian or salt - and - pepper noises can add Gaussian noise (transparent random noise) or salt - and - pepper noise (random black and white dots) to the image, which can help the model better process noise and thus improve the robustness of the model; image normalization is to adjust the pixel values of the image to a specific range (such as 0 to 1 or - 1 to 1), which helps to improve the stability of the optimization process and make the training process more efficient; Gaussian bilateral filtering is a smoothing operation that removes noise while retaining edge details, taking into account spatial distance and pixel value differences, making neighboring pixels more influential during smoothing, thus making the edges of the image clearer. In this application, by performing at least one randomly combined pre - processing on the image to be processed, the initial image after pre - processing is obtained, so that when the initial image is subsequently input into the layout analysis model based on the two - stream network and the graph neural network, the layout recognition result of the matching between the title and the figure can be obtained quickly and accurately.

[0033] It should be noted that specifically, the pre - processing of the image to be processed, including at least one randomly combined item, can be determined according to the actual situation, and this application does not make any limitations on this.

[0034] In the embodiments of this application, by performing pre - processing on the image to be processed to obtain the initial image after pre - processing, the image to be processed can be quickly converted into an initial image that is more suitable for analysis and processing, making the processed image clearer, which helps subsequent analysis and processing, thereby improving the efficiency and effect of the entire image - processing process.

[0035] S202: Input the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result that matches between the title and the illustration. Among them, the graph neural network includes multiple convolutional neural network structures and an illustration cross-page association network structure, and each of the convolutional neural network structures has different receptive fields from large to small in sequence.

[0036] In the embodiments of the present application, the two-stream network is a network structure used for processing videos and dynamic scene analysis. The graph neural network is a deep learning model that can directly process graph-structured data. The graph-structured data (such as social networks, traffic networks, molecular structures, etc.) consists of nodes (elements in the graph) and edges (relationships between nodes). Since the existing technical solutions tend to use NLP technology to analyze the associations between paragraphs in document layouts but lack the realization of the illustration cross-page scenario, the present application adds an illustration cross-page association network structure to the graph neural network. Given the aggregated feature FM = {FM2, FM3, FM4, FM5}, and then connect a target detection or instance segmentation model, such as the Mask-RCNN model, which can be used to generate candidate components in the document (not just using the RPN), thereby refining the prediction result. The illustration cross-page association network structure updates the hidden representation of each node by participating in the neighbors of each node, that is, by updating the node features, we can predict its precise label and position coordinates. Each node uses a representation that contains position information and depth features, and finally combines the two types of nodes to represent the constructed new node features. By inputting the initial image into a layout analysis model based on a two-stream network and a graph neural network, a layout recognition result that matches between the title and the illustration is obtained.

[0037] It should be noted that the receptive field is the area in the input space that affects a specific unit of the network. This input area can be the input of the network or the output of other units in the network, and no limitation is made here.

[0038] As an alternative embodiment, inputting the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result that matches between the title and the illustration includes:

[0039] Extract features from the initial image through the two-stream network in the image segmentation model to obtain semantic features and visual features;

[0040] Perform feature fusion on the semantic feature map and the visual feature map based on the attention mechanism to obtain a fused feature map;

[0041] Use the graph neural network to encode the fused feature map into a multi-scenario element relationship vector corresponding to the fused feature map;

[0042] According to the multi-scenario element relationship vector, perform layout analysis on the image to be processed to obtain an identification result of the matching between the title and the illustration.

[0043] Specifically, the dual-stream network (dual-stream vision grid transformer based on the VGT model) in the image segmentation model is used to extract visual features and semantic features respectively. These features include modality-specific visual features and language features, which are used for subsequent fusion and relationship modeling, providing a basis for subsequent layout analysis. Based on the attention mechanism, the extracted visual features and semantic features are adaptively fused to obtain a fused feature map, making full use of the complementary information between the two modalities. This step may involve complex fusion strategies, and fusion can also be performed by means of feature weighting, etc. The present application does not make any limitations in this regard. A relationship module based on a graph neural network (GNN) is introduced to model the relationships between components, that is, the graph neural network is used to encode the fused feature map into a multi-scenario element relationship vector corresponding to the fused feature map, thereby providing a useful information representation for subsequent tasks. It should be noted that the graph neural network can learn and reason about the nodes and edges in the multi-scenario element interconnection graph, and is used to make predictions based on the feature representation extracted from the fused feature map. In addition, the graph neural network updates the representation of each element node by iteratively aggregating the neighbor information of the element nodes according to the connection relationship between the element nodes in the fused feature map. The multi-scenario element relationship vector is used to characterize the features of each element node and the connection relationship between the element nodes. Since the multi-scenario element interconnection graph integrates the interaction relationships and interaction actions of multiple scene objects, the multi-scenario element relationship vector also integrates the information of multiple scenes. Therefore, making task predictions based on the multi-scenario element relationship vector can improve the efficiency and accuracy of task predictions. Especially for the task of cross-page association of illustrations, by participating in the neighbors of each node to update the hidden representation of each node, accurate labels and position coordinates can be predicted. Define a specific cross-page association network for illustrations (such as DeConw with Relation Language Modeling). Based on the results of GNN relationship modeling, according to the multi-scenario element relationship vector, the cross-page association network for illustrations is used to perform layout analysis on the image to be processed, perform cross-page association on the illustration and the title, and achieve precise matching between the illustration and the title by updating the hidden association representation of the node, such as sectional association of question numbers, and output the final layout recognition result, including the result of cross-page association of illustrations, which may include the precise positions, category labels of document components, and the association relationships between them.

[0044] Through the cross-page association network of the title figure in this application, the neighbor information of each node can be effectively aggregated, and the hidden association representation of each node can be updated. By stacking multiple graph convolutional layers, more and more complex feature representations can be learned to adapt to more complex task requirements. Through cross-page association and merging processing, the problems of information fragmentation and incompleteness caused by page separation can be overcome, thereby achieving a good cross-page association effect of the title figure and improving the accuracy and integrity of information extraction.

[0045] As an alternative embodiment, according to the multi-scenario element relationship vector, perform layout analysis on the to-be-processed image to obtain a recognition result of the matching between the title and the title figure, including:

[0046] According to the multi-scenario element relationship vector, perform layout analysis on each page of the to-be-processed image to determine a content list and a directory list, where at least a part of the paragraphs in the content list are consistent with at least a part of the titles in the directory list;

[0047] Perform random occlusion and prediction processing on the to-be-processed image in sequence to determine the positioning information and type information of a plurality of sub-region images included in the to-be-processed image;

[0048] According to the positioning information and type information of the plurality of sub-region images, traverse the directory list layer by layer, and perform text matching on the title indicated by the current node being traversed and the titles indicated by its adjacent nodes at the same layer, respectively, with at least a part of the paragraphs in the content list to determine two matching paragraphs;

[0049] Stitch all the paragraphs in the content list located between the two matching paragraphs, and use the stitching result as the associated text block of the current node to obtain a recognition result of the matching between the title and the title figure.

[0050] Specifically, for a to-be-processed image containing multiple pages, the multiple pages usually have a preset page sequence. According to the multi-scenario element relationship vector, page layout analysis is performed on each page of the to-be-processed image to determine a content list and a table of contents list. Among them, at least a part of the paragraphs in the content list are consistent with at least a part of the topics in the table of contents list. By sequentially performing random occlusion and prediction processing on the to-be-processed image, the positioning information and type information of multiple sub-region images included in the to-be-processed image are determined. According to the positioning information and type information of the multiple sub-region images, the table of contents list is traversed layer by layer. For the topic indicated by the current node being traversed and the topics indicated by its adjacent nodes at the same layer, text similarity matching is respectively performed with at least a part of the paragraphs in the content list to determine two matching paragraphs. All the paragraphs between the two matching paragraphs in the content list are spliced, and the splicing result is used as the associated text block of the current node to obtain the recognition result of the matching between the topic and the figure. Compared with performing matching with all the paragraphs in the content list each time, it helps to greatly narrow the scope of matching and improve the efficiency of cross-page association between the figure and the topic.

[0051] As an alternative embodiment, sequentially performing random occlusion and prediction processing on the to-be-processed image to determine the positioning information and type information of multiple sub-region images included in the to-be-processed image includes:

[0052] Selecting N grid data in the to-be-processed image for occlusion processing to obtain the occluded grid data;

[0053] Performing center point prediction and surrounding node prediction on the occluded grid data to obtain center point prediction data and surrounding node prediction data;

[0054] Performing category judgment on the center point prediction data and the surrounding node prediction data;

[0055] According to the judgment result, determining the positioning information and type information of multiple sub-region images included in the to-be-processed image.

[0056] Specifically, N grid data in the image to be processed are selected for occlusion processing to obtain the occluded grid data, which can enable the unoccluded area to better learn the connection of the occluded area. As a result, the unoccluded area not only has the connection of the data and position in its own area but also can have the connection of the occluded area, thereby increasing the generalization ability of the subsequent model. For example, when a picture with a partially blurred area or black dots is input, the model can also well learn the relationship of the blurred area or black dot area. Center point prediction and surrounding node prediction are performed on the occluded grid data. First, the category of the center point of each occluded grid is predicted, and the local maximum value is searched for each occluded grid, and the non-maximum values are suppressed, only the output of the maximum value is retained. Then, by setting a threshold, it is judged whether the category of the predicted occluded area grid is the same as that of its surrounding grids. If they are the same, the type information of the multiple sub-region images included in the image to be processed is determined. By statistically analyzing the boundary information, starting from the position of the corresponding center point, combined with the predicted category and node information, the positioning information of each sub-region image is determined. For example, each region can be represented by a bounding box. By using the information of the center point and the surrounding nodes, the boundary and category of each object in the image can be accurately obtained, improving the accuracy and efficiency of image processing and analysis, and also providing important support for the applications of machine learning and deep learning, making the overall process more efficient.

[0057] As an alternative embodiment, the layout analysis model is trained in the following manner:

[0058] Obtain a training sample data set;

[0059] Input the sample images in the training sample data set into a two-stream network to obtain the fused features in the sample images;

[0060] Input the fused features into a graph neural network to obtain aggregated neighbor features;

[0061] Calculate the loss function of the aggregated neighbor features in the sample images;

[0062] Train the layout analysis model based on the loss function until the loss value output by the trained layout analysis model is less than a preset loss threshold.

[0063] Specifically, several sample images can be obtained from images publicly available online or collected offline, and the obtained sample images include but are not limited to educational scene documents, such as test papers, and the obtained sample images are processed to obtain a training sample data set. The training sample data set can be composed of multiple sample subsets, and a sample subset includes a sample image. It should be noted that in the embodiment of the present disclosure, the size of the sample image is set in advance to facilitate subsequent processing. Since there are differences in the sizes of different sample images, in order to facilitate model training, the collected sample images can be resized, and sample images of different sizes can be corrected to a uniform size, such as uniformly scaling the sample images to a size of 512*512. Afterwards, the sample images in the training sample data set are input into the two-stream network to obtain the fusion features in the sample images, and the fusion features are input into the graph neural network to obtain the aggregated neighbor features. Since in the graph neural network, the hidden representation of each node is updated by aggregating the information of its neighbor nodes, that is, according to the neighbor node features of each node, aggregation is performed, and then the hidden representation of itself is updated by combining its own features and the aggregated neighbor information. Then, the loss function of the aggregated neighbor features in the sample image is calculated, and the network parameters of the layout analysis model are updated according to the loss function to perform iterative training until the loss value of the layout analysis model to be trained is less than or equal to the preset loss threshold, so as to obtain the target layout analysis model. It should be noted that the preset loss threshold can be set according to the experience of the technician, and this embodiment does not limit the value of the preset loss threshold.

[0064] It should be noted that model training is an iterative process. By calculating the loss function value after each iteration and continuously adjusting the model's network parameters for iterative training when the loss function value is greater than the preset value, the model converges to obtain a trained model until the overall loss function value of the model is less than or equal to the preset value, or the overall loss function value of the model no longer changes or changes slowly.

[0065] In the disclosed embodiment, a sample image in a training sample data set is input into a dual-stream network to obtain fused features in the sample image, and the fused features are input into a graph neural network to obtain aggregated neighbor features. Then, a layout analysis model is trained according to a loss function of the aggregated neighbor features in the sample image, until the loss value output by the trained layout analysis model is less than a preset loss threshold. Thus, the trained layout analysis model can effectively analyze the layout of the test paper, thereby better achieving the layout effect of cross-page association between questions and images, and improving the clarity of the image in visual perception.

[0066] In summary, the present invention provides a layout analysis method, apparatus, device, and medium. By preprocessing the image to be processed, an initial image after preprocessing is obtained, and the initial image is input into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result that matches between the title and the figure. Among them, the graph neural network includes multiple convolutional neural network structures and a figure cross-page association network structure, and each convolutional neural network structure has different receptive fields from large to small. Compared with the prior art, the present application inputs the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result that matches between the title and the figure, which can effectively analyze the layout of the test paper. Through the figure cross-page association network, the figure and the title are cross-page associated, so as to better achieve the layout effect of figure cross-page association, improving the accuracy and efficiency of layout analysis.

[0067] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the layout analysis apparatus provided by an embodiment of the present invention. In this embodiment, each unit included in the terminal is used to execute Figure 3 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figure 2 and Figure 2 the relevant descriptions in the corresponding embodiment. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 3 , the layout analysis apparatus 30 includes: a processing module 31 and a obtaining module 32.

[0068] The processing module 31 is used to preprocess the image to be processed to obtain an initial image after preprocessing;

[0069] The obtaining module 32 is used to input the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result that matches between the title and the figure. Among them, the graph neural network includes multiple convolutional neural network structures and a figure cross-page association network structure, and each of the convolutional neural network structures has different receptive fields from large to small.

[0070] Optionally, the above-mentioned processing module 31 is specifically used for:

[0071] Preprocessing the image to be processed includes at least one random combination: random rotation, random scaling, random cropping, random flipping, random contrast / brightness enhancement, random RGB-gray-RGB color space conversion, random addition of different Gaussian or salt-and-pepper noises, image normalization, and Gaussian bilateral filtering.

[0072] Optionally, the above-mentioned obtaining module 32 is specifically used for:

[0073] Extracting features from the initial image through a two-stream network in the image segmentation model to obtain semantic features and visual features;

[0074] Performing feature fusion on the semantic feature map and the visual feature map based on an attention mechanism to obtain a fused feature map;

[0075] Using a graph neural network to encode the fused feature map into a multi-scene element relationship vector corresponding to the fused feature map;

[0076] According to the multi-scene element relationship vector, the layout analysis of the image to be processed is performed to obtain a recognition result of the match between the title and the title image.

[0077] Optionally, the obtaining module 32 is further used for:

[0078] Performing layout analysis on each page of the to-be-processed image according to the multi-scene element relationship vector to determine a content list and a directory list, wherein at least a portion of the paragraphs in the content list is consistent with at least a portion of the titles in the directory list;

[0079] Performing random occlusion and prediction processing on the image to be processed in sequence to determine positioning information and type information of a plurality of sub-region images included in the image to be processed;

[0080] According to the location information and type information of the multiple sub-region images, the directory list is traversed layer by layer, and the title indicated by the traversed current node and the title indicated by the adjacent node in the same layer are respectively matched with at least a part of the paragraphs in the content list to determine two matching paragraphs;

[0081] All paragraphs between the two matching paragraphs in the content list are spliced, and the splicing result is used as the associated text block of the current node to obtain a recognition result of the match between the title and the title image.

[0082] Optionally, the obtaining module 32 is further used for:

[0083] Selecting N grid data in the image to be processed to perform occlusion processing to obtain occluded grid data;

[0084] Perform center point prediction and surrounding node prediction on the blocked grid data to obtain center point prediction data and surrounding node prediction data;

[0085] Make category judgment on the prediction data of the center point and the prediction data of the surrounding nodes;

[0086] According to the judgment result, the positioning information and type information of the multiple sub-region images included in the image to be processed are determined.

[0087] Optionally, the above-mentioned obtaining module 32 is further configured to:

[0088] Obtain a training sample data set;

[0089] Input the sample images in the training sample data set into a two-stream network to obtain the fused features in the sample images;

[0090] Input the fused features into a graph neural network to obtain aggregated neighbor features;

[0091] Calculate the loss function of the aggregated neighbor features in the sample images;

[0092] Train the layout analysis model based on the loss function until the loss value output by the trained layout analysis model is less than a preset loss threshold.

[0093] It should be noted that for the information interaction, execution process, etc. between the above units, since they are based on the same concept as the method embodiment of the present invention, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0094] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps in the method embodiment of the above layout analysis are implemented.

[0095] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 this is only an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.

[0096] In one embodiment, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by the processor in the computer device, the computer device can execute the steps of any embodiment of a layout analysis disclosed in the present invention, which will not be repeated here. The computer-readable storage medium may be non-volatile or volatile.

[0097] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0098] The memory includes a readable storage medium, internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation device and the running of computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operation device, application programs, a boot loader, data, and other programs, such as the program code of computer programs. The memory may also be used to temporarily store the data that has been output or will be output.

[0099] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0100] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0101] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A layout analysis method, characterized in that, including: preprocessing the image to be processed to obtain an initial preprocessed image; inputting the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result of the matching between the title and the figure, wherein the graph neural network includes a plurality of convolutional neural network structures and a figure cross-page association network structure, and each of the convolutional neural network structures has different receptive fields from large to small in sequence; wherein, the inputting the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result of the matching between the title and the figure includes: extracting features from the initial image through the two-stream network in the image segmentation model to obtain semantic features and visual features; performing feature fusion on the semantic feature map and the visual feature map based on an attention mechanism to obtain a fused feature map; encoding the fused feature map into a multi-scene element relationship vector corresponding to the fused feature map by using a graph neural network; performing layout analysis on the image to be processed according to the multi-scene element relationship vector to obtain a recognition result of the matching between the title and the figure; the performing layout analysis on the image to be processed according to the multi-scene element relationship vector to obtain a recognition result of the matching between the title and the figure includes: performing layout analysis on each page of the image to be processed according to the multi-scene element relationship vector to determine a content list and a table of contents list, wherein at least a part of the paragraphs in the content list are consistent with at least a part of the titles in the table of contents list; performing random occlusion and prediction processing on the image to be processed in sequence to determine the positioning information and type information of a plurality of sub-region images included in the image to be processed; traversing the table of contents list layer by layer according to the positioning information and type information of the plurality of sub-region images, and respectively performing text matching on the title indicated by the current node traversed and the titles indicated by the adjacent nodes in the same layer with at least a part of the paragraphs in the content list to determine two matching paragraphs; concatenating all the paragraphs between the two matching paragraphs in the content list, and taking the concatenation result as the associated text block of the current node to obtain a recognition result of the matching between the title and the figure.

2. The layout analysis method according to claim 1, wherein the performing random occlusion and prediction processing on the image to be processed in sequence to determine the positioning information and type information of a plurality of sub-region images included in the image to be processed includes: selecting N grid data in the image to be processed for occlusion processing to obtain occluded grid data; performing center point prediction and surrounding node prediction on the occluded grid data to obtain center point prediction data and surrounding node prediction data; performing category judgment on the center point prediction data and the surrounding node prediction data; determining the positioning information and type information of a plurality of sub-region images included in the image to be processed according to the judgment result.

3. The layout analysis method according to claim 1, wherein The layout analysis model is trained in the following manner: obtaining a training sample data set; inputting the sample images in the training sample data set into the two-stream network to obtain the fused features in the sample images; Input the fusion features into a graph neural network to obtain aggregated neighbor features; Calculate the loss function of the aggregated neighbor features in the sample image; Train the layout analysis model based on the loss function until the loss value output by the trained layout analysis model is less than a preset loss threshold.

4. The layout analysis method according to claim 1, wherein The preprocessing of the image to be processed includes at least one random combination: random rotation, random scaling, random cropping, random flipping, random contrast / brightness enhancement, random RGB-gray-RGB color space conversion, random addition of different Gaussian or salt-and-pepper noises, image normalization, and Gaussian bilateral filtering.

5. A layout analysis device, characterized in that, It includes: A processing module for preprocessing the image to be processed to obtain an initial preprocessed image; An obtaining module for inputting the initial image into a layout analysis model based on a two-stream network and a graph neural network to obtain a layout recognition result of the matching between the title and the figure, wherein the graph neural network includes a plurality of convolutional neural network structures and a figure cross-page association network structure, and each of the convolutional neural network structures has different receptive fields from large to small in sequence; Wherein, the obtaining module is further configured to: Extract features from the initial image through the two-stream network in the image segmentation model to obtain semantic features and visual features; Perform feature fusion on the semantic feature map and the visual feature map based on the attention mechanism to obtain a fusion feature map; Use the graph neural network to encode the fusion feature map into a multi-scene element relationship vector corresponding to the fusion feature map; Perform layout analysis on the image to be processed according to the multi-scene element relationship vector to obtain a recognition result of the matching between the title and the figure; The performing layout analysis on the image to be processed according to the multi-scene element relationship vector to obtain a recognition result of the matching between the title and the figure includes: Perform layout analysis on each page of the image to be processed according to the multi-scene element relationship vector to determine a content list and a table of contents list, wherein at least a part of the paragraphs in the content list are consistent with at least a part of the titles in the table of contents list; Perform random occlusion and prediction processing on the image to be processed in sequence to determine the positioning information and type information of a plurality of sub-region images included in the image to be processed; According to the positioning information and type information of the plurality of sub-region images, traverse the table of contents list layer by layer, and perform text matching on the title indicated by the current node traversed and the titles indicated by the adjacent nodes in the same layer with at least a part of the paragraphs in the content list respectively to determine two matching paragraphs; Concatenate all the paragraphs between the two matching paragraphs in the content list, and use the concatenation result as the associated text block of the current node to obtain a recognition result of the matching between the title and the figure.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the layout analysis method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the layout analysis method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • PDF document cross-page table merging method and device, electronic equipment and storage medium

    CN112380825A

  • Document layout analysis method, model training method and device and equipment

    CN113361247A