Sorting Method, Device, Electronic Device and Storage Medium for Document Picture Content

By extracting and integrating the image and position characteristics of document elements and determining their sorting number, the problem of restoring paper document layout to electronic documents is solved, and the accurate sorting and reading experience of electronic documents is achieved.

CN115114468BActive Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210351128.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-02
Publication Date
2025-07-29
Estimated Expiration
2042-04-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively restore the layout content of paper documents, resulting in poor reading experience of electronic documents.

Method used

By extracting the image features and position features of document elements, fusing them into fusion features, determining the positional relationship between document elements, and then determining the sorting number to achieve accurate sorting of document elements.

Benefits of technology

It realizes accurate sorting of document pictures in different layout formats, and generates editable electronic documents that are the same as paper document layout, improving the reading experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114468B_ABST
    Figure CN115114468B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, an apparatus, an electronic device, and a storage medium for sorting document picture content, which can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. The method includes: fusing the image features and position features of multiple document elements to obtain the fusion features of the multiple document elements; determining the positional relationship between any two of the multiple document elements based on the fusion features of the multiple document elements; and determining the sorting serial numbers of the multiple document elements based on the positional relationship between any two of the multiple document elements. The method of the present disclosure can obtain a better sorting effect for document pictures with different layout formats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to a method, apparatus, electronic device, and storage medium for sorting document picture content. Background Art

[0002] With the development of artificial intelligence technology and image processing technology, paper documents have gradually been replaced by electronic documents due to their disadvantages such as inconvenience in carrying and difficulty in storage. In actual application scenarios, by photographing paper documents, the paper documents are converted into document pictures, and then the content of the document pictures is sorted, and finally the paper documents are converted into electronic documents. In order to maximize the restoration of the layout content of paper documents and improve the user's reading experience, how to sort the content of document pictures has become a concern for those skilled in the art. Summary of the Invention

[0003] Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for sorting document picture content, which can determine relatively accurate sorting numbers for each document element in a document picture, so as to obtain better sorting effects for document pictures with different layout formats. The technical solutions are as follows:

[0004] In a first aspect, a method for sorting document picture content is provided. The method includes:

[0005] Extract the image features and position features of multiple document elements from the document picture to be sorted;

[0006] Fuse the image features and position features of each document element among the multiple document elements to obtain the fused features of the multiple document elements. The fused features are features representing the layout format of the document picture and the correlation between the document elements;

[0007] Based on the fused features of each document element among the multiple document elements, determine the positional relationship between any two document elements among the multiple document elements;

[0008] Based on the positional relationship between any two document elements among the multiple document elements, determine the sorting numbers of the multiple document elements.

[0009] In another embodiment of the present disclosure, before extracting the image features and position features of multiple document elements from the document picture to be sorted, the method further includes:

[0010] Identify the document picture to be sorted to obtain the document detection frames of the multiple document elements;

[0011] Extract the image features and position features of multiple document elements from the document detection frames of the multiple document elements.

[0012] In another embodiment of the present disclosure, the recognition of the document picture to be sorted to obtain the document detection frames of the multiple document elements includes:

[0013] Invoking a document element recognition model to recognize the document picture to obtain the document detection frames of the multiple document elements, where the document element recognition model is used to recognize the document detection frames of document elements in any document picture.

[0014] In another embodiment of the present disclosure, the training process of the document element recognition model is as follows:

[0015] Obtain a plurality of document picture samples, where the document picture samples are labeled with document detection frames;

[0016] Based on the plurality of document picture samples, train an initial document element recognition model to obtain the document element recognition model.

[0017] In a second aspect, a device for sorting the content of a document picture is provided, and the device includes:

[0018] An extraction module, configured to extract the image features and position features of multiple document elements from a document picture to be sorted;

[0019] A fusion module, configured to fuse the image features and position features of each document element among the multiple document elements to obtain the fusion features of each document element among the multiple document elements, where the fusion features are features characterizing the layout format of the document picture and the correlation between the document elements;

[0020] A determination module, configured to determine the positional relationship between any two document elements among the multiple document elements based on the fusion features of each document element among the multiple document elements;

[0021] The determination module is further configured to determine the sorting sequence numbers of the multiple document elements based on the positional relationship between any two document elements among the multiple document elements.

[0022] In a third aspect, an electronic device is provided, where the electronic device includes a processor and a memory, and at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to implement the method for sorting the content of a document picture as described in the first aspect.

[0023] In a fourth aspect, a computer-readable storage medium is provided, where at least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to implement the method for sorting the content of a document picture as described in the first aspect.

[0024] In a fifth aspect, a computer program product is provided. The computer program product includes computer program code stored in a computer-readable storage medium. A processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code to cause the electronic device to execute the method for sorting the content of a document picture as described in the first aspect.

[0025] The beneficial effects brought by the technical solutions provided in the embodiments of the present disclosure are as follows:

[0026] By fusing the image features and position features of each document element, a fused feature of each document element is obtained. This fused feature takes into account both the content and position characteristics of the document element. Through the fused features of each document element, the layout format of the document picture can be known. Considering that there will be a certain correlation in both content and position between the document elements located before and after the document picture, taking into account the layout format and the correlation of the document elements, based on the fused features of each document element, the positional relationship between each document element and other document elements in the content of the document picture to be sorted can be determined. Based on the positional relationships between the document elements, a relatively accurate sorting sequence number can be determined for each document element. This method has strong generalization ability and does not depend on the column information of the document picture. For document pictures with regular layout formats and document pictures with irregular layout formats, better sorting effects can be obtained. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0028] Figure 1 is a schematic diagram of the implementation environment involved in a method for sorting the content of a document picture provided by an embodiment of the present disclosure;

[0029] Figure 2 is a flowchart of a method for sorting the content of a document picture provided by an embodiment of the present disclosure;

[0030] Figure 3 is a flowchart of a method for sorting the content of a document picture provided by an embodiment of the present disclosure;

[0031] Figure 4 is a flowchart of a method for sorting the content of a document picture provided by an embodiment of the present disclosure;

[0032] Figure 5It is a schematic diagram of a constructed simulation data provided by an embodiment of the present disclosure;

[0033] Figure 6 It is a schematic diagram of a picture input by a method for sorting document picture content provided by an embodiment of the present disclosure;

[0034] Figure 7 It is a schematic diagram of a process of fusing image features and position features provided by an embodiment of the present disclosure;

[0035] Figure 8 It is a schematic diagram of a process for judging the positional relationship between document elements provided by an embodiment of the present disclosure;

[0036] Figure 9 It is a schematic diagram of a global relationship matrix provided by an embodiment of the present disclosure;

[0037] Figure 10 It is a schematic diagram of a sorting result of document picture content provided by an embodiment of the present disclosure;

[0038] Figure 11 It is a schematic diagram of a sorting result of document picture content provided by an embodiment of the present disclosure;

[0039] Figure 12 It is a schematic diagram of the structure of a device for sorting document picture content provided by an embodiment of the present disclosure;

[0040] Figure 13 It shows a structural block diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0041] To make the objectives, technical solutions and advantages of the present disclosure clearer, the following will further describe the embodiments of the present disclosure in detail with reference to the accompanying drawings.

[0042] It can be understood that the terms "each", "multiple" and "any one" used in the embodiments of the present disclosure, multiple includes two or more, each refers to each one of the corresponding multiple, and any one refers to any one of the corresponding multiple. For example, multiple words include 10 words, and each word refers to each of these 10 words, and any one word refers to any one of the 10 words.

[0043] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0044] Before further explaining the solutions in the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are first explained.

[0045] GNN (Graph Neural Networks): A method for processing graph domain information based on deep learning. It captures the dependencies of the graph through message passing between the nodes of the graph.

[0046] CNN (Convolutional Neural Networks): A class of feedforward neural networks that contain convolutional or related computations and have a deep structure.

[0047] OCR (Optical Character Recognition): The process of extracting text in an image using image algorithms.

[0048] KNN (K-NearestNeighbor): The core idea of the KNN algorithm is that if most of the K nearest samples of a sample in the feature space belong to a certain category, then the sample also belongs to that category and has the characteristics of the samples in that category.

[0049] The technologies involved in the embodiments of the present disclosure are introduced.

[0050] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0051] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, natural language processing technology, and machine learning / deep learning.

[0052] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing image processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies relevant theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0053] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that combines linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, knowledge graph, etc.

[0054] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0055] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool and be used on demand, being flexible and convenient.

[0056] Please refer to Figure 1, which shows the implementation environment involved in the sorting method of document picture content provided by the embodiments of the present disclosure. The implementation environment includes: a terminal 100 and a server 200. The terminal 100 communicates through a network 300, and the network 300 can be a wired network or a wireless network.

[0057] Among them, the terminal 100 can be a device with a photographing or scanning function such as a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a fax machine, a scanner, etc., but is not limited thereto. The server 200 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0058] In some embodiments, the sorting method of document picture content provided by the embodiments of the present disclosure can be implemented by the terminal 100. Specifically, the terminal 100 activates the photographing or scanning function, takes a photo of a paper document to obtain a document picture, and then implements the sorting method of document picture content provided by the embodiments of the present disclosure to determine sorting serial numbers for each document element in the document picture. Furthermore, based on the sorting serial numbers of each document element, the content of each document element is added to the corresponding position of a blank document, and finally an editable electronic document with the same layout content as the paper document is generated.

[0059] In other embodiments, the sorting method of document picture content provided by the embodiments of the present disclosure can be implemented by the server 200. Specifically, the terminal 100 activates the photographing or scanning function, takes a photo of a paper document to obtain a document picture, and then sends the document picture to the server 200. The server receives the document picture sent by the terminal 100, and then implements the sorting method of document picture content provided by the embodiments of the present disclosure to determine sorting serial numbers for each document element in the document picture. Then, based on the sorting serial numbers of each document element, the content of each document element is added to the corresponding position of a blank document, and finally an editable electronic document with the same layout content as the paper document is generated.

[0060] An embodiment of the present disclosure provides a method for sorting document picture content. This method can accurately sort various document elements such as pictures, tables, and text lines, and finally return the sorting information in a predetermined format according to the determined sorting serial numbers. To achieve the above goals, the hardware environment required for the embodiments of the present disclosure is divided into two parts according to model training and model output results. During model training, an electronic device (terminal 100 or server 200) needs to be equipped with a GPU (Graphics Processing Unit, image processor) chip to support GPU parallel computing; when the model outputs results, there are no special hardware requirements for the electronic device.

[0061] The method for sorting document picture content provided by the embodiments of the present disclosure has the sorting ability in a general document recognition scenario, and can rearrange the text content of optical character recognition and other document element content (including pictures, tables, etc.) according to a predetermined reading order. This sorting ability can be used for the layout restoration of document picture content, and can also provide valuable input for higher-order document understanding capabilities (such as fine-grained text recognition).

[0062] Figure 2 The sorting process of the document picture content provided by the embodiments of the present disclosure is shown. Refer to Figure 2 , for any document picture to be sorted, extract document elements from the document picture. The document elements include at least one of pictures, tables, text blocks, etc. Then, use the optical character recognition method to recognize the document lines contained in the text blocks, and further extract the image features and position features of multiple document elements such as pictures, tables, and text lines. Then, fuse the image features and position features of each document element to obtain the fusion feature of each document element. Furthermore, based on the fusion feature of each document element, determine the positional relationship between any two document elements among the multiple document elements. Then, based on the positional relationship between any two document elements among each document element, determine the sorting serial numbers of each document element. Finally, based on the sorting serial numbers of each document element, add the content of each document element to the corresponding position of a blank document to obtain an editable electronic document with the same content as the document picture.

[0063] An embodiment of the present disclosure provides a method for sorting document picture content. This method can be executed by an electronic device, and the electronic device is Figure 1 at least one of the terminal 100 or the server 200 in Figure 3 . The embodiments of the present disclosure do not make specific limitations on this. The embodiments of the present disclosure can be applied to various scenarios, including but not limited to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. Refer to

[0064] 301. Extract the image features and position features of multiple document elements from the document picture to be sorted.

[0065] Among them, the document images to be sorted can be document images obtained by photographing paper documents such as books, papers, periodicals, certificates, etc. These document images cannot be edited and are not convenient for users to use. To facilitate user use, the electronic device needs to convert the document images to be sorted into editable electronic documents. After converting the document images into editable electronic texts, it is necessary to restore the content of the document images on a blank document. To restore the content of the document images, it is necessary to fill the content of the document images into the corresponding positions in the original order. Considering that the document images are composed of multiple document elements, these document elements are the smallest units that make up the content of the document images, and their types include pictures, tables, text blocks, headers and footers, dividing lines, etc. When the sorting numbers of the various document elements in the document images are determined, the content of the document images is also sorted. It can be seen that the process of sorting the content of the document images by the electronic device is actually a process of determining the sorting numbers of the various document elements in the content of the document images.

[0066] For any document image to be sorted, after the electronic device identifies multiple document elements from the document image to be sorted, it extracts the image features and position features of each document element, so as to determine the sorting number of each document element based on the image features and position features of each document element in the subsequent steps. Among them, the image features are used to characterize the local image features of the corresponding image of the document element on the document image to be sorted, including texture features, brightness, contrast, etc.; the position features are used to characterize the local position where the corresponding image of the document element is located on the document image to be sorted.

[0067] 302. Fuse the image features and position features of each document element among multiple document elements to obtain the fusion features of the multiple document elements.

[0068] For each document element among multiple document elements, the electronic device fuses the image features and position features of each document element to obtain the fusion feature of each document element. This fusion feature is a feature that characterizes the layout format of the document image and the correlation between document elements.

[0069] 303. Based on the fusion features of each document element among multiple document elements, determine the positional relationship between any two document elements among the multiple document elements.

[0070] The electronic device processes the fusion feature of each document element to obtain the processing result of each document element, and then determines the positional relationship between any two document elements among the multiple document elements based on the processing result of each document element.

[0071] 304. Based on the positional relationship between any two document elements among multiple document elements, determine the sorting numbers of the multiple document elements.

[0072] Based on the positional relationship between any two document elements among multiple document elements, the electronic device can determine the order of each document element relative to other document elements in the document, and thus determine a corresponding sorting serial number for each document element.

[0073] The method provided by the embodiments of the present disclosure fuses the image features and positional features of each document element to obtain the fusion features of each document element. The fusion features take into account both the content and position characteristics of the document elements. Through the fusion features of each document element, the layout format of the document picture can be obtained. Considering that there will be a certain correlation in both content and position between the document elements located before and after the document picture, taking into account the layout format and the correlation of document elements, based on the fusion features of each document element, the positional relationship between each document element and other document elements in the content of the document picture to be sorted can be determined. Based on the positional relationship between the document elements, a relatively accurate sorting serial number can be determined for each document element. This method has strong generalization ability and does not depend on the column information of the document picture. For document pictures with regular layout formats and document pictures with irregular layout formats, better sorting effects can be obtained.

[0074] The embodiments of the present disclosure provide a method for sorting the content of a document picture. This method can be executed by an electronic device, and the electronic device is Figure 1 at least one of the terminal 100 or the server 200 in []. The embodiments of the present disclosure do not make specific limitations in this regard. The embodiments of the present disclosure can be applied to various scenarios, including but not limited to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. Refer to Figure 4 and the method flow provided by the embodiments of the present disclosure includes:

[0075] 401. The electronic device identifies the document picture to be sorted and obtains the document detection frames of multiple document elements.

[0076] Among them, the document detection frame is the position area of the document element on the document picture to be sorted. The shape of the document detection frame can be a regular shape such as a rectangle, square, circle, rhombus, hexagon, etc., or an irregular shape. For the convenience of subsequent sorting of the document elements within the document detection frame, the embodiments of the present disclosure set the shape of the document detection frame to a rectangle. If the shape of the document detection frame of the identified document element is not a rectangle, the document detection frame of other shapes can be converted into a rectangle.

[0077] After obtaining the document image to be sorted, the electronic device calls the document element recognition model to recognize the document image and obtain the document detection frames of multiple document elements. Among them, the document element recognition model is used to recognize the document detection frames of document elements in any document image. The training process of the document element recognition model is as follows:

[0078] 4011. The electronic device obtains multiple document image samples.

[0079] Among them, the document image samples are marked with document detection frames, sorting serial numbers of document elements, etc. Based on the marked document detection frames in the document image samples, the initial document element recognition model is trained so that the trained document element recognition model can learn how to recognize the document detection frames of document elements from the document image. Based on the sorting serial numbers corresponding to each document detection frame, the trained document element recognition model can learn the information related to the layout format of the document image, etc.

[0080] Among them, multiple document image samples can be obtained by manually annotating the document detection frames and sorting serial numbers of the document elements of the document image. However, due to the high cost and low efficiency of manual annotation, the document image samples obtained by this method cannot meet the data coverage requirements in the general document scenario in terms of both quantity and quality. In order to obtain sufficient and high-quality document image samples to improve the accuracy of the trained document element recognition model, the embodiments of the present disclosure will, based on the data construction engine, according to the templates of multiple public data sets, fill the collected open-source document elements into the document according to the real arrangement logic (the filling categories include text, pictures, tables, headers and footers, dividing lines, figure captions and table notes, titles, formulas, etc.), thereby constructing a large number of quasi-real layouts and expanding the quantity of document image samples. During the process of filling the above-mentioned document elements, the sorting serial numbers of the document elements will also be marked in the normal reading order. For example, for a template containing a title, text, picture, table, and figure caption annotation, the server randomly obtains a title, text, picture, table, picture annotation, table annotation, etc. from the open-source documents provided by the Internet, and then, based on the data construction engine, fills the obtained document elements such as the title, text, picture, table, picture annotation, and table annotation into the corresponding positions of the template to obtain Figure 5 the shown document image sample, and then the server adds document detection frames and sorting serial numbers to each document element according to the real reading order of each document element to obtain Figure 6 the shown document image sample.

[0081] 4012. The electronic device trains the initial document element recognition model based on multiple document image samples to obtain the document element recognition model.

[0082] Among them, the initial document original recognition model can be any GNN model, and the present disclosure embodiment does not limit the initial document element recognition model. The electronic device sets initial model parameters for the initial document element recognition model, calls the initial document element recognition model, recognizes multiple document picture samples, obtains the document detection frames of each document element in each document picture sample, and then adjusts the initial model parameters based on the document detection frames marked for each document element and the recognized document detection frames, and finally obtains a trained document original recognition model.

[0083] It should be noted that the types of document elements recognized by the document element recognition model include pictures, tables, text blocks, etc. To achieve more fine-grained text line classification, the present disclosure embodiment will use OCR to recognize the text blocks to obtain the accurate positions of each text line included in the text blocks, so as to generate the document detection frames of each text line.

[0084] Furthermore, after the electronic device obtains the document detection frames of each document element, for the content of each document element, if the type of the document element is a text line, the OCR method can be used for extraction, and if the type of the document element is a non-text line, the semantic segmentation method of ResNet-101 can be used for extraction. Among them, when using the OCR method to extract the content of the text line, the sequence recognition method of CRNN (Convolutional Recurrent Neural Network) is adopted.

[0085] 402. The electronic device extracts the image features and position features of multiple document elements from the document detection frames of multiple document elements.

[0086] For each of multiple document elements, when an electronic device extracts the image features of a document element from the document detection box of the document element, it can input the document detection box of the document element into a CNN model for extracting image features for extraction. The CNN model can be a ResNet-18 backbone network or the like. Since the types of document elements include pictures, tables, text lines, etc., the sizes of the document detection boxes of different document elements are not the same, resulting in different sizes of the feature maps where the image features extracted from the document detection boxes of each document element are located, thus affecting the subsequent judgment of the positional relationship between each document element. Therefore, before extracting the image features and position features from the document detection boxes of each document element, the electronic device will perform size normalization processing on the document detection boxes of each document element so that the document detection boxes of each document element have the same size. When the electronic device performs size normalization processing on the document detection boxes of each document element, it can scale the resolution of the document element to H*W to obtain a normalized document detection box. Among them, H and W are fixed pixel values and can be determined according to the actual processing ability of the electronic device.

[0087] Based on the normalized document detection box, the electronic device inputs the normalized document detection box into the CNN model and outputs image features of a preset dimension. Among them, the preset dimension is related to the network structure of the CNN model used and can be 16 dimensions, 32 dimensions, etc. The embodiments of the present disclosure do not limit the preset dimension. For example, the electronic device inputs the normalized document detection box into the ResNet-18 backbone network and outputs a feature map with a resolution of H / 32*W / 32 and a channel number of 16. For the feature map corresponding to the document detection box, the electronic device uses a preset feature extraction algorithm to extract its relevant regional features to obtain the image features of the document detection box. The preset feature extraction algorithm can be an ROI Align algorithm or the like. Among them, the ROI Align algorithm is a regional feature aggregation algorithm used to solve the problem of regional mismatch caused by two quantizations in the ROI Pooling operation. The principle of the ROI Align algorithm is: using the method of bilinear interpolation to obtain the image values at pixel points with floating-point coordinates, thereby converting the entire feature aggregation process into a continuous operation.

[0088] When the method provided by the embodiments of the present disclosure extracts the position features of the document detection frame, it is different from the coordinate representation method adopted in the related art. Instead, it adopts a vector embedding method to map the 4-dimensional original position features of the normalized document detection frame into a feature vector space of a preset dimension, and obtains position features with the same dimension as the image features. When mapping the 4-dimensional original position features of the document detection frame into position features of a preset dimension, a mapping strategy for converting the original position features into position features of a preset dimension can be set. This mapping strategy can be reflected in a separate feature dimension conversion model, which is used to convert 4-dimensional features into preset dimension features. The feature dimension conversion model can be trained using training data. Based on this feature dimension conversion model, this step can be executed to map the 4-dimensional original position features into position features of a preset dimension. This mapping strategy can also be integrated into the fusion network for fusing the image features and position features of the document detection frame. That is, what is obtained by executing this step is the 4-dimensional original position features. When executing step 403 to fuse the image features and position features of the document detection frame, the fusion network for fusing the image features and position features of the document detection frame will convert the original position features into position features with the same dimension as the image features, and then fuse the converted position features with the image features. By adopting this processing method, data with different dimensions can be placed in the same dimension for learning.

[0089] 403. The electronic device fuses the image features and position features of each document element among multiple document elements to obtain the fusion features of the multiple document elements.

[0090] The embodiments of the present disclosure provide a GNN model based on a depth map. This GNN model can fuse the image features and position features extracted from the document detection frame of the document element to obtain the fusion features of the document element. It should be noted here that in the general document recognition scenario, the semantics between document paragraphs are not obvious. Therefore, the input of the GNN model here is the image features and position features of the document element, and does not include the semantic features of the text within the document detection frame. Since the fusion features obtained by adopting this processing method can represent the layout format of the document picture and the correlation between multiple document elements, this makes the GNN model pay more attention to the inherent layout format of the document element and not overly focus on the semantic correlation between document elements.

[0091] Since different types of document elements have different characteristics, when the electronic device fuses the image features and position features of the document element for different types of document elements, different fusion methods will be adopted:

[0092] First method: When the type of the document element is a non-text line, the electronic device splices the image features and position features of the document element to obtain the fusion features of the document element.

[0093] When the type of the document element is a non-document line such as a figure or a table, the document element itself is an individual entity. At this time, the image features and position features of the document element are directly spliced to obtain the fused features of the document element. For example, the image features of the document element are image features of H / 32*W / 32*16 dimensions, and the position features are position features of H / 32*W / 32*16 dimensions. By splicing the image features and position features together, fused features of H / 32*W / 32*(16 + 16) can be obtained.

[0094] The second method: When the type of the document element is a text line, the electronic device obtains the image features and position features of the text block where the document element is located, and splices the image features of the document element, the position features of the document element, the image features of the text block, and the position features of the text block to obtain the fused features of the document element.

[0095] When the type of the document element is a text line, in order to better improve the sorting effect of the document element and enable the model to learn the group relationship between text lines, the electronic device also introduces the features of the text block where the text line is located into the model. By splicing the image features of the document element, the position features of the document element, the image features of the text block, and the position features of the text block, the fused features of the document element are obtained. The fused features include the features of the text line itself and the features of the text block where the text line is located, so that the common features of the same text block can be better learned. For example, a certain text block includes five text lines, and these five text lines belong to the same text block. The electronic device splices the 16-dimensional image features and 16-dimensional position features of the text block and the 16-dimensional image features and 16-dimensional position features of the corresponding text lines together to obtain the fused features of the text line as H / 32*W / 32*64.

[0096] Figure 7 Fig. shows a schematic structural diagram of a fusion network 710 for fusing the image features and position features of a document element. Refer to Figure 7 , the fusion network 710 downsamples the original coordinates of the document detection box of the document element to obtain the original position features of the document detection box, and then converts the original position features into position features with the same dimension as the image features. Then, the image features of the document element are fused with the position features of the same dimension. If the type of the document element is a non-text line such as a picture or a table, the image features of the document element can be directly fused with the position features of the same-dimension features. If the type of the document element is a text line, the image features of the document element need to be fused with the position features of the same-dimension features of the document element and the image features and position features of the text block where the document element is located, and finally the fused features of the document element are obtained.

[0097] 404. The electronic device determines the positional relationship between any two of the multiple document elements based on the fusion features of each of the multiple document elements.

[0098] When the electronic device determines the positional relationship between any two of the multiple document elements based on the fusion features of each of the multiple document elements, the following method can be adopted:

[0099] 4041. The electronic device determines the multiple global stitching features corresponding to any two of the multiple document elements based on the fusion features of each of the multiple document elements.

[0100] Among them, the global stitching feature is a feature characterizing the positional relationship between two document elements in the document picture to be sorted. For any document element, in order to obtain the performance of this document element globally in the document picture, the electronic device obtains the neighboring document elements of this document element, fuses the features of these neighboring document elements into itself to obtain the global feature of this document element, and then stitches the global features of the two document elements whose relationship needs to be judged, so as to judge the positional relationship between the two document elements based on the obtained global stitching feature.

[0101] When the electronic device obtains the global stitching features of any two of the multiple document elements, the following method can be adopted:

[0102] In the first step, the electronic device obtains multiple neighboring document elements whose distance from this document element is less than a preset distance.

[0103] The electronic device can use the KNN algorithm to obtain the K neighboring document elements with the closest distance to the document element. When obtaining the K neighboring document elements with the closest distance to the document element, the electronic device calculates the document elements whose distance between the fusion feature and the fusion feature of this document element is less than the preset distance according to the fusion features of the multiple document elements, and then uses the document elements that meet the preset distance condition as the neighboring document elements of this document element. The distance mentioned here can be the Euclidean distance, etc., and the preset distance can be set by technicians according to requirements.

[0104] In the second step, the electronic device fuses the fusion features of the multiple neighboring document elements corresponding to the document element with the fusion feature of the document element to obtain the global feature of the document element.

[0105] The electronic device stitches the fusion features of each neighboring document element corresponding to this document element with the fusion feature of this document element to obtain the global feature of this document element.

[0106] In the third step, the electronic device stitches the global features of any two of the multiple document elements to obtain multiple global stitching features.

[0107] For any two document elements among multiple document elements, the electronic device splices the global features of the two document elements to obtain the global splicing feature of the two document elements. By pairwise splicing the global features of each document element among multiple document elements with other document elements except itself, multiple global splicing features are obtained.

[0108] In another embodiment of the present disclosure, to facilitate representing the proximity relationship between a document element and its neighboring document elements, the electronic device takes each document element as a node, and then connects the node corresponding to the document element with the nodes corresponding to its neighboring document elements to obtain a node relationship network. The style of the node relationship network can be seen in Figure 8 .

[0109] 4042. The electronic device invokes a position relationship recognition model to process multiple global splicing features and obtain the position relationship between any two document elements among multiple document elements.

[0110] Among them, the position relationship recognition model is used to determine the position relationship between two document elements based on the global splicing features of the two document elements. The electronic device invokes the position relationship recognition model to process multiple global splicing features, obtains the position relationship value corresponding to each global splicing feature, and then determines the position relationship between any two document elements among multiple document elements according to the position relationship value. Among them, the position relationship value includes 0 or 1. When the position relationship value is 0, it means that the document element spliced earlier in the global splicing feature is ranked behind the document element spliced later; when the position relationship value is 1, it means that the document element spliced earlier in the global splicing feature is ranked in front of the document element spliced later.

[0111] 405. The electronic device determines the sorting serial numbers of multiple document elements based on the position relationship between any two document elements among multiple document elements.

[0112] Based on the positional relationship between any two of multiple document elements, the electronic device determines the priorities of the multiple document elements by counting the number of times each document element appears in front of other document elements, and arranges the multiple document elements in descending order of the number of times. Then, based on the priorities of the multiple document elements, the electronic device determines the sorting numbers of the multiple document elements in descending order of priority. For example, from the document pictures to be sorted, document element A, document element B, document element C, and document element D are identified. Among them, the number of times document element A appears in front of other document elements is 3 times, the number of times document element B appears in front of other document elements is 2 times, the number of times document element C appears in front of other document elements is 1 time, and the number of times document element D appears in front of other document elements is 0 times. Then, in descending order of the number of times, the priority of document element A is determined to be the first level, the priority of document element B is determined to be the second level, the priority of document element C is determined to be the third level, and the priority of document element D is determined to be the fourth level. Then, in descending order of priority, the sorting number of document element A is determined to be 1, the sorting number of document element B is determined to be 2, the sorting number of document element C is determined to be 3, and the sorting number of document element D is determined to be 4.

[0113] In another embodiment of the present disclosure, in order to more intuitively display the positional relationship between each document element, the electronic device also constructs a global relationship matrix based on the positional relationship between each document element. The number of rows and columns of the global relationship matrix is the same as the number of document elements, and each row and each column corresponds to a document element. Each value in the global relationship matrix represents the positional relationship value of the global splicing feature obtained by splicing the document element corresponding to the row as the previous document element and the document element corresponding to the column as the subsequent document element. For example, Figure 9 As shown in the global relationship matrix, the global relationship matrix has eighteen rows and eighteen columns. The first row to the eighteenth row respectively represent document elements A to R. Based on the constructed global relationship matrix, the electronic device traverses each row of the global relationship matrix, adds up the elements with a value of 1 in each row, and uses the sum as the number of times the document element corresponding to the row appears in front of other document elements. The priorities of each document element are determined in descending order of the number of times. Finally, the sorting number of the document element with the highest priority is set to 1, and the document element is sorted at the front of the document. The sorting number of the document element with the second highest priority is set to 2, and the document element is sorted at the second position of the document, and so on, until all document elements are sorted. For example, the global relationship matrix The first row of the global relationship matrix corresponds to document element A, the second row corresponds to document element B, and the third row corresponds to document element C. Add up the elements with a value of 1 in the first row, and the number of times document element A ranks ahead of other document elements is 3 times. Add up the elements with a value of 1 in the second row, and the number of times document element B ranks ahead of other document elements is 2 times. Add up the elements with a value of 1 in the third row, and the number of times document element C ranks ahead of other document elements is 1 time. Determine the priority of each document element in the order of the number of times from more to less. Determine that the priority of document element A is the first level, the priority of document element B is the second level, and the priority of document element C is the third level. Then, in the order of priority from high to low, determine that the sorting number of document element A is 1, the sorting number of document element B is 2, and the sorting number of document element C is 3.

[0114] In another embodiment of the present disclosure, when the document layout is complex or there are many elements, the learning difficulty of the network is relatively large, and local text reverse order errors may occur, resulting in the document detection box that should be sorted in the front being sorted in the back. For example Figure 10 A local text reverse order error occurs between the document detection box numbered 7 and the document detection box numbered 6, resulting in the document detection box that should be sorted in the front being sorted in the back. To solve this problem, based on the general prior knowledge that the order of each text line element falling into the same text block in the document scene is close, the present disclosure embodiment merges each document element belonging to the same text block to obtain multiple text blocks. Then, based on the sorting numbers of each document element in the multiple text blocks, determine the sorting numbers of the multiple text blocks. Then, in the order of the sorting numbers from small to large, re-determine the sorting numbers for the document elements of non-text line type among the multiple document elements and the multiple document blocks. When the electronic device determines the sorting number of the text block based on the sorting numbers of each document element in the text block, it can calculate the median value of the sorting numbers of each document element in the text block, and then use this median value as the sorting number of the text block. Using this processing method ensures that local reverse order errors are smoothed by text block-level information, improving the sorting accuracy. After using this method for processing, if the electronic device is a terminal, directly output all document elements and their corresponding sorting numbers; if the electronic device is a server, provide the sorting information of each document element to the terminal, and the terminal outputs all document elements and their corresponding sorting numbers. The sorting effect displayed on the terminal can be seen in Figure 11 .

[0115] The method provided by the embodiments of the present disclosure fuses the image features and position features of each document element to obtain the fusion features of each document element. The fusion features take into account both the content and position characteristics of the document elements. Through the fusion features of each document element, the layout format of the document picture can be known. Considering that there will be a certain correlation in both the content and position of the document elements located before and after the document picture, taking into account the layout format and the correlation of the document elements, based on the fusion features of each document element, the positional relationship between each document element and other document elements in the content of the document picture to be sorted can be determined. Based on the positional relationship between the document elements, a relatively accurate sorting number can be determined for each document element. This method has strong generalization ability and does not depend on the column information of the document picture. For document pictures with regular layout formats and document pictures with irregular layout formats, better sorting effects can be obtained.

[0116] See Figure 12 , the embodiments of the present disclosure provide a sorting device for the content of a document picture, and the device includes:

[0117] An extraction module 1201, configured to extract the image features and position features of multiple document elements from the document picture to be sorted;

[0118] A fusion module 1202, configured to fuse the image features and position features of each document element among the multiple document elements to obtain the fusion features of each document element among the multiple document elements, and the fusion features are features representing the layout format of the document picture and the correlation between document elements;

[0119] A determination module 1203, configured to determine the positional relationship between any two document elements among the multiple document elements based on the fusion features of each document element among the multiple document elements;

[0120] The determination module 1203 is further configured to determine the sorting numbers of the multiple document elements based on the positional relationship between any two document elements among the multiple document elements.

[0121] In another embodiment of the present disclosure, the device further includes:

[0122] An identification module, configured to identify the document picture to be sorted to obtain the document detection frames of multiple document elements;

[0123] The extraction module is further configured to extract the image features and position features of the multiple document elements from the document detection frames of the multiple document elements.

[0124] In another embodiment of the present disclosure, an identification module is configured to call a document element identification model to identify a document picture, and obtain document detection frames of multiple document elements. The document element identification model is used to identify document detection frames of document elements in any document picture.

[0125] In another embodiment of the present disclosure, the apparatus for training a document element identification model includes:

[0126] An acquisition module, configured to acquire a plurality of document picture samples, and the document picture samples are labeled with document detection frames;

[0127] A training module, configured to train an initial document element identification model based on the plurality of document picture samples to obtain a document element identification model.

[0128] In another embodiment of the present disclosure, a fusion module 1202 is configured to, for any document element, when the type of the document element is a non-text line, splice the image feature and the position feature of the document element to obtain a fusion feature of the document element; when the type of the document element is a text line, obtain the image feature and the position feature of the text block where the document element is located, and splice the image feature of the document element, the position feature of the document element, the image feature of the text block, and the position feature of the text block to obtain a fusion feature of the document element.

[0129] In another embodiment of the present disclosure, a determination module 1203 is configured to, based on the fusion features of each document element among the multiple document elements, determine a plurality of global splicing features corresponding to any two document elements among the multiple document elements. The global splicing feature is a feature representing the positional relationship between the two document elements in the document picture to be sorted; call a positional relationship identification model to process the plurality of global splicing features to obtain the positional relationship between any two document elements among the multiple document elements. The positional relationship identification model is used to determine the positional relationship between two document elements based on the global splicing features of the two document elements.

[0130] In another embodiment of the present disclosure, a determination module 1203 is configured to, for any document element, obtain a plurality of neighbor document elements whose distances from the document element are less than a preset distance; fuse the fusion features of the plurality of neighbor document elements corresponding to the document element with the fusion feature of the document element to obtain a global feature of the document element; splice the global features of any two document elements among the multiple document elements to obtain a plurality of global splicing features.

[0131] In another embodiment of the present disclosure, a determination module 1203 is configured to determine the priorities of multiple document elements by counting the number of times a document element ranks in front of other document elements based on the positional relationship between any two of the multiple document elements; and determine sorting serial numbers for the multiple document elements in descending order of priority.

[0132] In another embodiment of the present disclosure, the apparatus further includes:

[0133] A merging module is configured to merge each document element belonging to the same text block to obtain multiple text blocks;

[0134] The determination module 1203 is further configured to determine the sorting serial numbers of the multiple text blocks based on the sorting serial numbers of each document element within the multiple text blocks;

[0135] The determination module 1203 is further configured to re-determine the sorting serial numbers of the document elements of non-text line type among the multiple document elements and the multiple document blocks in ascending order of the sorting serial numbers.

[0136] In another embodiment of the present disclosure, the determination module 1203 is configured to calculate the median value of the serial numbers of the sorting serial numbers of each document element within the text block; and use the median value as the sorting serial number of the text block.

[0137] In another embodiment of the present disclosure, the apparatus further includes:

[0138] An adding module is configured to add the contents of the multiple document elements to corresponding positions of a blank document based on the sorting serial numbers of the multiple document elements to obtain an editable electronic document identical to the content of the document picture.

[0139] In summary, the apparatus provided by the embodiments of the present disclosure fuses the image features and position features of each document element to obtain the fusion feature of each document element. This fusion feature takes into account both the content and position characteristics of the document element. Through the fusion features of each document element, the layout format of the document picture can be known. Considering that there will be a certain correlation in both content and position for the document elements located before and after the document picture, taking into account the layout format and the correlation of the document elements, based on the fusion feature of each document element, the positional relationship between each document element and other document elements in the content of the document picture to be sorted can be determined. Based on the positional relationship between each document element, a relatively accurate sorting serial number can be determined for each document element. This method has strong generalization ability and does not rely on the column information of the document picture. For document pictures with regular layout formats and document pictures with irregular layout formats, better sorting effects can be obtained.

[0140] Figure 13The block diagram of an electronic device 1300 provided by an exemplary embodiment of the present disclosure is shown. Generally, the electronic device 1300 includes: a processor 1301 and a memory 1302.

[0141] The processor 1301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0142] The memory 1302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1302 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1301 to implement the document picture content sorting method provided in the method embodiments of the present disclosure.

[0143] In some embodiments, the electronic device 1300 may further optionally include: a peripheral device interface 1303 and at least one peripheral device. The processor 1301, the memory 1302, and the peripheral device interface 1303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1303 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes: a power supply 1304.

[0144] The peripheral device interface 1303 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0145] The power supply 1304 is used to supply power to each component in the electronic device 1300. The power supply 1304 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 1304 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0146] Those skilled in the art can understand that Figure 13 the structure shown in does not constitute a limitation on the electronic device 1300, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0147] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by the processor of the electronic device 1300 to complete the method for sorting document picture content. Optionally, the storage medium can be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium can be a CD-ROM (Compact Disc Read-Only Memory), ROM, RAM (Random Access Memory), magnetic tape, floppy disk, and optical data storage device, etc.

[0148] The embodiments of the present disclosure provide a computer-readable storage medium. At least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to implement the method for sorting document picture content.

[0149] The embodiments of the present disclosure provide a computer program product. The computer program product includes computer program code. The computer program code is stored in a computer-readable storage medium. The processor of the electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the electronic device executes the method for sorting document picture content.

[0150] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.

[0151] The above are only optional embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for sorting the content of a document picture, characterized in that, The method includes: Invoking a document element recognition model to recognize a document image to be sorted, obtaining document detection frames of multiple document elements, where the document elements are the smallest units that make up the content of the document image, and the document elements include at least one of pictures, tables, text lines, headers / footers, and dividing lines. The document image to be sorted is obtained by photographing a paper document; Performing size normalization processing on the document detection frames of each document element so that the document detection frames of different document elements have the same size; Based on the normalized document detection frames, extracting image features of a preset dimension and original position features of an original dimension of multiple document elements, where the image features include at least one of texture features, brightness, and contrast; Through a fusion network, converting the original dimension of the original position features to the preset dimension to obtain position features of multiple document elements; For any document element, when the type of the document element is a non-text line, splicing the image features and position features of the document element to obtain the fusion features of the document element; when the type of the document element is a text line, obtaining the image features and position features of the text block where the document element is located, and splicing the image features of the document element, the position features of the document element, the image features of the text block, and the position features of the text block to obtain the fusion features of the document element; the fusion features are features that characterize the layout format of the document image and the correlation between document elements; Based on the fusion features of each document element among the multiple document elements, determining the position relationship between any two document elements among the multiple document elements; Based on the position relationship between any two document elements among the multiple document elements, determining the sorting serial numbers of the multiple document elements; Merging the document elements belonging to the same text block to obtain multiple text blocks; For any text block among the multiple text blocks, calculating the median value of the sorting serial numbers of the document elements within the text block; using the median value as the sorting serial number of the text block; In the order from smallest to largest sorting serial number, re-determining the sorting serial numbers for the document elements of the non-text line type among the multiple document elements and the multiple text blocks; Based on the sorting serial numbers of the multiple document elements, adding the contents of the multiple document elements to the corresponding positions of a blank document to obtain an editable electronic document with the same content as the paper document.

2. The method according to claim 1, wherein The determining the position relationship between any two document elements among the multiple document elements based on the fusion features of each document element among the multiple document elements includes: Based on the fusion features of each document element among the multiple document elements, determining multiple global splicing features corresponding to any two document elements among the multiple document elements, where the global splicing features are features that characterize the position relationship between the two document elements in the document image to be sorted; Call the position relationship recognition model to process the multiple global splicing features, and obtain the position relationship between any two document elements among the multiple document elements. The position relationship recognition model is used to determine the position relationship between the two document elements based on the global splicing features of the two document elements.

3. The method according to claim 2, wherein The determining the multiple global splicing features corresponding to any two document elements among the multiple document elements based on the fusion features of each document element among the multiple document elements includes: For any document element, obtain multiple neighbor document elements whose distance from the document element is less than a preset distance; Fuse the fusion features of the multiple neighbor document elements corresponding to the document element with the fusion feature of the document element to obtain the global feature of the document element; Splice the global features of any two document elements among the multiple document elements to obtain the multiple global splicing features.

4. The method according to claim 1, wherein The determining the sorting serial numbers of the multiple document elements based on the position relationship between any two document elements among the multiple document elements includes: Based on the position relationship between any two document elements among the multiple document elements, determine the priorities of the multiple document elements by counting the number of times the document element is ranked in front of other document elements; Based on the priorities of the multiple document elements, determine the sorting serial numbers for the multiple document elements in the order from high to low priority.

5. A sorting device for the content of document pictures, characterized in that, The device includes: An extraction module, configured to call a document element recognition model to recognize a document image to be sorted, and obtain document detection frames of multiple document elements. The document element is the smallest unit constituting the content of the document image, and the document element includes at least one of a picture, a table, a text line, a header / footer, and a dividing line. The document image to be sorted is obtained by photographing a paper document; perform size normalization processing on the document detection frames of each document element so that the document detection frames of different document elements have the same size; based on the normalized document detection frames, extract image features of a preset dimension and original position features of an original dimension of the multiple document elements. The image features include at least one of texture features, brightness, and contrast; A fusion module, configured to convert the original dimension of the original position feature to the preset dimension through a fusion network to obtain position features of multiple document elements; for any document element, when the type of the document element is a non-text line, splice the image feature and the position feature of the document element to obtain the fusion feature of the document element; when the type of the document element is a text line, obtain the image feature and the position feature of the text block where the document element is located, and splice the image feature of the document element, the position feature of the document element, the image feature of the text block, and the position feature of the text block to obtain the fusion feature of the document element; the fusion feature is a feature characterizing the layout format of the document image and the correlation between document elements; A determination module, configured to determine the positional relationship between any two of the multiple document elements based on the fusion features of each document element in the multiple document elements; The determination module is further configured to determine the sorting serial numbers of the multiple document elements based on the positional relationship between any two of the multiple document elements; A merging module, configured to merge each document element belonging to the same text block to obtain multiple text blocks; The determination module is further configured to, for any one of the multiple text blocks, calculate the median of the sorting serial numbers of the document elements within the text block; and use the median as the sorting serial number of the text block; The determination module is further configured to re-determine the sorting serial numbers of the document elements of non-text line type among the multiple document elements and the multiple text blocks in ascending order of the sorting serial numbers; An adding module, configured to add the contents of the multiple document elements to corresponding positions of a blank document based on the sorting serial numbers of the multiple document elements, to obtain an editable electronic document with the same content as the paper document.

6. The device according to claim 5, characterized in that The determination module is configured to: Based on the fusion features of each document element in the multiple document elements, determine multiple global splicing features corresponding to any two of the multiple document elements, where the global splicing feature is a feature characterizing the positional relationship between the two document elements in the document picture to be sorted; Call a positional relationship recognition model to process the multiple global splicing features, to obtain the positional relationship between any two of the multiple document elements, where the positional relationship recognition model is configured to determine the positional relationship between the two document elements based on the global splicing features of the two document elements.

7. The device according to claim 6, characterized in that, The determination module is configured to: For any one document element, obtain multiple neighbor document elements whose distances from the document element are less than a preset distance; Fuse the fusion features of the multiple neighbor document elements corresponding to the document element with the fusion feature of the document element to obtain the global feature of the document element; Splice the global features of any two of the multiple document elements to obtain the multiple global splicing features.

8. The device according to claim 5, characterized in that, The determination module is configured to: Based on the positional relationship between any two of the multiple document elements, determine the priorities of the multiple document elements by counting the number of times the document element ranks in front of other document elements; Based on the priorities of the multiple document elements, determine the sorting serial numbers of the multiple document elements in descending order of the priorities.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, and at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to implement the method for sorting the content of a document picture according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, At least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to implement the method for sorting the content of a document picture according to any one of claims 1 to 4.

11. A computer program product, characterized in that, The computer program product includes computer program code, the computer program code is stored in a computer-readable storage medium, a processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code so that the electronic device executes the method for sorting the content of a document picture according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Text content processing method and device, computer equipment and storage medium

    CN113822283A