Document processing method and device, equipment and storage medium

By segmenting the document image and describing information generation model processing, document content that meets user needs is automatically filtered out, solving the problem of low efficiency and accuracy of document content acquisition in the prior art, and achieving efficient and accurate document content extraction.

CN120356231AInactive Publication Date: 2025-07-22CHINA UNICOM (GUANGDONG) IND INTERNET CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510842371.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, it takes time and effort to filter specified content from a large number of documents, and is prone to omissions, which affects the efficiency and accuracy of document content acquisition.

Method used

By acquiring document images and performing segmentation processing, the description information generation model is used to generate description information of each document area, and the target document content is filtered out based on the similarity matching of user demand information, so as to realize automated processing and accurate extraction of document areas.

Benefits of technology

It improves the accuracy of document processing and content extraction efficiency, reduces user operation steps, and improves the completeness and accuracy of user experience and document content acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356231A_ABST
    Figure CN120356231A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a document processing method and apparatus, a device and a storage medium. The method comprises the steps of obtaining a document image and demand information of a user; performing segmentation processing on the document image to obtain m document areas; according to the m document areas and a description information generation model, description information corresponding to each document area in the n document areas is obtained, the description information generation model is obtained through training according to the multiple sample images and the multiple pieces of sample description information, and n is smaller than or equal to m; target document content corresponding to target description information is obtained according to the demand information and the n pieces of description information, and the target description information comprises description information, of which the matching similarity with the demand information is larger than a preset similarity threshold value, in the n pieces of description information. According to the method, the document image corresponding to the to-be-processed document can be processed based on the user demand information, the target document content meeting the user demand in the document content of the to-be-processed document is obtained, and the document processing accuracy and the content extraction efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of data processing, including but not limited to a document processing method, device, equipment, and storage medium. Background Art

[0002] With the popularization of digital office, the number of documents that enterprises and individuals need to process in their work is increasing day by day. These documents usually include various types of document content such as text, images, and tables. In some scenarios, such as when users summarize work or compile reports based on existing documents, users often need to select specified document content from a large number of documents to meet the needs of different businesses.

[0003] In the related art, usually, keyword matching or type screening is first used to preliminarily screen the documents, and then the user manually verifies the preliminarily screened documents to obtain the required document content. This process not only consumes the user's time and energy, but also the required document content is easily omitted in the manual verification link, affecting the efficiency and accuracy of document content acquisition. Summary of the Invention

[0004] In view of this, the document processing method, device, equipment, and storage medium provided by the embodiments of the present application can process the document image corresponding to the document to be processed based on the user demand information, and obtain the target document content that meets the user demand in the document content of the document to be processed, improving the accuracy of document processing and the content extraction efficiency. The document processing method, device, equipment, and storage medium provided by the embodiments of the present application are implemented as follows: The first aspect of the present application provides a document processing method, including: Obtain a document image and the user's demand information, where the document image includes the document content of the document to be processed, and the document content includes at least one type of content among text type, image type, or table type; Perform segmentation processing on the document image to obtain m document regions, where each document region includes one type of content in the document content, and m is an integer greater than or equal to 1; Generate a description information generation model according to the m document regions and description information, and obtain the description information corresponding to each of the n document regions, where the description information generation model is trained according to a plurality of sample images and a plurality of sample description information, n is an integer greater than or equal to 1, and n is less than or equal to m; According to the demand information and the n description information, obtain the target document content corresponding to the target description information, where the target description information includes the description information among the n description information whose matching similarity with the demand information is greater than a preset similarity threshold.

[0005] In the above technical solution, first, a document image and the user's requirement information are obtained. Then, the document image is segmented, and the document image is segmented into m relatively independent document regions, ensuring that the document content corresponding to each document region is of a single type, thereby improving the accuracy of region segmentation. Next, a pre-trained description information generation model is used to generate respective corresponding description information for different document regions, that is, n pieces of description information corresponding to n document regions are obtained, so that different types of document content can obtain uniformly formatted description information and be matched with the user's requirement information. Finally, according to the user's requirement information and the obtained n pieces of description information, the target description information is determined, and the target document content corresponding to the target description information is obtained. Among them, the target description information can be the description information among the n pieces of description information whose matching similarity with the requirement information is greater than a preset similarity threshold. In this way, the operations required for the user to search for the specified document content are reduced, and the accuracy of document processing and the content extraction efficiency are improved.

[0006] As a possible implementation manner, in the first aspect of the present application, the number of the document images is multiple, and the segmenting the document image to obtain m document regions includes: According to the order of the multiple document images, it is determined whether there is cross-page document content between any two adjacent document images, where the cross-page document content includes one type of content in the document content; Performing splicing processing on any two adjacent document images with the cross-page document content to obtain a spliced document image; Performing the segmenting processing on the spliced document image and the unspliced document images to obtain the m document regions.

[0007] In the above technical solution, the number of document images can be multiple. If there is cross-page document content between any two adjacent document images, such as a cross-page text paragraph or a cross-page table, etc., performing splicing processing on any two adjacent document images with the cross-page document content to obtain a spliced document image, and then performing segmenting processing on the spliced document image and the unspliced document images to obtain m document regions, so as to ensure the integrity of the cross-page document content and avoid content fragmentation caused by segmenting each document image separately.

[0008] As a possible implementation manner, in the first aspect of the present application, the determining whether there is cross-page document content between any two adjacent document images according to the order of the multiple document images includes: When the types of the first document content at the bottom of the first document image are the same as those of the second document content at the top of the second document image, determine whether there is an association relationship between the first document content and the second document content, where the first document image is the image with a prior order among any two adjacent document images, and the second document image is the image with a subsequent order among any two adjacent document images; When there is an association relationship between the first document content and the second document content, determine that there is cross-page document content in any two adjacent document images.

[0009] In the above technical solution, by determining whether there is an association relationship between the two document contents only when the types of the first document content and the second document content are the same, the recognition efficiency of cross-page document content can be effectively improved. For example, when the types of the first document content and the second document content are different, there is no need to check the association relationship, which can reduce unnecessary calculations and thus improve the overall processing efficiency.

[0010] As a possible implementation manner, in the first aspect of the present application, obtaining the description information corresponding to each of the n document regions according to the m document regions and the description information generation model includes: Obtain the n document regions among the m document regions whose corresponding document content types are the same as the target type according to the target type in the requirement information and the document content types corresponding to each of the m document regions; Obtain the n description information according to the description information generation model and the n document regions.

[0011] In the above technical solution, according to the target type in the user's requirement information, for example, the user's requirement information indicates that the user needs to extract the table type in the document, n document regions whose corresponding document content types are the same as the target type are screened out from the m document regions. In this way, unnecessary calculations are reduced, and it is ensured that the generated description information is more in line with the user's requirements.

[0012] As a possible implementation manner, in the first aspect of the present application, obtaining the target document content corresponding to the target description information according to the requirement information and the n description information includes: Calculate the matching similarity between the requirement information and each of the n description information, and determine the description information with a matching similarity greater than the preset similarity threshold as the target description information; Convert the document region corresponding to the target description information to obtain the target document content in the target format.

[0013] In the above technical solution, by calculating the matching similarity between the demand information and the description information, it is possible to screen out the document content corresponding to the description information highly relevant to the user's demand, ensuring the accuracy of the screening result and improving the efficiency and reliability of document processing.

[0014] As a possible implementation manner, in the first aspect of the present application, after generating the model according to the m document regions and the description information and obtaining the description information corresponding to each of the n document regions, the method further includes: Group the n pieces of description information according to the types of the document content corresponding to each of the n document regions, to obtain at least one group, where each group in the at least one group includes at least one piece of description information with the same type of corresponding document content; Display a directory interface, where the directory interface includes a grouping control corresponding to each group in the at least one group; In response to a trigger operation on a target grouping control, display a grouping interface corresponding to the target grouping control, where the grouping interface includes a description control corresponding to each piece of description information in the group corresponding to the target grouping control; In response to a trigger operation on a target description control, obtain the document content of the description information corresponding to the target description control.

[0015] In the above technical solution, by displaying the directory interface and according to the user's trigger operation, the user can select a suitable type according to their own needs, and then select specific description information, so as to quickly obtain the required document content. This hierarchical interaction design not only improves the retrieval efficiency, but also significantly improves the user experience and reduces the complexity of the operation.

[0016] As a possible implementation manner, in the first aspect of the present application, the obtaining the document image and the user's demand information includes: Obtain the document to be processed and the user's voice data; Perform speech recognition on the voice data to obtain the demand information; Convert the document to be processed to obtain the document image.

[0017] In the above technical solution, it is possible to combine the user's voice data with the document processing process to provide a convenient document processing experience. The user can obtain the required document content without complex operations, significantly improving the work efficiency and the user experience.

[0018] The second aspect of the present application provides a document processing device, including: An acquisition module, configured to acquire a document image and requirement information of a user, where the document image includes document content of a document to be processed, and the document content includes at least one type of content among text type, image type, or table type; A segmentation module, configured to perform segmentation processing on the document image to obtain m document regions, where each document region includes one type of content in the document content, and m is an integer greater than or equal to 1; A description module, configured to generate a model according to the m document regions and description information, and obtain description information corresponding to each of the n document regions, where the description information generation model is trained according to a plurality of sample images and a plurality of sample description information, n is an integer greater than or equal to 1, and n is less than or equal to m; A matching module, configured to obtain target document content corresponding to target description information according to the requirement information and the n description information, where the target description information includes description information among the n description information whose matching similarity with the requirement information is greater than a preset similarity threshold.

[0019] A third aspect of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, the method provided in the first aspect of the present application is implemented.

[0020] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method provided in the first aspect of the present application is implemented. Description of the Drawings

[0021] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.

[0022] Figure 1 It is a schematic diagram of an application scenario of the document processing method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a document processing method provided by an embodiment of the present application; Figure 3 It is another schematic flowchart of the document processing method provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the document processing method provided by an embodiment of the present application for splicing adjacent document images with cross-page document content; Figure 5 It is yet another schematic flowchart of the document processing method provided by an embodiment of the present application; Figure 6A schematic flowchart for extracting document content according to a user's trigger operation in the document processing method provided by an embodiment of this application; Figure 7 A schematic diagram for displaying a directory interface in the document processing method provided by an embodiment of this application; Figure 8 A schematic diagram for displaying a grouping interface in the document processing method provided by an embodiment of this application; Figure 9 A schematic structural diagram of a document processing device provided by an embodiment of this application; Figure 10 A schematic structural diagram of a computer device provided by an embodiment of this application. Detailed implementation manners

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will further describe the specific technical solutions of this application in detail with reference to the accompanying drawings in the embodiments of this application. The following embodiments are used to illustrate this application, but are not used to limit the scope of this application.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0025] In the following descriptions, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0026] It should be noted that the terms "first / second / third" involved in the embodiments of this application are used to distinguish similar or different objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0027] In some scenarios, when a user needs to obtain specific content in a document, they can first perform a preliminary screening through methods such as keyword matching or type filtering. For example, the user can use keywords such as "financial data" or "product specifications" to filter out the document content containing these words, or select one or more types of tables, images, or text in the document through the type filtering function to narrow down the scope of the documents that need to be checked. However, although the preliminary screening can reduce the number of documents, the subsequent manual secondary screening still requires the user to check the content of the screened documents one by one. This not only consumes time and energy but also easily leads to the omission of the required document content, thus affecting the efficiency and accuracy of document content acquisition.

[0028] In view of this, embodiments of the present application provide a document processing method, apparatus, device, and storage medium that can process a document image corresponding to a document to be processed based on user demand information to obtain target document content that meets the user's needs in the document content of the document to be processed, improving the accuracy of document processing and the efficiency of content extraction.

[0029] Among them, the document processing method provided by the embodiments of the present application can be applied to electronic devices such as mobile phones, wearable devices (such as smart watches, smart bracelets, smart glasses, etc.), tablet computers, laptop computers, vehicle-mounted terminals, and PCs (Personal Computers), which are not limited herein. The functions implemented by this method can be achieved by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.

[0030] To make the purpose and technical solutions of the present application clearer and more intuitive, the application scenarios of the document processing method provided by the present application will be introduced below with reference to the accompanying drawings.

[0031] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application scenario of the document processing method provided by the embodiments of the present application. The scenario indicated by this application scenario schematic diagram includes a document image 10. Among them, the document image 10 can be obtained by converting the format of an electronic document to be processed, or by scanning a paper document to be processed, or by directly photographing a paper document to be processed, etc., which are not limited herein.

[0032] It should be noted that the document image 10 can include various types of document content, such as Figure 1As shown in the figure, the document image 10 may include, but is not limited to, document content of text type 11 (such as paragraph text, headings, etc.), document content of image type 12 (such as photos, illustrations, etc.), and document content of table type 13 (such as data tables, etc.). Among them, there may be one or more of each type of document content. The image document content 12 may be visual charts such as pie charts, bar charts, line charts, etc., or various images such as product pictures, flowcharts, schematic diagrams, etc., which are not specifically limited herein.

[0033] Through the method provided by the embodiments of the present application, according to the document image 10 and the obtained user requirement information, the target document content that meets the user requirements can be quickly and accurately located and extracted. The target document content is the document part highly relevant to the user requirements. For example, the target document content may be the image document content 12. In this way, the user can obtain the required content without having to flip through a large amount of document content, improving the efficiency and accuracy of document processing and enhancing the user experience.

[0034] To facilitate understanding of how to process the document image corresponding to the document to be processed based on the user requirement information and improve the accuracy of document processing and the efficiency of content extraction, the following introduces an implementation manner of the document processing method provided by the present application.

[0035] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an implementation of the document processing method provided by the embodiments of the present application. As shown in Figure 2 the figure, the method may include the following steps: S201, obtain the document image and the user requirement information.

[0036] In the embodiments of the present application, the document image includes the document content of the document to be processed, and the document content includes at least one type of content among text type, image type, or table type. The user requirement information indicates the document content that the user hopes to obtain from the document to be processed. For example, the user may need to find a certain table data, a specific picture, or a certain text description in the document to be processed.

[0037] It should be noted that the document image may be an image obtained by scanning or photographing a paper-based document to be processed through a scanning device or a photographing device, or an image obtained by converting an electronic document to be processed.

[0038] Exemplarily, when the document to be processed is an electronic document in formats such as PPT, Word, Excel, etc., it can be saved as a document image in formats such as PNG, JPEG, etc., so as to perform subsequent steps such as segmenting the obtained document image, dividing the document content of the document to be processed into multiple document regions containing a single content type, and improving the accuracy of extracting the content required by the user.

[0039] In some possible embodiments, the user's requirement information may be natural language text, such as "Extract the financial data table in the document" or "Find the photos of Product A", etc., so that the user can express their requirements in an intuitive and convenient manner. Exemplarily, when the user's requirement information is natural language text, the requirement information of the user can be obtained through an input device of an electronic device that applies the method provided in this application, such as a mouse, a keyboard, or a touch screen, etc., or can be obtained through instructions or data sent by other devices communicatively connected to the electronic device, which is not limited herein.

[0040] In some possible embodiments, the user requirement information may also be a structured instruction. For example, a formatted command defined by JSON, XML, or specific syntax rules (such as {"action": "extract_table", "keywords": "financial data"}), or a standardized operation template is generated in a graphical operation interface (such as checking the "Table Extraction" option and entering the keyword "financial data"), which is not limited herein. Such structured instructions can be batch-generated through a programming interface (API) and are applicable to automated processing scenarios.

[0041] S202, perform a segmentation process on the document image to obtain m document regions.

[0042] In the embodiments of this application, each document region includes one type of content in the document content, and m is an integer greater than or equal to 1. It can be understood that any document region obtained through the segmentation process may only contain text-type content, may only contain image-type content, may also only contain table-type content, or other single-type content. This division method facilitates subsequent recognition and processing, ensures the singularity and accuracy of the document region division, and meets the user's precise extraction and processing requirements for different types of document content.

[0043] In some possible embodiments, performing a segmentation process on the document image may be implemented through a segmentation algorithm, such as a segmentation algorithm based on edge detection, a segmentation algorithm based on region growing, or a semantic segmentation algorithm based on deep learning (such as a convolutional neural network), etc., which is not limited herein.

[0044] Exemplarily, the edge detection-based segmentation algorithm distinguishes content such as text, images, and tables by detecting lines and contours in the document image. This method utilizes the differences in visual features of different content types. For example, text and table areas usually have relatively regular edges, while image areas may contain more complex contours. The region-growing-based segmentation algorithm can start from seed points and gradually expand to identify coherent regions. This method analyzes the pixel attributes of the image, clusters similar pixels together to form regions, and is suitable for identifying content blocks with similar features. The deep learning-based semantic segmentation algorithm, such as using a convolutional neural network (CNN), can learn the features of the document image and achieve automatic classification and segmentation of different types of content. This method trains the model to understand the semantic information in the image, thereby more intelligently identifying and distinguishing different content types.

[0045] After performing segmentation processing on the document image, the obtained m document regions can be content regions of multiple different types or content regions of the same type, depending on the actual content of the document image. For example, in the case of the document image 10 as shown Figure 1 After performing segmentation processing on the document image, 4 document regions can be obtained, including 2 document regions containing text-type content, 1 document region containing image-type content, and 1 document region containing table-type content. Such a segmentation method can ensure that the content of each document region is single and complete, facilitating subsequent identification and processing.

[0046] In some possible embodiments, after performing segmentation processing on the document image to obtain m document regions, the document processing method provided by this application may further include: When the type of the document content corresponding to the target document region is text type, it is determined whether there are multiple text paragraphs in the document content corresponding to the target document region, where the target document region is one of the m document regions; When there are multiple text paragraphs in the document content corresponding to the target document region, segmentation processing is performed on the target document region to obtain document regions corresponding to each text paragraph.

[0047] It can be understood that compared with document content of types such as images and tables, document content of text type usually includes multiple consecutive paragraphs. Performing paragraph segmentation on the text content to obtain finer document regions can better meet the user's requirements for detailed processing and accurate extraction of text content. Exemplarily, the boundaries of text paragraphs can be determined by detecting paragraph spacing, symbols, the number of indented characters, etc., and the target document region can be segmented.

[0048] S203. Generate a model based on m document regions and description information to obtain the description information corresponding to each of the n document regions among the n document regions.

[0049] In the embodiments of the present application, the description information generation model is trained based on multiple sample images and multiple sample description information. n is an integer greater than or equal to 1 and less than or equal to m.

[0050] It can be understood that among the obtained m document regions, there may be some document regions whose document content is blank regions, decorative graphics, footer numbers, or other non-substantive regions in the document to be processed. By screening the document regions, the document regions containing substantive content can be retained, thereby reducing the number of document regions to be processed and improving the efficiency and accuracy of subsequent processing.

[0051] In some possible embodiments, if all m document regions contain substantive content, all m document regions can be input into the description information generation model to obtain the description information corresponding to each document region respectively. That is, at this time, the value of n can be the same as the value of m.

[0052] In some possible embodiments, the m document regions can be screened by means such as the optical character recognition OCR algorithm and setting a pixel area threshold.

[0053] Exemplarily, when the document content corresponding to the document region is text content, the text can be recognized by the OCR algorithm. If the number of characters of the recognized text is less than a preset character quantity threshold (such as 5 characters, 10 characters, etc.), or the recognized text contains preset text content (such as "header", "footer", "copyright information", etc.), it is determined that the document region is a non-substantive content region and no further processing is required.

[0054] When the document content corresponding to the document region is image content, the pixel area of the document region can be obtained. If the obtained pixel area is less than a preset area threshold, it can be understood that the document region may be a decorative pattern or background rather than a substantive content region and no further processing is required.

[0055] In the embodiments of the present application, the description information generation model can use the document region as the input of the model and input the description information of the document region.

[0056] In some possible embodiments, the description information generation model can be a multimodal model and can be trained through the following steps: First, obtain the training data of the description information generation model. The training data can include multiple sample images and multiple sample description information. The multiple sample images can include various types of document content, such as text, images, tables, etc.

[0057] Then, preprocess the obtained training data, such as downsampling, noise reduction, etc., and classify each document area, annotating its content type so that the subsequent model can perform feature extraction and learning according to different types of content.

[0058] Finally, train the preset initial model according to the preprocessed training data. The initial model can be a deep learning model, such as a convolutional neural network (CNN) or a Transformer-based model, etc. During the training process, the model parameters can be adjusted by setting a loss function or adopting an optimization algorithm, etc., to minimize the difference between the predicted description information and the sample description information, thereby optimizing the performance of the model and obtaining the finally trained multi-modal model, that is, the description information generation model, which can process document areas containing various document content types such as text, images, and tables, and generate description information in a unified format for subsequent matching with user requirement information.

[0059] In some possible embodiments, the description information generation model can also be a hybrid architecture model, which can include multiple sub-models. For example, the multiple sub-models can be machine learning algorithms, such as support vector machines, decision trees, etc., or deep learning models, such as CNN or Transformer-based models, etc. Exemplarily, traditional machine learning algorithms can be used to process specific types of document content, such as document areas of text or table types, while deep learning models can be used to process complex document content, such as document areas of image types, and can generate picture descriptive information according to prompts. Through this hybrid architecture, the description information generation model can exert the efficiency of traditional machine learning algorithms and the representation ability of deep learning models to achieve comprehensive processing of document areas and obtain the description information of each document area.

[0060] In the embodiments of the present application, the description information corresponding to each document area can be natural language text, such as "This area contains the summary part of the 2024 financial report" or "This area contains photos of the product launch site", etc. Such text descriptions can intuitively reflect the content of the document area and are convenient for users to understand and identify. The description information can also be structured instructions, such as {"action": "extract_table", "keywords": "financial quarter"}, etc., which are not limited herein.

[0061] It should be noted that to ensure the efficiency and accuracy of document processing, the user's requirement information and the description information corresponding to each document area can be in the same form to improve the matching efficiency. For example, if the user's requirement information is in natural language text, then the description information should also be generated as natural language text; if the user's requirement information is a structured instruction, then the description information should also be generated in the format of a structured instruction. In this way, it is ensured that the target document content required by the user can be identified quickly and accurately.

[0062] S204. Obtain the target document content corresponding to the target description information according to the requirement information and n pieces of description information.

[0063] In the embodiments of the present application, the target description information includes the description information among the n pieces of description information whose matching similarity with the requirement information is greater than a preset similarity threshold. The target document content refers to the document content of the document area corresponding to the target description information.

[0064] In some possible embodiments, when the type of the target document content is other types except the image type, such as the text type or the table type, the target document content in the target format can be obtained by processing the document area corresponding to the target description information. For example, for the document content of the text type, an editable text file can be obtained through text extraction and formatting processing, while for the document content of the table type, the reconstruction of the table structure and data extraction can be performed to obtain structured table data. In this way, the user can obtain the target document content in the required format, improving the flexibility and practicality of document processing.

[0065] In some possible embodiments, when the requirement information and the description information are natural language text or structured instructions, the matching similarity between the requirement information and each of the n pieces of description information can be calculated respectively through similarity calculation methods such as cosine similarity, Jaccard similarity, or semantic-based similarity, to obtain n matching similarities, and the description information with a matching similarity greater than the preset similarity threshold is determined as the target description information, where the preset similarity threshold can be set according to the actual application scenario and accuracy requirements, and can be set to values such as 0.7 or 0.8, for example.

[0066] In some possible embodiments, when n is an integer greater than or equal to 2, if there are 2 or more pieces of description information among the n pieces of description information whose matching similarity with the requirement information is greater than the preset similarity threshold, the description information with the highest matching similarity can be determined as the target description information.

[0067] In some possible embodiments, when n is an integer greater than or equal to 2, there may be a situation where two or more of the n pieces of description information have a matching similarity greater than a preset similarity threshold with the demand information. In this case, these pieces of description information with a matching similarity greater than the preset similarity threshold can be determined as target description information, and multiple corresponding document contents can be provided for the user to select based on these target description information. In this way, the user can obtain multiple potentially relevant document contents, further improving the flexibility of document content extraction and the user experience.

[0068] The document processing method provided by the embodiments of the present application, first, obtains a document image and the user's demand information. Then, performs segmentation processing on the document image, divides the document image into m relatively independent document regions, and ensures that the document content corresponding to each document region is of a single type, improving the accuracy of region segmentation. Next, uses a pre-trained description information generation model to generate respective corresponding description information for different document regions, that is, obtains n pieces of description information corresponding to n document regions, so that different types of document contents can obtain description information with a unified format and match with the user's demand information. Finally, according to the user's demand information and the obtained n pieces of description information, determines the target description information, and obtains the target document content corresponding to the target description information, where the target description information may be the description information among the n pieces of description information with a matching similarity greater than the preset similarity threshold. In this way, the operations required for the user to search for the specified document content are reduced, and the accuracy of document processing and the content extraction efficiency are improved.

[0069] The following will introduce the method of performing segmentation processing on the document image in the document processing method in conjunction with the accompanying drawings, so as to better understand the implementation process of this document processing method.

[0070] Please refer to Figure 3 , Figure 3 which is another schematic flowchart of the document processing method provided by the embodiments of the present application. As Figure 3 shown, the method may include the following steps: S301, obtain a document image and the user's demand information.

[0071] S302, according to the order of multiple document images, determine whether there is cross-page document content between any two adjacent document images.

[0072] In some possible embodiments, the number of document images of the document to be processed may be multiple, and the multiple document images may be sorted in a certain order, such as page number order, shooting time order, etc., for subsequent processing and analysis.

[0073] Sequentially adjacent document images among multiple document images may have content that spans multiple pages. The content that spans multiple pages includes a type of content in the document content. The content that spans multiple pages indicates a complete content unit, such as a table, a paragraph of text, or an image that spans two or more document contents. If the content that spans multiple pages is not processed and each document image is directly segmented, it may lead to inaccurate generated description information or incomplete or fragmented target content.

[0074] To solve the problems that may be caused by the content that spans multiple pages, the document processing method provided in this application can determine whether there is content that spans multiple pages between any two adjacent document images according to the order of the document images. Exemplarily, the determination method can be based on the continuity analysis of the content at the bottom and top of the page. For example, check whether the text at the bottom of the page is an unfinished sentence, or whether the table has an obvious truncation, etc.

[0075] Exemplarily, for text-type content, it can be checked whether the text line at the bottom of the page ends with a hyphen or an incomplete word, or whether the text line at the top of the page starts with an incomplete sentence. For table-type content, it can be checked whether the table row at the bottom of the page is truncated, or whether the table column at the top of the page does not match the column on the previous page. For image-type content, it can be checked whether the image at the bottom or top of the page is truncated. For example, a part of the image is at the bottom of one page and another part is at the top of the next page, or the boundary features of the document area are detected to determine whether they are different parts of the same image. This is not limited here.

[0076] In some possible embodiments, determining whether there is content that spans multiple pages between any two adjacent document images according to the order of multiple document images includes: When the types of the first document content at the bottom of the first document image and the second document content at the top of the second document image are the same, determining whether there is an association relationship between the first document content and the second document content, where the first document image is the image with the earlier order among any two adjacent document images, and the second document image is the image with the later order among any two adjacent document images; When there is an association relationship between the first document content and the second document content, determining that there is content that spans multiple pages between any two adjacent document images.

[0077] It is understandable that the cross-page document content in adjacent document images is of the same type of document content. By comparing the type of the first document content at the bottom of the first document image with the type of the second document content at the top of the second document image, the efficiency of determining whether there is cross-page document content can be improved. For example, when the types of the first document content and the second document content are different, there is no need to judge the association relationship, and it can be directly determined that there is no cross-page document content in these two adjacent document images.

[0078] Among them, the association relationship can indicate whether the first document content and the second document content belong to the same content, such as the same paragraph of text, the same image, or the same table, etc. When the types of the first document content at the bottom of the first document image and the second document content at the top of the second document image are the same, the above judgment methods, such as the coherence check of text lines, the continuity check of table rows, the integrity check of images, etc., can be used to judge whether there is an association relationship between the first document content and the second document content.

[0079] In some possible embodiments, the type of the document content in each document area can be generated by a description information generation model, or can be determined by other image recognition or text recognition technologies. For example, text content can be recognized by OCR technology, and image or table content can be recognized by image recognition algorithms to ensure the accuracy of the type information of the document area and provide a reliable basis for subsequent processing.

[0080] S303, perform splicing processing on any two adjacent document images with cross-page document content to obtain a spliced document image.

[0081] Please refer to Figure 4 , Figure 4 FIG. is a schematic diagram of the splicing processing of adjacent document images with cross-page document content by the document processing method provided by the embodiment of the present application. As Figure 4 shown, the types of the first document content 41 at the bottom of the first document image and the second document content 42 at the top of the second document image are the same, both are of the table type. If it is detected that the first document content 41 and the second document content 42 belong to the same table, the first document image and the second document image can be spliced to obtain a spliced document image. In this way, when performing segmentation processing on the spliced document area subsequently, a complete document area 43 can be obtained, improving the accuracy and integrity of content extraction.

[0082] In some possible embodiments, when performing splicing processing on any two adjacent document images with cross-page document content, technologies such as image alignment and image fusion can be used to improve the quality of the spliced document image after splicing processing. For example, by cropping blank areas, adjusting image brightness and contrast, or performing gradient blurring on the splicing position, etc., so that as Figure 4The splicing position 44 shown has a natural transition, improving the visual effect of the spliced document image and the accuracy of subsequent processing.

[0083] In some possible embodiments, after splicing is completed, it is possible to further check whether there is still a situation of cross-page document content between the spliced document image and other adjacent document images. If so, the above splicing process needs to be repeated until all cross-page document content is completely spliced to ensure the integrity of the document content. In this way, when processing a document to be processed with a large amount of cross-page content, such as a table with a large number of rows, this multiple splicing method can ensure the integrity and accuracy of the target document content finally obtained.

[0084] S304. Perform segmentation processing on the spliced document image and the unspliced document images to obtain m document regions.

[0085] In some possible embodiments, when there are multiple document images corresponding to the document to be processed and there is cross-page document content between any two adjacent document images, it is possible to perform segmentation processing on the spliced document image and the unspliced document images among the multiple document images to obtain m document regions.

[0086] Exemplarily, if there are 4 document images of the document to be processed in a sequential order, namely document image 1, document image 2, document image 3, and document image 4, and if there is cross-page document content between document image 2 and document image 3, splicing is performed to obtain the spliced document image 5. At this time, only the spliced document image 5, and the unspliced document images 1 and 4 need to be segmented to obtain the required document regions.

[0087] S305. Generate a model based on the m document regions and the description information to obtain the description information corresponding to each of the n document regions.

[0088] S306. According to the requirement information and the n description information, obtain the target document content corresponding to the target description information.

[0089] By implementing the above technical solutions, it is possible to effectively identify cross-page document content, avoid incomplete or broken document content caused by cross-page problems, thereby ensuring the continuity and integrity of the document content, and improving the efficiency and accuracy of document processing.

[0090] Please refer to Figure 5 , Figure 5 which is another schematic flowchart of the document processing method provided by the embodiment of the present application. As Figure 5 shown, the method may include the following steps: S501. Obtain the document image and the requirement information of the user.

[0091] In some possible embodiments, obtaining a document image and user requirement information includes: Obtaining a document to be processed and the user's voice data; Performing speech recognition on the voice data to obtain requirement information; Converting the document to be processed to obtain a document image.

[0092] Exemplarily, a speech recognition model based on deep learning can be used to convert a speech signal into text. By combining the user's voice data with the document processing flow, a convenient document processing experience is provided. The user can obtain the required document content without complex operations, significantly improving work efficiency and user experience.

[0093] S502. Performing segmentation processing on the document image to obtain m document regions.

[0094] S503. According to the target type in the requirement information and the types of the document content corresponding to each of the m document regions, obtaining n document regions among the m document regions where the types of the corresponding document content are the same as the target type.

[0095] In some possible embodiments, the requirement information can be in the form of a natural language text such as "need financial forms" or a structured instruction such as {"action": "extract_table", "keywords": "financial quarter"}. If it is detected that the requirement information contains a user-specified target type (such as "table" or "extract_table"), then all document regions with the document content type of table can be filtered out from the m document regions to obtain n document regions.

[0096] It can be understood that if the target type in the user's requirement information indicates that a table type in the document needs to be extracted, the system will filter out all document regions with the document content type of table from the m document regions as the n document regions. This reduces unnecessary calculations and ensures that the generated description information is more in line with the user's requirements.

[0097] S504. According to the description information generation model and the n document regions, obtaining n pieces of description information.

[0098] S505. Calculating the matching similarity between the requirement information and each of the n pieces of description information, and determining the description information with the matching similarity greater than the preset similarity threshold as the target description information.

[0099] In some possible embodiments, similarity calculation methods such as cosine similarity, Jaccard similarity, or semantic-based similarity can be used to calculate the matching similarity between the requirement information and each piece of description information among the n pieces of description information, and the description information with a matching similarity greater than a preset similarity threshold is determined as the target description information.

[0100] S506. Convert the document area corresponding to the target description information to obtain the target document content in the target format.

[0101] By implementing the above technical solutions, it is possible to conveniently obtain the user's requirement information and document images through speech recognition and document conversion, simplifying the user's operation steps. At the same time, according to the target type in the user's requirement information, n document areas with the same document content type as the target type are selected from the m document areas, improving the efficiency and accuracy of document content extraction.

[0102] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of a process for extracting document content according to a user's trigger operation in the document processing method provided by the embodiments of the present application. As Figure 6 shown, the method may include the following steps: S601. Group the n pieces of description information according to the type of the document content corresponding to each document area among the n document areas to obtain at least one group.

[0103] In some possible embodiments, after generating a model based on the m document areas and the description information to obtain the description information corresponding to each document area among the n document areas, the document processing method provided by the present application can provide a more intuitive and convenient interactive method for document content extraction for the user through the display interface and corresponding controls.

[0104] It should be noted that each group in at least one group includes at least one piece of description information with the same type of corresponding document content. Exemplarily, when the document content of the document to be processed includes text type, image type, and table type, all the description information corresponding to the text type can be divided into one group, all the description information corresponding to the image type can be divided into one group, and all the description information corresponding to the table type can be divided into one group.

[0105] S602. Display the directory interface.

[0106] In some possible embodiments, the directory interface includes group controls corresponding to each group in at least one group.

[0107] Please refer to Figure 7 , Figure 7A schematic diagram showing the directory interface for the document processing method provided by an embodiment of this application, as Figure 7 shown, the directory interface 70 includes three grouping controls, namely a text grouping control 71, an image grouping control 72, and a table grouping control 73.

[0108] In some possible embodiments, the description information included in each grouping can also be displayed at the position of the grouping control, that is, the number of document areas in the corresponding category of the document, as Figure 7 shown, a total of 9 document areas in the text category, 4 document areas in the image category, and 3 document areas in the table category are extracted. This facilitates the user to quickly understand how many document areas there are in each category, thereby improving the efficiency of document processing and the user experience.

[0109] S603. In response to a trigger operation on the target grouping control, display the grouping interface corresponding to the target grouping control.

[0110] In some possible embodiments, the grouping interface includes description controls corresponding to each description information in the grouping corresponding to the target grouping control.

[0111] Please refer to Figure 8 , Figure 8 a schematic diagram showing the grouping interface for the document processing method provided by an embodiment of this application. In some possible embodiments, when the target grouping control triggered by the user is the table grouping control among multiple grouping controls, the grouping interface 80 as Figure 8 shown can be displayed. Among them, the grouping interface 80 includes three description controls corresponding to three description information in this grouping, namely a first description control 81 corresponding to "Current-year financial table", a second description control 82 corresponding to "Historical financial tables", and a third description control 83 corresponding to "Business growth table".

[0112] By displaying the description controls corresponding to each description information, the user can quickly locate the required document content, improving the efficiency of document content extraction and the user experience.

[0113] In some possible embodiments, the grouping interface can also include the document content corresponding to each description information, so that the user can more intuitively understand the content of each document area. For example, in the grouping interface, in addition to displaying the description controls, the complete image or thumbnail of the document area corresponding to each description information can also be displayed respectively. In this way, the user can quickly browse and filter the required content without opening the document, further improving the efficiency and convenience of document processing.

[0114] S604. In response to a trigger operation on the target description control, obtain the document content of the description information corresponding to the target description control.

[0115] In some possible embodiments, the target description control may be one of multiple description controls in a grouped interface. By triggering the target description control, the user can flexibly select the required document content. The obtained document content may be in the form of an image, or may be converted into the format corresponding to the type of the document content, which is not limited herein.

[0116] By implementing the above technical solution, by displaying a directory interface and according to the user's triggering operation, the user can select a suitable type according to their own needs, and then select specific description information, so as to quickly obtain the required document content. This hierarchical interaction design not only improves the retrieval efficiency, but also significantly improves the user experience and reduces the complexity of operations.

[0117] It should be understood that although the steps in the above flowcharts are sequentially displayed according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0118] Based on the foregoing embodiments, an embodiment of the present application provides a document processing device. The device includes each module included and each unit included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the process of implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0119] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the document processing device provided by the embodiment of the present application. As Figure 9 shown, the document processing device includes an acquisition module 901, a segmentation module 902, a description module 903, and a matching module 904, where: The acquisition module 901 is configured to acquire a document image and the user's requirement information, where the document image includes the document content of the document to be processed, and the document content includes at least one type of content among text type, image type, or table type.

[0120] A splitting module 902 is configured to perform splitting processing on a document image to obtain m document regions, where each document region includes one type of content in the document content, and m is an integer greater than or equal to 1.

[0121] A description module 903 is configured to generate a model based on the m document regions and description information to obtain description information corresponding to each of the n document regions, where the description information generation model is trained based on a plurality of sample images and a plurality of sample description information, n is an integer greater than or equal to 1, and n is less than or equal to m.

[0122] A matching module 904 is configured to obtain target document content corresponding to target description information according to requirement information and the n description information, where the target description information includes description information in the n description information whose matching similarity with the requirement information is greater than a preset similarity threshold.

[0123] In some possible embodiments, the number of document images is multiple. The splitting module 902 is further configured to determine whether there is cross-page document content between any two adjacent document images according to the order of the multiple document images, where the cross-page document content includes one type of content in the document content; perform splicing processing on any two adjacent document images with cross-page document content to obtain a spliced document image; perform splitting processing on the spliced document image and the unspliced document images to obtain m document regions.

[0124] In some possible embodiments, the splitting module 902 is further configured to determine whether there is an association relationship between a first document content and a second document content when the type of the first document content at the bottom of a first document image is the same as the type of the second document content at the top of a second document image, where the first document image is the image with a previous order among any two adjacent document images, and the second document image is the image with a subsequent order among any two adjacent document images; When there is an association relationship between the first document content and the second document content, it is determined that there is cross-page document content between any two adjacent document images.

[0125] In some possible embodiments, the description module 903 is further configured to obtain n document regions among the m document regions whose corresponding document content types are the same as the target type according to the target type in the requirement information and the types of the document content corresponding to each of the m document regions; obtain the n description information based on the description information generation model and the n document regions.

[0126] In some possible embodiments, the matching module 904 is further configured to calculate the matching similarity between the requirement information and each of the n description information, determine the description information with the matching similarity greater than the preset similarity threshold as the target description information; and convert the document area corresponding to the target description information to obtain the target document content in the target format.

[0127] In some possible embodiments, the document processing apparatus further includes a display module, configured to group the n description information according to the type of the document content corresponding to each of the n document areas to obtain at least one group, where each group in the at least one group includes at least one description information with the same type of corresponding document content; display a directory interface, where the directory interface includes a group control corresponding to each group in the at least one group; in response to a trigger operation on the target group control, display a group interface corresponding to the target group control, where the group interface includes a description control corresponding to each description information in the group corresponding to the target group control; and in response to a trigger operation on the target description control, obtain the document content of the description information corresponding to the target description control.

[0128] In some possible embodiments, the obtaining module 901 is configured to obtain a document to be processed and the voice data of the user; perform voice recognition on the voice data to obtain requirement information; and convert the document to be processed to obtain a document image.

[0129] The description of the above apparatus embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the apparatus embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0130] It should be noted that in the embodiments of the present application Figure 9 The division of the modules of the document processing apparatus shown is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. In addition, each functional unit in the various embodiments of the present application may be integrated in one processing unit, may exist separately physically, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware, or in the form of a software functional unit, or in the form of a combination of software and hardware.

[0131] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0132] The embodiments of the present application provide a computer device, which may be a server, and its internal structural diagram may be as Figure 10 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0133] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiments are implemented.

[0134] The embodiments of the present application provide a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the steps in the method provided in the above method embodiments.

[0135] Those skilled in the art can understand that Figure 10 the structure shown in

[0136] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 10It runs on the computer device shown. In the memory of the computer device, each program module that makes up the above device can be stored. The computer program composed of each program module enables the processor to execute the steps in the methods of the various embodiments of the present application described in this specification. It should be noted here that the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium, storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0137] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments. The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated here.

[0138] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0139] In the several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be electrical, mechanical or other forms.

[0140] The modules described above as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network elements; some or all of the modules can be selected according to actual needs to achieve the objectives of the solution of this embodiment. Additionally, in each embodiment of this application, the various functional modules can all be integrated in one processing unit, or each module can be a separate unit individually, or two or more modules can be integrated in one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0141] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical discs, and other various media that can store program code. Alternatively, if the above-mentioned integrated unit of this application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of this application essentially or the part that contributes to the relevant technology can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device to execute all or part of the methods described in the various embodiments of this application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical discs, and other various media that can store program code.

[0142] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0143] As described above, the above is only the implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A document processing method, characterized in that Including: Obtaining a document image and user requirement information, where the document image includes the document content of the document to be processed, and the document content includes at least one type of content among text type, image type, or table type; Performing a segmentation process on the document image to obtain m document regions, where each document region includes one type of content in the document content, and m is an integer greater than or equal to 1; Generating a model based on the m document regions and description information generation model to obtain the description information corresponding to each of the n document regions, where the description information generation model is trained based on a plurality of sample images and a plurality of sample description information, n is an integer greater than or equal to 1, and n is less than or equal to m; Based on the requirement information and the n description information, obtaining the target document content corresponding to the target description information, where the target description information includes the description information among the n description information whose matching similarity with the requirement information is greater than a preset similarity threshold.

2. The method according to claim 1, wherein The number of the document images is multiple, and the performing a segmentation process on the document image to obtain m document regions includes: According to the order of the multiple document images, determining whether there is cross-page document content between any two adjacent document images, where the cross-page document content includes one type of content in the document content; Performing a splicing process on any two adjacent document images with the cross-page document content to obtain a spliced document image; Performing the segmentation process on the spliced document image and the unspliced document images to obtain the m document regions.

3. The method according to claim 2, wherein The determining whether there is cross-page document content between any two adjacent document images according to the order of the multiple document images includes: When the types of the first document content at the bottom of the first document image and the second document content at the top of the second document image are the same, determining whether there is an association relationship between the first document content and the second document content, where the first document image is the image with a previous order among the any two adjacent document images, and the second document image is the image with a subsequent order among the any two adjacent document images; When there is an association relationship between the first document content and the second document content, determining that there is the cross-page document content between the any two adjacent document images.

4. The method according to claim 1, wherein The obtaining the description information corresponding to each of the n document regions according to the m document regions and the description information generation model includes: According to the target type in the requirement information and the types of the document content corresponding to each of the m document regions, obtaining the n document regions among the m document regions whose types of the corresponding document content are the same as the target type; Based on the description information generation model and the n document regions, obtaining the n description information.

5. The method according to claim 1, characterized in that, The obtaining the target document content corresponding to the target description information according to the requirement information and the n description information includes: Calculate the matching similarity between the required information and each piece of the n pieces of description information, and determine the description information with a matching similarity greater than the preset similarity threshold as the target description information; Convert the document area corresponding to the target description information to obtain the target document content in the target format.

6. The method according to claim 1, characterized in that, After obtaining, according to the m document areas and the description information generation model, the description information corresponding to each of the n document areas, the method further includes: Group the n pieces of description information according to the types of the document content corresponding to each of the n document areas to obtain at least one group, where each group in the at least one group includes at least one piece of description information with the same type of corresponding document content; Display a directory interface, where the directory interface includes a grouping control corresponding to each group in the at least one group; In response to a trigger operation on a target grouping control, display a grouping interface corresponding to the target grouping control, where the grouping interface includes a description control corresponding to each piece of description information in the group corresponding to the target grouping control; In response to a trigger operation on a target description control, obtain the document content of the description information corresponding to the target description control.

7. The method according to claim 1, wherein The obtaining the document image and the user's required information includes: Obtain the document to be processed and the user's voice data; Perform speech recognition on the voice data to obtain the required information; Convert the document to be processed to obtain the document image.

8. A document processing apparatus, characterized in that, Includes: An obtaining module, configured to obtain a document image and the user's required information, where the document image includes the document content of the document to be processed, and the document content includes at least one type of content among text type, image type, or table type; A splitting module, configured to perform splitting processing on the document image to obtain m document areas, where each document area includes one type of content in the document content, and m is an integer greater than or equal to 1; A description module, configured to obtain, according to the m document areas and the description information generation model, the description information corresponding to each of the n document areas, where the description information generation model is trained according to a plurality of sample images and a plurality of sample description information, n is an integer greater than or equal to 1, and n is less than or equal to m; A matching module, configured to obtain, according to the required information and the n pieces of description information, the target document content corresponding to the target description information, where the target description information includes the description information among the n pieces of description information with a matching similarity greater than the preset similarity threshold with the required information.

9. A computer device, comprising a memory and a processor, the memory storing a computer program that can be run on the processor, characterized in that, When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Document content classification method, system and device and computer readable storage medium

    CN114863408A

  • PDF extraction method and system based on deep learning and layout analysis

    CN119598971A

  • Chart analysis method and device and electronic equipment

    CN119621675A

  • Information retrieval method and device, electronic equipment and storage medium

    CN120179886A

  • Method for extracting data from a structured graphic document, program product and recording medium for implementing such a method

    EP4531005A1