Document layout detection method, training method and device for text processing model
The text processing model enhances document layout detection by using large language models and transformer architectures to improve the accuracy and generalization of layout detection, particularly for ambiguous texts.
Patent Information
- Application Number
- CN202410501879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-04-24
AI Technical Summary
In the prior art, the effectiveness of document layout detection is poor, resulting in low quality of detection results, especially difficult to accurately detect text with ambiguity.
By using a text processing model to perform document layout detection based on multiple texts in the document and corresponding position information, construct target input, and use a large-scale language model to perform document layout detection, adjust the parameters of the text processing model to improve detection accuracy.
It improves the generalization performance and accuracy of document layout detection, especially for ambiguity text, which can achieve better detection results, and realizes effective restoration of document images.
Smart Images

Figure CN118334684B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, in particular to technologies such as natural language processing, computer vision, and deep learning. Specifically, the present disclosure relates to a method for detecting document layout, a method for training a text processing model, a device for detecting document layout, a device for training a text processing model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, training of neural network models, thinking, planning, etc.). It includes both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as natural language processing technology, computer vision technology, speech recognition technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] The document layout detection task requires extracting different components in a document to obtain the text content contained in different components, so as to support restoring the document image into an editable format.
[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, any method described in this section should not be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a method for detecting document layout, a method for training a text processing model, a device for detecting document layout, a device for training a text processing model, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to an aspect of the present disclosure, there is provided a method for detecting document layout, including: determining a plurality of texts in a document and position information of the plurality of texts; constructing a target input based on the plurality of texts and the position information of the plurality of texts; and using a text processing model to process the target input to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the plurality of texts and the order of the at least one text in the document.
[0007] According to another aspect of the present disclosure, there is provided a method for training a text processing model, including: determining a plurality of sample texts and position information of the plurality of sample texts in a sample document; determining true document layout information of the sample document; constructing a sample input based on the plurality of sample texts and the position information of the plurality of sample texts; processing the sample input by using an initial text processing model to obtain a sample document layout detection result, where the sample document layout detection result indicates at least one sample text corresponding to a preset document component among the plurality of sample texts and the order of the at least one sample text in the sample document; and adjusting parameters of the initial text processing model based on the sample document layout detection result and the true document layout information to obtain a target text processing model.
[0008] According to another aspect of the present disclosure, there is provided a document layout detection device, including: a text determination unit configured to determine a plurality of texts and position information of the plurality of texts in a document; a target input construction unit configured to construct a target input based on the plurality of texts and the position information of the plurality of texts; and a first processing unit configured to process the target input by using a text processing model to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the plurality of texts and the order of the at least one text in the document.
[0009] According to another aspect of the present disclosure, there is provided a training device for a text processing model, including: a sample text determination unit configured to determine a plurality of sample texts and position information of the plurality of sample texts in a sample document; a true information determination unit configured to determine true document layout information of the sample document; a sample input construction unit configured to construct a sample input based on the plurality of sample texts and the position information of the plurality of sample texts; a second processing unit configured to process the sample input by using an initial text processing model to obtain a sample document layout detection result, where the sample document layout detection result indicates at least one sample text corresponding to a preset document component among the plurality of sample texts and the order of the at least one sample text in the sample document; and a parameter adjustment unit configured to adjust parameters of the initial text processing model based on the sample document layout detection result and the true document layout information to obtain a target text processing model.
[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and these instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the above method.
[0012] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program implements the above method when executed by a processor.
[0013] According to one or more embodiments of the present disclosure, the present disclosure improves the generalization performance of document layout detection by using a text processing model to perform document layout detection based on multiple texts and corresponding position information in a document, and the text processing model can extract more effective semantic information, thereby improving the accuracy of the document layout detection result. In particular, better detection results can be obtained for texts with ambiguity.
[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0016] Figure 1 A schematic diagram of an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure is shown;
[0017] Figure 2 A flowchart of a document layout detection method according to an exemplary embodiment of the present disclosure is shown;
[0018] Figure 3 A flowchart of a process for constructing a target input according to an exemplary embodiment of the present disclosure is shown;
[0019] Figure 4 A flowchart of a process for constructing a target input according to an exemplary embodiment of the present disclosure is shown;
[0020] Figure 5 A schematic diagram of document layout detection according to an exemplary embodiment of the present disclosure is shown;
[0021] Figure 6 A flowchart of a training method for a document layout detection model according to an exemplary embodiment of the present disclosure is shown;
[0022] Figure 7 A block diagram of the structure of a document layout detection device according to an exemplary embodiment of the present disclosure is shown;
[0023] Figure 8 shows a structural block diagram of a training apparatus for a text processing model according to an exemplary embodiment of the present disclosure; and
[0024] Figure 9 shows a structural block diagram of an exemplary electronic device capable of implementing an embodiment of the present disclosure. Detailed implementation manners
[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0027] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0028] In the related art, some exemplary document layout detection methods have poor effects, and the quality of the obtained detection results is low.
[0029] To solve the above problems, by using a text processing model to perform document layout detection based on multiple texts and corresponding position information in a document, the generalization performance of document layout detection is improved, and the text processing model can extract more effective semantic information, thereby improving the accuracy of document layout detection results. In particular, better detection results can be obtained for texts with ambiguity.
[0030] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Refer to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0032] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of the methods of the present disclosure.
[0033] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0034] In Figure 1 the configuration shown, the server 120 can include one or more components that implement the functions performed by the server 120. These components can include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0035] Users can use the client devices 101, 102, 103, 104, 105, and / or 106 to perform human-computer interactions. The client devices can provide an interface that enables the users of the client devices to interact with the client devices. The client devices can also output information to the users via this interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure can support any number of client devices.
[0036] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0037] Network 110 may be any type of network known to those skilled in the art, which may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0038] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that may be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.
[0039] The computing units in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0040] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0041] In some embodiments, server 120 can be a server of a distributed system or a server incorporating a blockchain. Server 120 can also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0042] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120 or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0043] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0044] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and apparatuses described in the present disclosure.
[0045] According to one aspect of the present disclosure, a method for detecting document layout is provided. Figure 2 A flowchart of a document layout detection method 200 according to an exemplary embodiment of the present disclosure is shown. As Figure 2 shown, the method 200 includes: step S201, determining multiple texts and position information of the multiple texts in the document; step S202, constructing a target input based on the multiple texts and the position information of the multiple texts; and step S203, processing the target input by using a text processing model to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the multiple texts and the order of the at least one text in the document.
[0046] Thus, by using a text processing model to perform document layout detection based on multiple texts and corresponding position information in the document, the generalization performance of document layout detection is improved, and more effective semantic information can be extracted by using the text processing model, thereby enhancing the accuracy of the document layout detection result, especially achieving better detection effects for ambiguous texts.
[0047] The document can be an image containing one or more preset document components, and each of these preset document components includes some text content. According to some embodiments, the preset document components can include at least one of a paragraph title, a text paragraph, a header, a footer, a footnote, a page number, a figure caption, a table name, and table content. By performing layout detection on the document, the text content corresponding to different preset document components in the document can be extracted, thereby supporting the restoration of the document image into an editable format.
[0048] In step S201, multiple texts and position information of the multiple texts in the document are determined.
[0049] According to some embodiments, the multiple texts and the position information of the multiple texts can be determined based on multiple detection frames obtained by performing optical character recognition (OCR) on the document. The multiple texts include the text recognition results of the multiple detection frames, and the position information of the multiple texts can include the coordinates of the multiple detection frames. Thus, by performing OCR on the document, the multiple texts and the corresponding position information in the document can be conveniently obtained.
[0050] In some embodiments, the coordinates of the detection frame can be the coordinates of the center position of the detection frame, or the coordinates of the four corners of the detection frame, or the coordinates of the upper left corner and the lower right corner of the detection frame, or coordinates determined by other means, which are not limited herein.
[0051] It is understandable that the multiple texts and the position information of the multiple texts in the document can also be determined by other means than OCR, which is not limited herein.
[0052] In step S202, based on the multiple texts and the position information of the multiple texts, a target input is constructed.
[0053] In some embodiments, the multiple texts and the position information of the multiple texts can be directly concatenated to obtain a target input for inputting into a text processing model.
[0054] In step S203, the text processing model is used to process the target input to obtain a document layout detection result.
[0055] In some embodiments, a document layout detection result output by the text processing model based on the above target input can be obtained. The document layout detection result can include at least one text among the multiple texts corresponding to a preset document component. The at least one text can have an order in the document layout detection result, and this order can be used to indicate the order of the at least one text in the document.
[0056] After obtaining the at least one text corresponding to the preset document component and its order, post-processing can be performed to restore the document image to an editable format.
[0057] According to some embodiments, the text processing model can be a large language model. A large language model (LLM) is an artificial intelligence system trained on a large dataset, aiming at tasks related to natural language or specific forms of text such as understanding, generating, translating, answering questions, etc. By using a large language model, the multiple input texts and the position information of the multiple texts can be more fully understood, so as to generate a more accurate document layout detection result.
[0058] In some embodiments, the text processing model can be a decoder-only structure. In an exemplary embodiment, the text processing model can include 32 decoder layers. The decoder layers can adopt a structure similar to that of a Transformer decoder, including Grouped Query Attention (GQA), a feed-forward neural network, Root Mean Square Normalization (RMSNorm), and residual connections. In addition, the decoder layers can adopt Relative Position Encoding (RoPE). It has been observed that such a setting method has brought an obvious improvement in the effect of using a large language model for document layout detection.
[0059] In some embodiments, the text processing model can be, for example, obtained by training using the text processing model training method 600 to be introduced below.
[0060] Return to step S202. Figure 3 FIG. shows a flowchart of a process 300 for constructing a target input according to an exemplary embodiment of the present disclosure. The process 300 can be used to implement step S202 in the method 200. As Figure 3 shown, the process 300 includes: step S301, determining numbers of multiple texts based on position information of the multiple texts; and step S302, constructing a target input based on the multiple texts, the numbers of the multiple texts, and the position information of the multiple texts. The document layout detection result can indicate at least one number corresponding to a preset document component among the numbers of the multiple texts and the order of the at least one number, the at least one number can indicate at least one text, and the order of the at least one number can indicate the order of the at least one text in the document.
[0061] Thus, by determining the numbers of multiple texts based on the position information, constructing the input of the text processing model based on the numbers, and then using the text processing model to predict the numbers and order corresponding to the preset document component, on the one hand, the position information and the number order are further utilized as prior information, improving the accuracy of the output result, and on the other hand, it is convenient for the text processing model to predict the document layout detection result, especially the order of at least one text corresponding to the preset document component in the document, thereby improving the generalization and robustness.
[0062] In some embodiments, the numbers of multiple texts can be determined based on the position information of the multiple texts according to a pre-determined rule. In an exemplary embodiment, the multiple texts can be numbered in the order of first from left to right and then from top to bottom. That is, the text closer to the upper edge of the document has a more forward number order; for texts with the same ordinate, the document closer to the left edge of the document has a more forward number order. It can be understood that the numbers of multiple texts can also be determined in other ways, which are not limited herein.
[0063] Figure 4 FIG. shows a flowchart of a process 400 for constructing a target input according to an exemplary embodiment of the present disclosure. The process 400 can be used to implement step S302 in the process 300. The process 400 can include: step S401, constructing input elements for each of the multiple texts, where the input element of each text in the multiple texts includes a first start identifier identifying the number of the text, a text identifier characterizing the text, an intermediate identifier, a position identifier characterizing the position information of the text, and a first end identifier identifying the number of the text; and step S402, constructing a target input based on the input elements of each of the multiple texts.
[0064] Thus, through the above method, multiple texts, the numbers of multiple texts, and the position information of multiple texts can be converted into identifiers (tokens) that are convenient for processing by a text processing model to construct a target input, so that the text processing model can output a more accurate document layout detection result based on the target input.
[0065] An identifier (token) is a unit or element used for inputting into a text processing model, especially a large language model. An identifier can express semantic information, such as a text identifier, or can be used to express other non-semantic content, such as the position identifier to be introduced below, and functional identifiers such as a start identifier, a middle identifier, and an end identifier.
[0066] In some embodiments, in step S401, the input element of the i-th text with the number i in multiple texts can be constructed as: <seg i >text i <bbox>Bbox i < / seg i >, where <seg i > is the first start identifier of identification number i; text i is the text identifier representing the i-th text, which may include one or more text identifiers; <bbox>is an intermediate identifier used to separate text identifiers and position identifiers, and can also be understood as indicating that the identifier between the intermediate identifier and the subsequent end identifier is a position identifier; Bbox i is a position identifier representing the position information of the i-th text, and can include one or more position identifiers; < / seg i > is the first end identifier for identifying the number i. It should be noted that the identifiers corresponding to the angle brackets <> (the first start identifier, intermediate identifier, and first end identifier, and also other identifiers such as the number identifier to be introduced below) are functional identifiers that do not express specific semantic information but are used to identify specific objects. The text identifier can be obtained by performing operations such as word segmentation and embedding on the corresponding text and can express the semantic information in the text. The position identifier can be obtained by embedding the position information and can express the position information corresponding to the text.
[0067] It can be understood that the first start identifier, intermediate identifier, first end identifier, and other functional identifiers below can also be defined in other ways, and the input elements can also be constructed in other ways, which are not limited here.
[0068] In some embodiments, in step S402, the text sequence formed by connecting the input elements of multiple texts end to end can be used as the constructed target input. In an exemplary embodiment, the constructed target input can be: <seg1>text1 <bbox>Bbox1< / bbox> …<seg n >text n <bbox>Bbox n < / seg n >. It can be understood that the target input can also be constructed in other ways, which are not limited herein.
[0069] According to some embodiments, the document layout detection result may include output elements of a preset document component. The output elements of the preset document component may include a second start identifier identifying the preset document component, at least one number identifier identifying at least one number, and a second end identifier identifying the preset document component. The order of at least one number identifier in the document layout detection result may indicate the order of at least one text in the document.
[0070] Thus, through the above method, the mode of outputting the text processing result by the identifier of the text processing model and the output content of the document layout detection task are organically combined, so that the text processing model can generate a more effective and accurate document layout detection result.
[0071] In some embodiments, the output element corresponding to the preset document component may be: <element><seg j ><seg k >< / element> , where <element>To identify a second starting identifier for a preset document component, <seg j ><seg k > is a number identifier for identifying at least one of the numbers j and k,< / element> is the second end identifier identifying the preset document component, where at least one number identifier <seg j ><seg k > in the order in the document layout detection result indicates the order of at least one text in the document, that is, the text corresponding to number j is before the text corresponding to number k in the document.
[0072] It can be understood that for a specific preset document component, the above <element>and< / element> can be replaced with corresponding identifiers. For example, the output element corresponding to the header may be <header><seg i >< / header> , and the output element corresponding to the footer may be <footer><seg i >< / footer> .
[0073] In some embodiments, there may be a nested relationship between preset document components. In an exemplary embodiment, the output element corresponding to the table name... and the output element <table_content>... < / table_content> corresponding to the table content may be nested in the output element corresponding to the table … in the middle. In addition, a segmentation identifier may be included among the output elements of some preset document components for grouping multiple text contents corresponding to the preset document component. For example, the output element corresponding to the table content may include a segmentation identifier <sep>, used to distinguish the content in different cells. Correspondingly, the output elements of an exemplary table can be:
[0074] j k <sep><seg m ><seg n >< / table_content>< / sep> <seg i > <table_content><seg><seg>
[0075] In some embodiments, the output elements corresponding to all the preset document components can be nested between the identifiers.... After obtaining the document layout detection result output by the text processing model, the output elements therein can be parsed to obtain multiple texts included in each preset document component and their corresponding order.
[0076] It can be understood that the output elements can also be constructed in other ways, which are not limited herein.
[0077] According to some embodiments, the preset document component can include multiple columns and text paragraphs corresponding to each column respectively. The document layout detection result can include output elements of each column respectively. The output element of each column in the multiple columns includes a third start identifier identifying the column, an output element of the text paragraph corresponding to the column, and a third end identifier identifying the column.
[0078] The output element of the text paragraph corresponding to the column can include a fourth start identifier identifying the text paragraph, one or more first number identifiers of one or more first numbers in the numbers of multiple texts, and a fourth end identifier identifying the text paragraph. One or more first numbers can indicate one or more first texts in the multiple texts corresponding to the text paragraph, and the order of one or more first number identifiers in the document layout detection result can indicate the order of one or more first texts in the document.
[0079] Thus, through the above method, the layout of a document with multiple columns, especially the text paragraphs corresponding to each column, is effectively and accurately detected.
[0080] In an exemplary embodiment, if the preset text component includes two columns, the document layout detection result can be:
[0081] <section>
[0082] <paragraph><seg i ><seg l >< / paragraph>
[0083] < / section>
[0084] <section>
[0085] <paragraph><seg j ><seg k >< / paragraph>
[0086] < / section>
[0087] Among them, the identifiers (including these two identifiers) between the first <section>and the first< / section> constitute the output element of the first column. Specifically, the first <section>For the third starting identifier that identifies the first column, <paragraph><seg i ><seg l >< / paragraph> is the output element of the text paragraph corresponding to the first column, the first< / section> is the third end identifier indicating the first column.
[0088] The output element for the text paragraph corresponding to the first column, the first <paragraph>As the fourth starting identifier for identifying this text paragraph, <seg i ><seg l > are two first number identifiers for identifying two first numbers i and l, the first< / paragraph> is the fourth end identifier for identifying this text paragraph. Among them, the two first numbers i and l indicate two first texts in multiple texts corresponding to this text paragraph, and the order of the two first number identifiers indicates that the first text corresponding to number i is before the first text corresponding to number l.
[0089] The second <section>and the second< / section> The identifiers between (including these two identifiers) constitute the output elements of the second column. The meanings of the identifiers included are similar to those of the corresponding identifiers in the first column, and will not be elaborated here.
[0090] It can be understood that for more columns, a similar method can be used to determine the output elements corresponding to each column.
[0091] According to some embodiments, the preset document component may include the paragraph title corresponding to the target column in multiple columns, and the output element corresponding to the target column further includes the output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier.
[0092] The output element of the paragraph title corresponding to the target column may include a fifth start identifier for identifying this paragraph title, one or more second number identifiers for identifying one or more second numbers in the numbers of multiple texts, and a fifth end identifier for identifying this text paragraph. One or more second numbers may indicate one or more second texts in multiple texts corresponding to this paragraph title, and the order of one or more second number identifiers in the document layout detection result may indicate the order of one or more second texts in the document.
[0093] Thus, through the above method, the document layout of columns including paragraph titles and text paragraphs is effectively and accurately detected.
[0094] The target column is the column among multiple columns that contains a paragraph title. In some embodiments, the output element corresponding to the target column may be:
[0095] <section>
[0096] <title><segm>< / title>
[0097] <paragraph><seg i ><seg l >< / paragraph>
[0098] < / section>
[0099] Among them, between the third start identifier for identifying the target column <section>and the third end identifier for identifying the target column< / section> , there is an output element of the paragraph title corresponding to the target column, that is <title><segm><segn>< / title> . Among them, <title>is the fifth start identifier for identifying the paragraph title, <segm><segn> are two second number identifiers for identifying the second numbers m and n,< / title> is the fifth end identifier for identifying this paragraph title. Among them, the numbers m and n indicate two second texts in multiple texts corresponding to this paragraph title, and the order of the two second number identifiers indicates that the second text corresponding to number m is before the second text corresponding to number n.
[0100] Figure 5 shows a schematic diagram of document layout detection according to an exemplary embodiment of the present disclosure. As Figure 5 shown, OCR 510 can be performed on the document 502 to obtain an OCR result 504. The text recognition results and position information of multiple detection frames in the OCR result can be used as the multiple texts and the position information of the multiple texts in the document 502, and are used to construct a target input 520. Furthermore, the target input can be processed by a text processing model 530, and a model output result 540 can be obtained. Finally, a document layout detection result can be obtained. Visualizing the document layout detection result in the document 502 can obtain a visualization result 506. It can be seen from the visualization result that the document layout detection result indicates at least one text corresponding to a preset document component and its order. For example, cyan is the text corresponding to the page number, purple is the text corresponding to the header, pink is the text corresponding to the figure caption, brown is the text corresponding to the table name, and orange is the text corresponding to the table content. In addition, the document layout detection result may further include paragraph headings and text paragraphs in different columns. For example, dark green is the paragraph heading in the left column, light green is the text paragraph in the left column, dark blue is the paragraph heading in the right column, and light blue is the text paragraph in the right column.
[0101] According to another aspect of the present disclosure, a method for training a document layout detection model is provided. Figure 6 shows a flowchart of a method 600 for training a document layout detection model according to an exemplary embodiment of the present disclosure. As Figure 6 shown, the method 600 may include: step S601, determining multiple sample texts and the position information of the multiple sample texts in the sample document; step S602, determining the true document layout information of the sample document; step S603, constructing a sample input based on the multiple sample texts and the position information of the multiple sample texts; step S604, processing the sample input by using an initial text processing model to obtain a sample document layout detection result, where the sample document layout detection result indicates at least one sample text corresponding to a preset document component in the multiple sample texts and the order of at least one sample text in the sample document; and step S605, adjusting the parameters of the initial text processing model based on the sample document layout detection result and the true document layout information to obtain a target text processing model.
[0102] It can be understood that the operations and effects of step S601, step S603, and step S604 in the method 600 can respectively refer to the descriptions of step S201-step S203 in the method 200 above, and will not be elaborated here.
[0103] Thus, by training the initial document layout detection model in the above manner, the obtained target layout detection model can be used for document layout detection based on multiple texts and corresponding position information in the document, improving the generalization performance of document layout detection and enhancing the accuracy of document layout detection results. In particular, better detection results can be obtained for texts with ambiguity.
[0104] In some embodiments, the true document layout information determined in step S602 may include at least one ground truth text corresponding to a preset document component in the sample document and the ground truth order of at least one ground truth text in the sample document. The true document layout information may be obtained by annotating the sample document, or by performing document layout detection using other models, or further by manual review and adjustment, or determined by other means, which is not limited herein.
[0105] According to some embodiments, the initial text processing model may be a large language model. By using a large language model, the trained target text processing model can more fully understand the input multiple texts and the position information of the multiple texts, thereby generating more accurate document layout detection results.
[0106] According to some embodiments, step S603, constructing a sample input based on multiple sample texts and the position information of the multiple sample texts may include: determining the numbers of the multiple sample texts based on the position information of the multiple sample texts; and constructing a sample input based on the multiple sample texts, the numbers of the multiple sample texts, and the position information of the multiple sample texts. The sample document layout detection result may indicate at least one number corresponding to a preset document component among the numbers of the multiple sample texts and the order of at least one number. At least one number may indicate at least one sample text, and the order of at least one number may indicate the order of at least one sample text in the sample document.
[0107] Thus, by determining the numbers of multiple sample texts based on the position information and constructing the input of the text processing model based on the numbers, and then using the initial text processing model to predict the numbers and order corresponding to the preset document component, on the one hand, the position information and number order are further utilized as prior information to improve the accuracy of the output result, and on the other hand, it is convenient for the initial text processing model to predict the document layout detection result, especially the order of at least one sample text corresponding to the preset document component in the document, thereby enhancing the generalization and robustness.
[0108] According to some embodiments, constructing a sample input based on multiple sample texts, the numbers of the multiple sample texts, and the position information of the multiple sample texts may include: constructing input elements for each of the multiple sample texts, where the input element of each sample text among the multiple sample texts includes a first start identifier identifying the number of the sample text, a text identifier characterizing the sample text, a middle identifier, a position identifier characterizing the position information of the sample text, and a first end identifier identifying the number of the sample text; and constructing a sample input based on the input elements of each of the multiple sample texts.
[0109] Thus, through the above manner, it is possible to convert multiple sample texts, the numbers of the multiple sample texts, and the position information of the multiple sample texts into identifiers convenient for processing by an initial text processing model to construct a target input, so that the trained target text processing model can output a more accurate document layout detection result based on the corresponding target input.
[0110] According to some embodiments, the sample document layout detection result may include output elements of a preset document component, and the output elements of the preset document component may include a second start identifier identifying the preset document component, at least one number identifier identifying at least one number, and a second end identifier identifying the preset document component. The order of at least one number identifier in the sample document layout detection result may indicate the order of at least one sample text in the sample document.
[0111] Thus, through the above manner, an organic combination is made between the mode of outputting a text processing result by the initial text processing model according to identifiers and the output content of the document layout detection task, so that the trained target text processing model can generate a more effective and accurate document layout detection result.
[0112] According to some embodiments, the preset document component may include multiple columns and text paragraphs corresponding to each of the multiple columns, and the sample document layout detection result may include output elements of each of the multiple columns. The output element of each column among the multiple columns may include a third start identifier identifying the column, the output element of the text paragraph corresponding to the column, and a third end identifier identifying the column.
[0113] The output element of the text paragraph corresponding to the column may include a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more of the numbers of the multiple sample texts, and a fourth end identifier identifying the text paragraph. One or more first numbers may indicate one or more first sample texts among the multiple sample texts corresponding to the text paragraph, and the order of one or more first number identifiers in the sample document layout detection result may indicate the order of one or more first sample texts in the sample document.
[0114] Thus, in the above-described manner, the layout of a document having multiple columns, particularly the text paragraphs corresponding to each column, is effectively and accurately detected.
[0115] According to some embodiments, the preset document component may include the paragraph title corresponding to the target column among the multiple columns, and the output element corresponding to the target column may further include the output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier.
[0116] The output element of the paragraph title corresponding to the target column may include a fifth start identifier for identifying the paragraph title, one or more second number identifiers for identifying one or more of the second numbers among the numbers of multiple sample texts, and a fifth end identifier for identifying the text paragraph. One or more of the second numbers may indicate one or more second sample texts corresponding to the paragraph title among the multiple sample texts, and the order of the one or more second number identifiers in the sample document layout detection result may indicate the order of the one or more second sample texts in the sample document.
[0117] Thus, in the above-described manner, the layout of a document including columns with paragraph titles and text paragraphs is effectively and accurately detected.
[0118] In some embodiments, in step S605, based on the sample document layout detection result, the true document layout information, and a pre-determined loss function, a corresponding loss value may be determined, and then based on the loss value, the parameters of the initial text processing model may be adjusted to obtain the target text processing model. It can be understood that the initial text processing model may also be parameter-tuned based on the sample document layout detection result and the true document layout information in other ways, which is not limited herein.
[0119] According to another aspect of the present disclosure, a document layout detection device is provided. Figure 7 The structural block diagram of a document layout detection device 700 according to an exemplary embodiment of the present disclosure is shown. As Figure 7 shown, the device 700 includes: a text determination unit 710 configured to determine multiple texts in the document and the position information of the multiple texts; a target input construction unit 720 configured to construct a target input based on the multiple texts and the position information of the multiple texts; and a first processing unit 730 configured to process the target input using a text processing model to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the multiple texts and the order of the at least one text in the document.
[0120] It can be understood that the operations and effects of units 710-730 in apparatus 700 can be respectively referred to the descriptions of steps S201-S203 in method 200 above, and will not be elaborated here.
[0121] According to some embodiments, the preset document component may include at least one of a paragraph title, a text paragraph, a header, a footer, a footnote, a page number, a figure caption, a table name, and table content.
[0122] According to some embodiments, the multiple texts and the position information of the multiple texts may be determined based on multiple detection frames obtained by performing optical character recognition on the document. The multiple texts include the text recognition results of the multiple detection frames, and the position information of the multiple texts includes the coordinates of the multiple detection frames.
[0123] According to some embodiments, the text processing model may be a large language model.
[0124] According to some embodiments, the target input construction unit may include: a first number determination subunit configured to determine numbers of the multiple texts based on the position information of the multiple texts; and a first target input construction subunit configured to construct a target input based on the multiple texts, the numbers of the multiple texts, and the position information of the multiple texts. The document layout detection result indicates at least one number corresponding to the preset document component among the numbers of the multiple texts and the order of the at least one number. The at least one number indicates at least one text, and the order of the at least one number indicates the order of the at least one text in the document.
[0125] According to some embodiments, the first target input construction subunit may include: a first input element construction subunit configured to construct input elements for each of the multiple texts. The input element of each text among the multiple texts includes a first start identifier identifying the number of the text, a text identifier characterizing the text, an intermediate identifier, a position identifier characterizing the position information of the text, and a first end identifier identifying the number of the text; and a second target input construction subunit configured to construct a target input based on the input elements of the multiple texts.
[0126] According to some embodiments, the document layout detection result may include output elements of the preset document component. The output elements of the preset document component include a second start identifier identifying the preset document component, at least one number identifier identifying the at least one number, and a second end identifier identifying the preset document component. The order of the at least one number identifier in the document layout detection result may indicate the order of the at least one text in the document.
[0127] According to some embodiments, a preset document component may include a plurality of columns and text paragraphs corresponding to each of the plurality of columns. The document layout detection result may include output elements for each of the plurality of columns. The output element for each column of the plurality of columns may include a third start identifier identifying the column, an output element of the text paragraph corresponding to the column, and a third end identifier identifying the column.
[0128] The output element of the text paragraph corresponding to the column may include a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more of the first numbers in the numbers of a plurality of texts, and a fourth end identifier identifying the text paragraph. One or more of the first numbers may indicate one or more of the first texts corresponding to the text paragraph in the plurality of texts, and the order of the one or more first number identifiers in the document layout detection result may indicate the order of the one or more first texts in the document.
[0129] According to some embodiments, a preset document component may include a paragraph title corresponding to a target column among a plurality of columns, and the output element corresponding to the target column may further include an output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier.
[0130] The output element of the paragraph title corresponding to the target column may include a fifth start identifier identifying the paragraph title, one or more second number identifiers identifying one or more of the second numbers in the numbers of a plurality of texts, and a fifth end identifier identifying the text paragraph. One or more of the second numbers may indicate one or more of the second texts corresponding to the paragraph title in the plurality of texts, and the order of the one or more second number identifiers in the document layout detection result may indicate the order of the one or more second texts in the document.
[0131] According to another aspect of the present disclosure, there is provided a training device for a document layout detection model. Figure 8 FIG. shows a structural block diagram of a training device 800 for a text processing model according to an exemplary embodiment of the present disclosure. As Figure 8 As shown, the apparatus 800 includes: a sample text determination unit 810 configured to determine a plurality of sample texts and position information of the plurality of sample texts in a sample document; a true information determination unit 820 configured to determine true document layout information of the sample document; a sample input construction unit 830 configured to construct a sample input based on the plurality of sample texts and the position information of the plurality of sample texts; a second processing unit 840 configured to process the sample input using an initial text processing model to obtain a sample document layout detection result, where the sample document layout detection result indicates at least one sample text among the plurality of sample texts corresponding to a preset document component and the order of the at least one sample text in the sample document; and a parameter adjustment unit 850 configured to adjust parameters of the initial text processing model based on the sample document layout detection result and the true document layout information to obtain a target text processing model.
[0132] It can be understood that the operations and effects of units 810-850 in the apparatus 800 can respectively refer to the descriptions of steps S601-S605 in the method 600 above, and will not be elaborated here.
[0133] According to some embodiments, the initial text processing model can be a large language model.
[0134] According to some embodiments, the sample input construction unit can include: a second number determination subunit configured to determine numbers of the plurality of sample texts based on the position information of the plurality of sample texts; and a first sample input construction subunit configured to construct a sample input based on the plurality of sample texts, the numbers of the plurality of sample texts, and the position information of the plurality of sample texts. The sample document layout detection result can indicate at least one number among the numbers of the plurality of sample texts corresponding to a preset document component and the order of the at least one number, the at least one number can indicate at least one sample text, and the order of the at least one number can indicate the order of the at least one sample text in the sample document.
[0135] According to some embodiments, the first sample input construction subunit can include: a second input element construction subunit configured to generate input elements for each of the plurality of sample texts, where the input element of each sample text among the plurality of sample texts includes a first start identifier identifying the number of the sample text, a text identifier characterizing the sample text, a middle identifier, a position identifier characterizing the position information of the sample text, and a first end identifier identifying the number of the sample text; and a second sample input construction subunit configured to construct a sample input based on the input elements of the plurality of sample texts.
[0136] According to some embodiments, the sample document layout detection result may include output elements of a preset document component. The output elements of the preset document component include a second start identifier identifying the preset document component, at least one number identifier identifying at least one number, and a second end identifier identifying the preset document component. The order of the at least one number identifier in the sample document layout detection result may indicate the order of at least one sample text in the sample document.
[0137] According to some embodiments, the preset document component may include a plurality of columns and text paragraphs corresponding to each of the plurality of columns. The sample document layout detection result may include output elements of each of the plurality of columns. The output element of each column of the plurality of columns may include a third start identifier identifying the column, the output element of the text paragraph corresponding to the column, and a third end identifier identifying the column.
[0138] The output element of the text paragraph corresponding to the column may include a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more of the numbers of a plurality of sample texts, and a fourth end identifier identifying the text paragraph. The one or more first numbers may indicate one or more first sample texts among the plurality of sample texts corresponding to the text paragraph. The order of the one or more first number identifiers in the sample document layout detection result may indicate the order of the one or more first sample texts in the sample document.
[0139] According to some embodiments, the preset document component may include a paragraph title corresponding to a target column among the plurality of columns. The output element corresponding to the target column may further include the output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier.
[0140] The output element of the paragraph title corresponding to the target column may include a fifth start identifier identifying the paragraph title, one or more second number identifiers identifying one or more of the numbers of a plurality of sample texts, and a fifth end identifier identifying the text paragraph. The one or more second numbers may indicate one or more second sample texts among the plurality of sample texts corresponding to the paragraph title. The order of the one or more second number identifiers in the sample document layout detection result may indicate the order of the one or more second sample texts in the sample document.
[0141] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0142] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.
[0143] Reference Figure 9 , a block diagram of an electronic device 900 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0144] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0145] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device that can input information into the electronic device 900. The input unit 906 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 907 can be any type of device that can present information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 908 can include but are not limited to a magnetic disk, an optical disk. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0146] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods, processes, and / or operations described above. For example, in some embodiments, these methods, processes, and / or operations can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods, processes, and / or operations described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute these methods, processes, and / or operations in any other suitable manner (e.g., by means of firmware).
[0147] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0148] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0149] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0150] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0151] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0152] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0153] It should be understood that various forms of the processes shown above can be used, and steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0154] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.< / sep> < / bbox> < / bbox> < / bbox>
Claims
1. A document layout detection method, comprising: Determining a plurality of texts in the document and location information of the plurality of texts; Constructing a target input based on the plurality of texts and the location information of the plurality of texts, wherein each of the plurality of texts has a corresponding number, and the target input includes input elements of each of the plurality of texts, and the input element of each text in the plurality of texts includes: a first start identifier identifying the number of the text, a text identifier characterizing the text, a middle identifier, a location identifier characterizing the location information of the text, and a first end identifier identifying the number of the text; and Processing the target input by using a text processing model to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the plurality of texts and the order of the at least one text in the document, wherein the document layout detection result includes output elements of the preset document component, and the output elements of the preset document component include: a second start identifier identifying the preset document component, at least one number identifier identifying at least one number corresponding to the at least one text, and a second end identifier identifying the preset document component, and wherein the order of the at least one number identifier in the document layout detection result indicates the order of the at least one text in the document.
2. The method according to claim 1, wherein Constructing a target input based on the plurality of texts and the location information of the plurality of texts includes: Determining the numbers of the plurality of texts based on the location information of the plurality of texts.
3. The method according to claim 1, wherein, The preset document component includes at least one of a paragraph title, a text paragraph, a header, a footer, a footnote, a page number, a figure caption, a table name, and table content.
4. The method according to claim 2, wherein The preset document component includes a plurality of columns and text paragraphs corresponding to each of the plurality of columns, and the document layout detection result includes output elements of each of the plurality of columns. The output element of each column in the plurality of columns includes a third start identifier identifying the column, an output element of the text paragraph corresponding to the column, and a third end identifier identifying the column, The output element of the text paragraph corresponding to the column includes a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more first numbers among the numbers of the plurality of texts, and a fourth end identifier identifying the text paragraph, wherein the one or more first numbers indicate one or more first texts corresponding to the text paragraph among the plurality of texts, and the order of the one or more first number identifiers in the document layout detection result indicates the order of the one or more first texts in the document.
5. The method according to claim 4, wherein The preset document component includes a paragraph title corresponding to a target column among the plurality of columns, and the output element corresponding to the target column further includes an output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier. The output elements of the paragraph title corresponding to the target column division include a fifth start identifier identifying the paragraph title, one or more second number identifiers identifying one or more of the second numbers in the numbers of the plurality of texts, and a fifth end identifier identifying the text paragraph, wherein the one or more second numbers indicate one or more second texts in the plurality of texts corresponding to the paragraph title, and the order of the one or more second number identifiers in the document layout detection result indicates the order of the one or more second texts in the document.
6. The method according to any one of claims 1-5, wherein The plurality of texts and the position information of the plurality of texts are determined based on a plurality of detection frames obtained by performing optical character recognition on the document. The plurality of texts include the text recognition results of the plurality of detection frames, and the position information of the plurality of texts includes the coordinates of the plurality of detection frames.
7. The method according to any one of claims 1-5, wherein The text processing model is a large language model.
8. A method for training a text processing model, comprising: Determining a plurality of sample texts in a sample document and the position information of the plurality of sample texts; Determining the true document layout information of the sample document; Constructing a sample input based on the plurality of sample texts and the position information of the plurality of sample texts, wherein each of the plurality of sample texts has a corresponding number, the sample input includes input elements of each of the plurality of sample texts, and the input element of each sample text in the plurality of sample texts includes a first start identifier identifying the number of the sample text, a text identifier representing the sample text, an intermediate identifier, a position identifier representing the position information of the sample text, and a first end identifier identifying the number of the sample text; Processing the sample input using an initial text processing model to obtain a sample document layout detection result, the sample document layout detection result indicating at least one sample text in the plurality of sample texts corresponding to a preset document component and the order of the at least one sample text in the sample document, wherein the sample document layout detection result includes output elements of the preset document component, the output elements of the preset document component include a second start identifier identifying the preset document component, at least one number identifier identifying the at least one number, and a second end identifier identifying the preset document component, and wherein the order of the at least one number identifier in the sample document layout detection result indicates the order of the at least one sample text in the sample document; and Adjusting the parameters of the initial text processing model based on the sample document layout detection result and the true document layout information to obtain a target text processing model.
9. The method according to claim 8, wherein Constructing a sample input based on the plurality of sample texts and the position information of the plurality of sample texts includes: Determining the numbers of the plurality of sample texts based on the position information of the plurality of sample texts.
10. The method according to claim 9, wherein, The preset document component includes a plurality of columns and text paragraphs corresponding to each of the plurality of columns. The sample document layout detection result includes output elements corresponding to each of the plurality of columns. The output element of each column in the plurality of columns includes a third start identifier identifying the column, an output element of the text paragraph corresponding to the column, and a third end identifier identifying the column. The output element of the text paragraph corresponding to the column includes a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more of the first numbers in the numbers of the plurality of sample texts, and a fourth end identifier identifying the text paragraph. Wherein, the one or more first numbers indicate one or more first sample texts corresponding to the text paragraph in the plurality of sample texts, and the order of the one or more first number identifiers in the sample document layout detection result indicates the order of the one or more first sample texts in the sample document.
11. The method according to claim 10, wherein The preset document component includes a paragraph title corresponding to a target column among the plurality of columns. The output element corresponding to the target column further includes an output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier. The output element of the paragraph title corresponding to the target column includes a fifth start identifier identifying the paragraph title, one or more second number identifiers identifying one or more of the second numbers in the numbers of the plurality of sample texts, and a fifth end identifier identifying the text paragraph. Wherein, the one or more second numbers indicate one or more second sample texts corresponding to the paragraph title in the plurality of sample texts, and the order of the one or more second number identifiers in the sample document layout detection result indicates the order of the one or more second sample texts in the sample document.
12. The method according to any one of claims 8-11, wherein, The initial text processing model is a large language model.
13. A document layout detection device, comprising: A text determination unit configured to determine a plurality of texts in a document and position information of the plurality of texts; A target input construction unit configured to construct a target input based on the plurality of texts and the position information of the plurality of texts. Wherein, each of the plurality of texts has a corresponding number, and the target input includes input elements of each of the plurality of texts. The input element of each text in the plurality of texts includes: a first start identifier identifying the number of the text, a text identifier characterizing the text, a middle identifier, a position identifier characterizing the position information of the text, and a first end identifier identifying the number of the text; and A first processing unit configured to process the target input using a text processing model to obtain a document layout detection result, where the document layout detection result indicates at least one text corresponding to a preset document component among the plurality of texts and the order of the at least one text in the document. Among them, the document layout detection result includes output elements of the preset document component, and the output elements of the preset document component include: a second start identifier identifying the preset document component, at least one number identifier identifying at least one number corresponding to the at least one text, and a second end identifier identifying the preset document component. And among them, the order of the at least one number identifier in the document layout detection result indicates the order of the at least one text in the document.
14. The apparatus according to claim 13, wherein, The target input construction unit includes: A first number determination sub-unit, configured to determine numbers of the plurality of texts based on position information of the plurality of texts.
15. The device according to claim 13, wherein, The preset document component includes at least one of a paragraph title, a text paragraph, a header, a footer, a footnote, a page number, a figure caption, a table name, and table content.
16. The apparatus according to claim 14, wherein, The preset document component includes a plurality of columns and text paragraphs corresponding to the plurality of columns respectively. The document layout detection result includes output elements of the plurality of columns respectively. The output element of each column in the plurality of columns includes a third start identifier identifying the column, an output element of the text paragraph corresponding to the column, and a third end identifier identifying the column. The output element of the text paragraph corresponding to the column includes a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more first numbers among the numbers of the plurality of texts, and a fourth end identifier identifying the text paragraph. Among them, the one or more first numbers indicate one or more first texts among the plurality of texts corresponding to the text paragraph, and the order of the one or more first number identifiers in the document layout detection result indicates the order of the one or more first texts in the document.
17. The apparatus according to claim 16, wherein, The preset document component includes a paragraph title corresponding to a target column among the plurality of columns. The output element corresponding to the target column further includes an output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier. The output element of the paragraph title corresponding to the target column includes a fifth start identifier identifying the paragraph title, one or more second number identifiers identifying one or more second numbers among the numbers of the plurality of texts, and a fifth end identifier identifying the text paragraph. Among them, the one or more second numbers indicate one or more second texts among the plurality of texts corresponding to the paragraph title, and the order of the one or more second number identifiers in the document layout detection result indicates the order of the one or more second texts in the document.
18. The apparatus according to any one of claims 13 - 17, wherein, The plurality of texts and the position information of the plurality of texts are determined based on a plurality of detection frames obtained by performing optical character recognition on the document. The plurality of texts include text recognition results of the plurality of detection frames, and the position information of the plurality of texts includes coordinates of the plurality of detection frames.
19. The apparatus according to any one of claims 13 - 17, wherein, The text processing model is a large language model.
20. A training device for a text processing model, including: A sample text determination unit, configured to determine a plurality of sample texts in a sample document and position information of the plurality of sample texts; A true information determination unit, configured to determine true document layout information of the sample document; A sample input construction unit, configured to construct a sample input based on the plurality of sample texts and the position information of the plurality of sample texts, wherein each of the plurality of sample texts has a corresponding number, and the sample input includes input elements of each of the plurality of sample texts. The input element of each sample text in the plurality of sample texts includes: a first start identifier identifying the number of the sample text, a text identifier characterizing the sample text, an intermediate identifier, a position identifier characterizing the position information of the sample text, and a first end identifier identifying the number of the sample text; A second processing unit, configured to process the sample input by using an initial text processing model to obtain a sample document layout detection result. The sample document layout detection result indicates at least one sample text corresponding to a preset document component among the plurality of sample texts and the order of the at least one sample text in the sample document. The sample document layout detection result includes output elements of the preset document component. The output elements of the preset document component include: a second start identifier identifying the preset document component, at least one number identifier identifying at least one number corresponding to the at least one sample text, and a second end identifier identifying the preset document component. And wherein the order of the at least one number identifier in the sample document layout detection result indicates the order of the at least one sample text in the sample document; and A parameter adjustment unit, configured to adjust parameters of the initial text processing model based on the sample document layout detection result and the true document layout information to obtain a target text processing model.
21. The apparatus according to claim 20, wherein, The sample input construction unit includes: A second number determination subunit, configured to determine numbers of the plurality of sample texts based on the position information of the plurality of sample texts.
22. The apparatus according to claim 21, wherein, The preset document component includes a plurality of columns and text paragraphs respectively corresponding to the plurality of columns. The sample document layout detection result includes output elements of each of the plurality of columns. The output element of each column in the plurality of columns includes a third start identifier identifying the column, output elements of the text paragraph corresponding to the column, and a third end identifier identifying the column, The output elements of the text paragraph corresponding to the column include a fourth start identifier identifying the text paragraph, one or more first number identifiers identifying one or more first numbers among the numbers of the plurality of sample texts, and a fourth end identifier identifying the text paragraph. The one or more first numbers indicate one or more first sample texts corresponding to the text paragraph among the plurality of sample texts, and the order of the one or more first number identifiers in the sample document layout detection result indicates the order of the one or more first sample texts in the sample document.
23. The apparatus according to claim 22, wherein, The preset document component includes a paragraph title corresponding to a target column among the multiple columns, and the output element corresponding to the target column further includes an output element of the paragraph title corresponding to the target column located between the third start identifier and the third end identifier. The output element of the paragraph title corresponding to the target column includes a fifth start identifier for identifying the paragraph title, one or more second number identifiers for identifying one or more second numbers among the numbers of the multiple sample texts, and a fifth end identifier for identifying the text paragraph. Among them, the one or more second numbers indicate one or more second sample texts among the multiple sample texts corresponding to the paragraph title, and the order of the one or more second number identifiers in the sample document layout detection result indicates the order of the one or more second sample texts in the sample document.
24. The apparatus according to any one of claims 20 - 23, wherein, The initial text processing model is a large language model.
25. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-12.
26. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
27. A computer program product includes a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-12.
Citation Information
Patent Citations
Document image processing method and device, equipment and medium
CN114663902A
Document processing model training method, document processing method, device and equipment
CN115809325A