Document Processing Method, Apparatus, Electronic Device, and Storage Medium

By analyzing and judging text line semantics and layout information in the document, the independent customization problem of continuous fold line judgment processing for different types of documents in the prior art is solved, and accurate and efficient judgment of different types of documents is achieved.

CN114282517BActive Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111422878.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-06-24
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

In the prior art, the processing rules for continuous fold judgments for different types of documents are independently customized, which makes it difficult to uniformly and effectively process them in different business scenarios, document formats or vertical documents.

Method used

By acquiring the document, the first text line and the second text line are parsed, their semantic information and layout information are determined, and based on these information, they are determined whether the two are continuous folded lines.

Benefits of technology

It realizes the accuracy and efficiency of continuous folding of text lines in different types of documents, and can more effectively handle documents in different business scenarios and document formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114282517B_ABST
    Figure CN114282517B_ABST
Patent Text Reader

Abstract

The present disclosure provides a document processing method, apparatus, electronic device, and storage medium, which relate to the field of artificial intelligence, specifically to the fields of deep learning and natural language processing technologies. The implementation solution is as follows: Obtain a document, parse a first text line and a second text line from the document, determine first semantic information of the first text line, and determine second semantic information of the second text line, determine first layout information of the first text line, and determine second layout information of the second text line, and determine whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information. Thus, it is possible to fully combine the semantic information of text lines with the layout information of text lines in a document to determine whether different text lines are consecutive line breaks, which can effectively improve the accuracy of determining whether different text lines in a document are consecutive line breaks and effectively improve the determination efficiency of consecutive line breaks for a document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to deep learning and natural language processing technologies, and particularly to a document processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as several major directions such as machine learning, deep learning, big data processing technology, and knowledge graph technology.

[0003] In related technologies, generally, corresponding determination processing rules for continuous line breaks in a document are independently customized for different types of documents. Different types of documents include, for example, documents in different business scenarios, or documents in different document formats, or documents in different vertical categories under the same document format. Summary of the Invention

[0004] The present disclosure provides a document processing method, apparatus, electronic device, storage medium, and computer program product.

[0005] According to a first aspect of the present disclosure, there is provided a document processing method, including: obtaining a document; parsing a first text line and a second text line from the document; determining first semantic information of the first text line and determining second semantic information of the second text line; determining first layout information of the first text line and determining second layout information of the second text line; and determining whether the first text line and the second text line are continuous line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0006] According to a second aspect of the present disclosure, there is provided a document processing apparatus, including: an obtaining module for obtaining a document; a parsing module for parsing a first text line and a second text line from the document; a first determining module for determining first semantic information of the first text line and determining second semantic information of the second text line; a second determining module for determining first layout information of the first text line and determining second layout information of the second text line; and a third determining module for determining whether the first text line and the second text line are continuous line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0007] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the document processing method according to the embodiment of the first aspect of the present disclosure.

[0008] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the document processing method according to the embodiment of the first aspect of the present disclosure.

[0009] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the document processing method according to the embodiment of the first aspect of the present disclosure.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram of the text line position of the document in the embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram according to the second embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram according to the third embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0017] Figure 6 is a schematic diagram of the text pre-training model structure in the embodiment of the present disclosure;

[0018] Figure 7 is a schematic diagram of the first consecutive line break recognition model structure in the embodiment of the present disclosure;

[0019] Figure 8 is a schematic diagram of the second consecutive line break recognition model structure in the embodiment of the present disclosure;

[0020] Figure 9 is a schematic diagram according to the fifth embodiment of the present disclosure;

[0021] Figure 10 is a schematic diagram according to the sixth embodiment of the present disclosure;

[0022] Figure 11 shows a schematic block diagram of an exemplary electronic device that can be used to implement the document processing method of the embodiments of the present disclosure. Detailed implementation manners

[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0024] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure.

[0025] It should be noted that the execution subject of the document processing method in this embodiment is a document processing device, which can be implemented in software and / or hardware, and the device can be configured in an electronic device. The electronic device can include, but is not limited to, a terminal, a server, etc.

[0026] The embodiments of the present disclosure relate to the field of artificial intelligence technology, specifically to the fields of deep learning and natural language processing technology.

[0027] Among them, artificial intelligence (AI) is an abbreviation in English. It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.

[0028] Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images and sounds. The ultimate goal of deep learning is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images and sounds.

[0029] Natural Language Processing (NLP) is a field in computer science, artificial intelligence, and linguistics that focuses on the interaction between computers and human (natural) languages. It takes language as the object and uses computer technology to analyze, understand and process natural language, that is, uses the computer as a language research tool to quantitatively study and process language information with the support of the computer, and provides a language description that can be used jointly by humans and computers.

[0030] As Figure 1 shown, the document processing method includes:

[0031] S101: Obtain a document.

[0032] Among them, the format of the document can be various formats including text content. For example, it can be a Portable Document Format (PDF) document, a text document, a scanned copy, or it can also be a document in the format of a picture, a web page, etc. There is no limitation on this.

[0033] In the embodiments of the present disclosure, when obtaining a document, a corresponding document acquisition device can be pre-configured for the document processing device. The document acquisition device can acquire a document containing text content. For example, it can obtain a PDF document, a text document, or a data transmission interface can be pre-configured for the document processing device, and a scanned copy collected by a device such as a scanner can be received via this data transmission interface. There is no limitation on this.

[0034] In other embodiments, a document can also be obtained by receiving pictures transmitted by other electronic devices via the data transmission interface, or a document can also be obtained by crawling and parsing a web page. There is no limitation on this.

[0035] S102: Parse a first text line and a second text line from the document.

[0036] Among them, a text line is composed of an arrangement of multiple characters and can be parsed from the document. The first text line can be a text line that is relatively earlier in position in the document than the second text line, and the second text line is a text line arranged after the first text line.

[0037] In the embodiments of the present disclosure, when parsing a first text line and a second text line from the document, the intelligent document analysis platform TextMind can be used to analyze and process the obtained document, extract the text content in the document and the position where the text content is located, so as to parse the document into the form of text lines, or the document can be processed into the form of multiple sample text lines. Two text lines can be included in the sample text line. The one that is relatively earlier in position in the document is the first text line, and the text line behind the first text line can be the second text line. There is no limitation on this.

[0038] As Figure 2 shown, Figure 2It is a schematic diagram of the text line position of a document in an embodiment of the present disclosure. The first line of the body text under the document title can be used as the first text line, and the second text line can be the next line of the first text line. At this time, the first text line and the second text line are continuously wrapped, or the first text line can be the last line of the body text before column splitting, and the second text line can be the first line of the text in the column splitting. In this case, the first text line and the second text line are not continuously wrapped.

[0039] S103: Determine the first semantic information of the first text line and determine the second semantic information of the second text line.

[0040] Among them, the semantic information is used to describe the meaning of the text content of the text line. The meaning of the text content of the first text line can be called the first semantic information, and the meaning of the text content of the second text line can be called the second semantic information.

[0041] In an embodiment of the present disclosure, when determining the first semantic information of the first text line and determining the second semantic information of the second text line, the first text line and the second text line can be respectively subjected to semantic analysis processing through a text pre-training language model to understand the semantic information of the first text line and the second text line, and the outputs of the model are respectively used as the first semantic information of the first text line and the second semantic information of the second text line.

[0042] For example, the first semantic information can be, for example, the text content meaning of the words and text structure in the first text line in the context of the corresponding document semantics, and the second semantic information can be, for example, the text content meaning of the words and text structure in the second text line in the context of the corresponding document semantics.

[0043] S104: Determine the first layout information of the first text line and determine the second layout information of the second text line.

[0044] Among them, the layout information is used to describe the spatial position information of the text line in the document. The spatial position information of the first text line in the document can be called the first layout information, and the spatial position information of the second text line in the document can be called the second layout information.

[0045] For example, the spatial position information can be, for example, the layout module position of the text line in the document and the text line sequence number, etc. Among them, the layout module in the document can be, for example, the title module, body text module, table module or column splitting module, etc. in the document. The first layout information can be, for example, that the first text line is in the column splitting module of the document and is located in the first line of the column splitting module, and the second layout information can be, for example, that the second text line is in the column splitting module of the document and is located in the second line of the column splitting module.

[0046] In the embodiments of the present disclosure, when determining the first layout information of the first text line and the second layout information of the second text line, the intelligent document analysis platform TextMind can be used to analyze and process the obtained document, so as to obtain the size of the document page, the text content of each character, and the relative coordinates in the two-dimensional space. Then, the spatial position information of the text line in the document can be determined according to the relative coordinates of the characters in the text line, and the spatial position information of the first text line in the document is determined as the first layout information of the first text line, and the spatial position information of the second text line in the document is determined as the second layout information.

[0047] In some other embodiments, the overall layout information of the document can be processed by using a text pre-trained language model according to the positions of the characters in the text line in the two-dimensional space, and the overall layout information of the document can be analyzed according to the size information and position information of each character, which is not limited herein.

[0048] For example, according to the size information and position information of the characters in the text line, it can be determined whether the text line is the overall layout information such as a title, a column, or a table. Then, according to the overall layout information of the document, the layout information of the first text line and the second text line is extracted, and the obtained layout information is respectively used as the first layout information corresponding to the first text line and the second layout information corresponding to the second text line. Or, any other possible method can also be used to determine the first layout information of the first text line and the second layout information of the second text line, which is not limited herein.

[0049] S105: Determine whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0050] After obtaining the first semantic information and the first layout information of the first text line, and the second semantic information and the second layout information of the second text line as described above, it can be determined whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0051] In an embodiment of the present disclosure, when determining whether a first text line and a second text line are consecutive line breaks according to first semantic information, first layout information, second semantic information, and second layout information, a text pre-trained language model may be used to determine whether the overall semantics is reasonable after connecting the first text line and the second text line based on the first semantic information and the second semantic information, and the first layout information and the second layout information may be jointly used to comprehensively determine whether the first text line and the second text line are text lines of different layout modules. The different layout modules may be, for example, a title or a column break, etc. If the overall semantics is reasonable after connecting the first text line and the second text line, and the first text line and the second text line are not in different modules, it indicates that the first text line and the second text line are consecutive line breaks. If the overall semantics is unreasonable after connecting the first text line and the second text line, and the first text line and the second text line are in different modules, it indicates that the first text line and the second text line are not consecutive line breaks.

[0052] In this embodiment, by obtaining a document, parsing the first text line and the second text line from the document, determining the first semantic information of the first text line, determining the second semantic information of the second text line, determining the first layout information of the first text line, and determining the second layout information of the second text line, and determining whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information, it is possible to fully combine the semantic information of the text line and the layout information of the text line in the document to determine whether different text lines are consecutive line breaks, which can effectively improve the determination accuracy of whether different text lines in the document are consecutive line breaks and effectively improve the determination efficiency of consecutive line breaks for the document.

[0053] Figure 3 It is a schematic diagram according to the second embodiment of the present disclosure.

[0054] As Figure 3 shown, the document processing method includes:

[0055] S301: Obtain a document.

[0056] S302: Parse the first text line and the second text line from the document.

[0057] For the descriptions of S301 - S302, reference may be made to the above embodiments for examples, and details are not described herein again.

[0058] S303: Determine the overall semantics of the document.

[0059] Among them, the overall semantics of the document is used to describe the overall text meaning of the document text.

[0060] In the embodiments of the present disclosure, when determining the overall semantics of a document, the text content in the document can be extracted to obtain streaming text content, and then the streaming text content is input into a pre-trained language model for text. The pre-trained language model for text is used to understand and process the text content of the document, and the output of the model is used as the overall semantics of the document.

[0061] S304: Determine the first local semantics of the first text line and determine the second local semantics of the second text line.

[0062] After determining the overall semantics of the document in the above embodiments of the present disclosure, the first local semantics of the first text line and the second local semantics of the second text line can be determined respectively.

[0063] Among them, the first local semantics refers to the meaning of the text content of the vocabulary and text structure in the first text line in the context of the text semantics of this line, and the second local semantics refers to the meaning of the text content of the vocabulary and text structure in the second text line in the context of the text semantics of this line.

[0064] In the embodiments of the present disclosure, when determining the first local semantics of the first text line and determining the second local semantics of the second text line, a natural language processing model can be used to perform separate semantic understanding processing on the text lines in the document, perform semantic understanding processing on the first text line to obtain the first local semantics of the first text line, and perform semantic understanding processing on the second text line to obtain the second local semantics of the second text line.

[0065] S305: Obtain the first context semantics of the first local semantics relative to the overall semantics, and obtain the second context semantics of the second local semantics relative to the overall semantics.

[0066] Among them, the context semantics refers to the meaning of the text content of the local semantics of the text line in the context of the overall semantics of the document. The text meaning of the first local semantics of the first text line in the context of the overall semantics of the document can be called the first context semantics, and the text meaning of the second local semantics of the second text line in the context of the overall semantics of the document can be called the second context semantics.

[0067] In the embodiments of the present disclosure, when obtaining the first context semantics of the first local semantics relative to the overall semantics and obtaining the second context semantics of the second local semantics relative to the overall semantics, a context search algorithm can be used to perform semantic analysis processing on the first local semantics of the first text line in the context of the overall semantics of the document to obtain the first context semantics of the first local semantics relative to the overall semantics, and use the context search algorithm to perform semantic analysis processing on the second local semantics of the second text line in the context of the overall semantics of the document to obtain the second context semantics of the second local semantics relative to the overall semantics.

[0068] S306: Take the first local semantics and the first context semantics as the first semantic information, and take the second local semantics and the second context semantics as the second semantic information.

[0069] After obtaining the first context semantics of the first local semantics relative to the overall semantics of the document and the second context semantics of the second local semantics relative to the overall semantics of the document in the embodiments of the present disclosure, the first local semantics and the first context semantics can be jointly used as the first semantic information of the first text line, and the second local semantics and the second context semantics can be jointly used as the second semantic information. The first semantic information and the second semantic information can be used to jointly determine whether the first text line and the second text line are continuous line breaks based on the first layout information and the second layout information.

[0070] In this embodiment, by determining the overall semantics of the document, determining the first local semantics of the first text line, and determining the second local semantics of the second text line, obtaining the first context semantics of the first local semantics relative to the overall semantics of the document, and obtaining the second context semantics of the second local semantics relative to the overall semantics of the document, and taking the first local semantics and the first context semantics as the first semantic information, and taking the second local semantics and the second context semantics as the second semantic information, it is possible to obtain the first local semantics and the second local semantics and obtain the context semantics of the first text line and the second text line according to the overall semantics of the document, enriching the semantic information of the obtained first text line and second text line, enabling the acquisition of semantic information in more dimensions, enhancing the representational ability of the semantic information when determining whether the first text line and the second text line are continuous, and assisting in improving the accuracy of the line break determination process for the text lines in the document.

[0071] S307: Determine the first one-dimensional position information of the first text line in the document, and determine the second one-dimensional position information of the second text line in the document.

[0072] The one-dimensional position information is used to represent the serial number information of the text line in the text line sequence of the document. For example, it can indicate that a text line in the document is the first text line of the document. The serial number information of the first text line in the text line sequence of the document can be the first one-dimensional position information of the first text line in the document, and the serial number information of the second text line in the text line sequence of the document can be the second one-dimensional position information of the second text line in the document.

[0073] In the embodiments of the present disclosure, the text lines in a document can be marked to obtain the numbering information of each text line in the document. When determining the first one-dimensional position information of the first text line in the document and the second one-dimensional position information of the second text line in the document, the first one-dimensional position information of the first text line in the document can be determined according to the numbering information of the first text line, and the second one-dimensional position information of the second text line in the document can be determined according to the numbering information of the second text line.

[0074] Optionally, in some other embodiments, the first one-dimensional position information may also be the layout module position of the first text line in the document. For example, the first text line may be located in the title module of the document, and the second one-dimensional position information may also be the layout module position of the second text line in the document. For example, the second text line may be located in the body module of the document, etc., and there is no limitation thereto.

[0075] S308: Determine the first spatial position information of the first text line in the document and determine the second spatial position information of the second text line in the document.

[0076] Among them, the first spatial position information of the first text line in the document can be determined by the spatial position information of each character in the first text line, and the second spatial position information of the second text line in the document can be determined by the spatial position information of each character in the second text line.

[0077] Among them, the spatial position information refers to the two-dimensional position information of the text line in the document, the first spatial position information refers to the two-dimensional position information of the first text line in the document, and the second spatial position information refers to the two-dimensional position information of the second text line in the document.

[0078] In the embodiments of the present disclosure, when determining the first spatial position information of the first text line in the document and the second spatial position information of the second text line in the document, the characters in the first text line can be box-selected, the coordinate information of the character box where the characters in the first text line are located can be extracted, and the coordinate information can be processed to obtain the first spatial position information of the first text line in the document. The characters in the second text line are box-selected, the coordinate information of the character box where the characters in the second text line are located is extracted, and the coordinate information is processed to obtain the second spatial position information of the second text line in the document.

[0079] S309: Take the first one-dimensional position information and the first spatial position information as the first layout information, and take the second one-dimensional position information and the second spatial position information as the second layout information.

[0080] In an embodiment of the present disclosure, after determining the first one-dimensional position information of the first text line in the document, the second one-dimensional position information of the second text line in the document, the first spatial position information of the first text line in the document, and the second spatial position information of the second text line in the document, the first one-dimensional position information and the first spatial position information may be used as the first layout information, and the second one-dimensional position information and the second spatial position information may be used as the second layout information.

[0081] Among them, the first layout information and the second layout information may be used to jointly determine whether the first text line and the second text line are consecutive line breaks based on the first semantic information and the second semantic information.

[0082] In this embodiment, by determining the first one-dimensional position information of the first text line in the document, determining the second one-dimensional position information of the second text line in the document, determining the first spatial position information of the first text line in the document, determining the second spatial position information of the second text line in the document, using the first one-dimensional position information and the first spatial position information as the first layout information, and using the second one-dimensional position information and the second spatial position information as the second layout information, the acquisition of the layout information of the text line can be enriched to a large extent. Therefore, the relatively rich layout information can be used to determine whether the first text line and the second text line are consecutive line breaks, so that the global features of the text lines in the document can be better learned, and the accuracy of the line break determination process for the text lines in the document can be effectively improved.

[0083] S310: Determine whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0084] For the description of S310, reference may be made to the above embodiments for examples, and details are not described herein again.

[0085] In this embodiment, by determining the overall semantics of the document, determining the first local semantics of the first text line, and determining the second local semantics of the second text line, obtaining the first context semantics of the first local semantics relative to the overall semantics, and obtaining the second context semantics of the second local semantics relative to the overall semantics, and taking the first local semantics and the first context semantics as the first semantic information, and taking the second local semantics and the second context semantics as the second semantic information, it is possible to obtain the first local semantics and the second local semantics, and obtain the context semantics of the first text line and the second text line according to the overall semantics of the document, enriching the semantic information of the first text line and the second text line, so that semantic information in multiple dimensions can be obtained, enhancing the representational ability of the semantic information when determining whether the first text line and the second text line are continuous, and assisting in improving the accuracy of the line-breaking determination process for the text lines in the document. By determining the first one-dimensional position information of the first text line in the document, and determining the second one-dimensional position information of the second text line in the document, determining the first spatial position information of the first text line in the document, and determining the second spatial position information of the second text line in the document, and taking the first one-dimensional position information and the first spatial position information as the first layout information, and taking the second one-dimensional position information and the second spatial position information as the second layout information, the acquisition of the layout information of the text lines can be enriched to a large extent, so as to determine whether the first text line and the second text line are continuous line breaks by using the relatively rich layout information, enabling better learning of the global features of the text lines in the document and effectively improving the accuracy of the line-breaking determination process for the text lines in the document.

[0086] Figure 4 It is a schematic diagram according to the third embodiment of the present disclosure.

[0087] As Figure 4 shown, the document processing method includes:

[0088] S401: Obtain a document.

[0089] S402: Parse the first text line and the second text line from the document.

[0090] S403: Determine the first semantic information of the first text line, and determine the second semantic information of the second text line.

[0091] S404: Determine the first one-dimensional position information of the first text line in the document, and determine the second one-dimensional position information of the second text line in the document.

[0092] The description of S401 - S404 can be exemplified by referring to the above embodiments and will not be elaborated here.

[0093] S405: Determine multiple first characters of the first text line, and determine multiple second characters of the second text line.

[0094] Among them, the text lines in the document are composed of multiple characters. The characters that make up the first text line can be called the first characters. The first characters can be, for example, Chinese characters, English characters, punctuation marks, or other characters in the first text line. The multiple characters that make up the second text line can be called the second characters. The second characters can be, for example, Chinese characters, English characters, punctuation marks, or other characters in the second text line.

[0095] In the embodiments of the present disclosure, when determining multiple first characters of the first text line, the first text line can be processed by splitting and bounding to obtain multiple characters wrapped by a character bounding box after processing, and the multiple characters after the splitting and bounding processing are used as the multiple first characters of the first text line. The second text line is processed by splitting and bounding to obtain multiple characters wrapped by a character bounding box after processing, and the multiple characters after the splitting and bounding processing are used as the multiple second characters of the second text line.

[0096] S406: Obtain multiple first character space information respectively corresponding to multiple first characters, and obtain multiple second character space information respectively corresponding to multiple second characters.

[0097] Among them, the character space information is used to describe the position information of the character in the document. The position information can be represented by the two-dimensional spatial coordinates of the character bounding box where the character is located and the size information of the character bounding box. The position information of the first character in the document can be called the first character space information, and the position information of the second character in the document can be called the second character space information.

[0098] For example, the first character space information can be the two-dimensional spatial coordinate information of the character bounding box of the first character and the size information of the character bounding box, and the second character space information can be the two-dimensional spatial coordinate information of the character bounding box of the second character and the size information of the character bounding box.

[0099] In the embodiments of the present disclosure, when obtaining multiple first character space information respectively corresponding to multiple first characters, the two-dimensional spatial coordinates of the character bounding boxes where the multiple first characters are located and the size information of the character bounding boxes can be obtained respectively, and the corresponding multiple two-dimensional spatial coordinates and size information are used as the multiple first character space information respectively corresponding to the multiple first characters. The two-dimensional spatial coordinates of the character bounding boxes where the multiple second characters are located and the size information of the character bounding boxes are obtained respectively, and the corresponding multiple two-dimensional spatial coordinates and size information are used as the multiple second character space information respectively corresponding to the multiple second characters.

[0100] Optionally, in some embodiments, obtaining multiple first-character spatial information respectively corresponding to multiple first characters, and obtaining multiple second-character spatial information respectively corresponding to multiple second characters, can obtain multiple first-character positions respectively corresponding to multiple first characters, and obtain multiple second-character positions respectively corresponding to multiple second characters, obtain multiple first-character sizes respectively corresponding to multiple first characters, and obtain multiple second-character sizes respectively corresponding to multiple second characters, and use the first-character position and the first-character size as the first-character spatial information of the corresponding first character, and use the second-character position and the second-character size as the second-character spatial information of the corresponding second character, so that the character position and the character size of the character can be obtained as the character spatial information, introducing relatively comprehensive character spatial information, enabling more comprehensive character spatial features to be obtained, and thus enabling the depth combination with semantic information to determine whether the text line in the document is a continuous line break, assisting in improving the accuracy of the line break determination process for the text line in the document.

[0101] Among them, the first-character position can be used to represent the spatial position information of the character box where the first character is located in the document, and the second-character position can be used to represent the spatial position information of the character box where the second character is located in the document.

[0102] For example, the first-character position can be the upper-left corner coordinates and the lower-right corner coordinates of the character box where the first character is located, and the second-character position can be the upper-left corner coordinates and the lower-right corner coordinates of the character box where the second character is located.

[0103] Among them, the first-character size refers to the width and height of the character box where the first character is located, and the second-character size refers to the width and height of the character box where the second character is located.

[0104] In the embodiments of the present disclosure, when obtaining multiple first-character positions respectively corresponding to multiple first characters, and obtaining multiple second-character positions respectively corresponding to multiple second characters, the two-dimensional spatial position coordinates of the upper left corner and the two-dimensional spatial position coordinates of the lower right corner of the character boxes where multiple first characters are located can be obtained respectively, and the two-dimensional spatial position coordinates of the upper left corner and the two-dimensional spatial position coordinates of the lower right corner are used as the first-character positions of the corresponding first characters, and the two-dimensional spatial position coordinates of the upper left corner and the two-dimensional spatial position coordinates of the lower right corner of the character boxes where multiple second characters are located can be obtained respectively, and the two-dimensional spatial position coordinates of the upper left corner and the two-dimensional spatial position coordinates of the lower right corner are used as the second-character positions of the corresponding second characters.

[0105] In an embodiment of the present disclosure, when obtaining multiple first character sizes respectively corresponding to multiple first characters and obtaining multiple second character sizes respectively corresponding to multiple second characters, the width information and height information of the character frames where the multiple first characters are located may be obtained respectively, and the width information and height information are used as the first character sizes of the multiple first characters. The width information and height information of the character frames where the multiple second characters are located may be obtained respectively, and the width information and height information are used as the second character sizes of the multiple second characters.

[0106] For example, the two-dimensional spatial position coordinates (x0, y0) of the upper left corner and the two-dimensional spatial position coordinates (x1, y1) of the lower right corner of the character frames where the multiple first characters are located, as well as the width w and height h of the character frames where the multiple first characters are located, may be obtained respectively. And x0, y0, x1, y1, w, and h are used as the multiple first character spatial information respectively corresponding to the multiple first characters. The two-dimensional spatial position coordinates (x0, y0) of the upper left corner and the two-dimensional spatial position coordinates (x1, y1) of the lower right corner of the character frames where the multiple second characters are located, as well as the width w and height h of the character frames where the multiple second characters are located, may be obtained respectively. And x0, y0, x1, y1, w, and h are used as the multiple second character spatial information respectively corresponding to the multiple second characters.

[0107] After obtaining the multiple first character positions respectively corresponding to the multiple first characters and the multiple second character positions respectively corresponding to the multiple second characters, and obtaining the multiple first character sizes respectively corresponding to the multiple first characters and the multiple second character sizes respectively corresponding to the multiple second characters, the first character position and the first character size may be used as the first character spatial information of the corresponding first character, and the second character position and the second character size may be used as the second character spatial information of the corresponding second character.

[0108] S407: Determine the first spatial position information according to the multiple first character spatial information, and determine the second spatial position information according to the multiple second character spatial information.

[0109] Among them, the first spatial position information is used to represent the spatial position information of the first text line in the document, and the second spatial position information is used to represent the spatial position information of the second text line in the document.

[0110] For example, the first spatial position information may be the two-dimensional spatial position information of the first text line in the document and the first text line size information, and the second spatial position information may be the two-dimensional spatial position information of the second text line in the document and the second text line size information.

[0111] After taking the first character position and the first character size as the first character space information of the corresponding first character, and taking the second character position and the second character size as the second character space information of the corresponding second character, the first space position information can be determined according to the multiple first character space information of multiple characters, and the second space position information can be determined according to the multiple second character space information.

[0112] In an embodiment of the present disclosure, when determining the first space position information according to the multiple first character space information of multiple characters and determining the second space position information according to the multiple second character space information, the first character position and the first character size in the first character space information may be processed to obtain the first character space information of the first character, and the second character position and the second character size in the second character space information may be processed to obtain the second character space information of the second character.

[0113] Optionally, in some embodiments, the first space position information is determined according to the multiple first character space information, and the second space position information is determined according to the multiple second character space information. The multiple first character positions are spliced to obtain the first text position, the multiple second character positions are spliced to obtain the second text position, the multiple first character sizes are spliced to obtain the first text size, and the multiple second character sizes are spliced to obtain the second text size. And the first text position and the first text size are used as the first space position information, and the second text position and the second text size are used as the second space position information, so that the space position information of the text line can be determined according to the multiple character space information. Since the first space position information of the first text line and the second space position information of the second text line can be used to obtain the corresponding text features, the accuracy of obtaining text features can be assisted in improving.

[0114] Among them, the first text position is used to represent the position information of the first text line, the first text size is used to represent the size information of the first text line, the first text position and the first text size can be used to represent the space position information of the first text line in the document, the second text position is used to represent the position information of the second text line, the second text size is used to represent the size information of the second text line, and the second text position and the second text size can be used to represent the space position information of the second text line in the document.

[0115] In an embodiment of the present disclosure, when splicing the multiple first character positions to obtain the first text position and splicing the multiple second character positions to obtain the second text position, the multiple first character positions may be vectorially added to obtain the first text position, and the multiple second character positions may be vectorially added to obtain the second text position.

[0116] In an embodiment of the present disclosure, when splicing a plurality of first character sizes to obtain a first text size and splicing a plurality of second character sizes to obtain a second text size, vector splicing processing can be performed on the widths and heights of the character frames in the plurality of first character sizes to obtain the corresponding first text size, and vector splicing processing can be performed on the widths and heights of the character frames in the plurality of second character sizes to obtain the corresponding second text size.

[0117] In an embodiment of the present disclosure, after splicing a plurality of first character positions to obtain a first text position, splicing a plurality of second character positions to obtain a second text position, splicing a plurality of first character sizes to obtain a first text size, and splicing a plurality of second character sizes to obtain a second text size, the first text position and the first text size can be used as first spatial position information, and the second text position and the second text size can be used as second spatial position information.

[0118] For example, the two-dimensional spatial position coordinates (x0, y0) of the upper left corner and the two-dimensional spatial position coordinates (x1, y1) of the lower right corner of the character frame where the upper left corners of a plurality of first characters are located, as well as the width w and height h of the character frames where the plurality of first characters are located, can be respectively established in a mapping table, and learning processing can be performed on them through text pre-training. Corresponding representation vectors can be established for x0, y0, x1, y1, w, and h of the first characters. Vector addition processing can be performed on the position representation vectors of the plurality of first characters to obtain the first text position, and vector addition processing can be performed on the size representation vectors of the plurality of first characters to obtain the first text size.

[0119] In this embodiment, by determining a plurality of first characters in the first text line and determining a plurality of second characters in the second text line, obtaining a plurality of first character spatial information respectively corresponding to the plurality of first characters, obtaining a plurality of second character spatial information respectively corresponding to the plurality of second characters, determining first spatial position information according to the plurality of first character spatial information, and determining second spatial position information according to the plurality of second character spatial information, the spatial position information of the text line can be determined by using the character spatial information of a plurality of characters in the text line. At the same time, the comprehensiveness of the spatial position information obtained by processing the text from the spatial dimension is also enriched. The spatial position information of the text line can be used to extract the text features of the text line to determine whether the text line in the document is a continuous line break, so as to assist in improving the accuracy of the line break determination processing of the text line in the document.

[0120] S408: Determine the first spatial position information of the first text line in the document and determine the second spatial position information of the second text line in the document.

[0121] S409: Use the first one-dimensional position information and the first spatial position information as the first layout information, and use the second one-dimensional position information and the second spatial position information as the second layout information.

[0122] S410: Determine whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0123] For the descriptions of S408 - S410, refer to the above embodiments for examples and will not be elaborated here.

[0124] In this embodiment, by determining multiple first characters of the first text line and multiple second characters of the second text line, obtaining multiple first character spatial information respectively corresponding to the multiple first characters, obtaining multiple second character spatial information respectively corresponding to the multiple second characters, determining the first spatial position information according to the multiple first character spatial information, and determining the second spatial position information according to the multiple second character spatial information, the spatial position information of the text line can be determined by using the character spatial information of multiple characters in the text line. At the same time, it also enriches the comprehensiveness of the spatial position information obtained by processing the text from the spatial dimension. The spatial position information of the text line can be used for extracting the text features of the text line to determine whether the text line in the document is a consecutive line break, so as to assist in improving the accuracy of the line break determination process for the text line in the document. By determining the first one-dimensional position information of the first text line in the document, determining the second one-dimensional position information of the second text line in the document, determining the first spatial position information of the first text line in the document, determining the second spatial position information of the second text line in the document, using the first one-dimensional position information and the first spatial position information as the first layout information, and using the second one-dimensional position information and the second spatial position information as the second layout information, the acquisition of the layout information of the text line can be enriched to a large extent, so as to determine whether the first text line and the second text line are consecutive line breaks by using the relatively rich layout information, so that the global features of the text line in the document can be better learned, and the accuracy of the line break determination process for the text line in the document can be effectively improved.

[0125] Figure 5 It is a schematic diagram according to the fourth embodiment of the present disclosure.

[0126] As Figure 5 shown, the document processing method includes:

[0127] S501: Obtain a document.

[0128] S502: Parse the first text line and the second text line from the document.

[0129] S503: Determine the first semantic information of the first text line and determine the second semantic information of the second text line.

[0130] S504: Determine the first one-dimensional position information of the first text line in the document and determine the second one-dimensional position information of the second text line in the document.

[0131] S505: Determine the first spatial position information of the first text line in the document and determine the second spatial position information of the second text line in the document.

[0132] S506: Take the first one-dimensional position information and the first spatial position information as the first layout information, and take the second one-dimensional position information and the second spatial position information as the second layout information.

[0133] For the descriptions of S501 - S506, refer to the above embodiments for examples and will not be elaborated here.

[0134] S507: Perform a fusion process on the first semantic information, the first one-dimensional position information, and the first spatial position information to obtain the first text fusion feature of the first text line.

[0135] Among them, the first text fusion feature is used to characterize the semantic and spatial position features of the first text line.

[0136] In the embodiments of the present disclosure, when performing a fusion process on the first semantic information, the first one-dimensional position information, and the first spatial position information to obtain the first text fusion feature of the first text line, multiple first characters in the first text line can be concatenated in sequence and separated by a separator symbol, and then the first semantic information, the first one-dimensional position information, and the first spatial position information of the first text line are input into the Enhanced Language Representation with Informative Entities (ERNIE). Using the built-in multi-layer conversion structure in ERNIE, perform a fusion process on the first semantic information, the first one-dimensional position information, and the first spatial position information of the first text line to extract the first text fusion feature of the first text line.

[0137] S508: Perform a fusion process on the second semantic information, the second one-dimensional position information, and the second spatial position information to obtain the second text fusion feature of the second text line.

[0138] Among them, the second text fusion feature is used to characterize the semantic and spatial position features of the second text line.

[0139] In an embodiment of the present disclosure, when fusing the second semantic information, the second one-dimensional position information, and the second spatial position information to obtain the second text fusion feature of the second text line, multiple second characters in the second text line can be concatenated in sequence and separated by a separator symbol. Then, the second semantic information, the second one-dimensional position information, and the second spatial position information of the second text line are input into ERNIE, and the built-in multi-layer conversion structure of ERNIE is used to fuse the second semantic information, the second one-dimensional position information, and the second spatial position information of the second text line to extract the second text fusion feature of the second text line.

[0140] S509: Obtain the tail character representation feature in the first text line and obtain the head character representation feature in the second text line.

[0141] After obtaining the first fusion feature of the first text line and the second fusion feature of the second text line as described above, the tail character representation feature in the first text line can be obtained, and the head character representation feature in the second text line can be obtained.

[0142] In an embodiment of the present disclosure, when obtaining the tail character representation feature in the first text line and obtaining the head character representation feature in the second text line, the tail character in the first text line and the head character in the second text line can be found according to the separator between the first text line and the second text line. Then, the tail character identification feature of the first text line can be extracted, and the head character feature of the second text line can be extracted to obtain the tail character representation feature in the first text line and the head character representation feature in the second text line.

[0143] S510: Determine whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature.

[0144] After obtaining the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature as described above, it can be determined whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature.

[0145] In an embodiment of the present disclosure, when determining whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature, the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature can be concatenated and input into the linear layer of ERNIE. Then, after being processed by the classifier in the model, the final binary classification result is obtained, and this binary classification result is the judgment result of whether the first text line and the second text line are consecutive line breaks.

[0146] For example, as Figure 6 shown, Figure 6 is a schematic diagram of the text pre-training model structure in an embodiment of the present disclosure. Multiple first characters of the first text line and multiple second characters of the second text line input can be concatenated in sequence, and interval symbols are used for interval processing. The position information of each character is separately added and input into the text pre-training model ERNIE model one by one. Through the multi-layer conversion structure built in the text pre-training model ERNIE model, the features of the first text line and the second text line are further fused and extracted, and the tail character representation feature of the first text line and the head character representation feature of the second text line are extracted. Multiple features are concatenated and input into the linear layer of the text pre-training model, and a softmax classifier is used to obtain a binary classification result. This binary classification result is the judgment result on whether the first text line and the second text line are continuous line breaks.

[0147] Optionally, in some embodiments, to determine whether the first text line and the second text line are continuous line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature, the first text fusion feature and the second text fusion feature can be concatenated to obtain a target fusion feature. The target fusion feature, the tail character representation feature, and the head character representation feature are input into the first continuous line break recognition model to obtain a first determination result output by the first continuous line break recognition model. The first continuous line break recognition model includes: a first neural network structure for processing the target fusion feature, so that the target fusion feature, the tail character representation feature, and the head character identification feature can be used to judge whether the first text line and the second text line are continuous line breaks. Since the target fusion feature is obtained by concatenating the first text fusion feature and the second text fusion feature, the overall features of the first text line and the second text line can be globally modeled, and at the same time, the generality of the line break recognition processing model and the accuracy of line break processing can be effectively improved.

[0148] Among them, the first continuous line break recognition model is used to judge whether the first text line and the second text line are continuous line breaks, and the first continuous line break recognition model includes a first neural network structure for processing the target fusion feature.

[0149] In an embodiment of the present disclosure, when determining whether a first text line and a second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the end character representation feature, and the start character representation feature, the first text fusion feature and the second text fusion feature may be concatenated to obtain a target fusion feature after the concatenation process, and the target fusion feature, the end character representation feature, and the start character representation feature are input into a first consecutive line break recognition model. The first neural network structure in the first consecutive line break recognition model is used to process the target fusion feature, and the start character representation feature and the end character representation feature are jointly used for judgment processing to obtain the output result of the first consecutive line break recognition model as the first determination result. The first determination result is the judgment result of whether the first text line and the second text line are consecutive line breaks.

[0150] For example, as Figure 7 shown, Figure 7 is a schematic diagram of the structure of the first consecutive line break recognition model in an embodiment of the present disclosure. The first consecutive line break recognition model is a single-tower structure. The first text line and the second text line can be input into the model in sequence, and the text features of the first text line and the second text line are modeled. Then, the first text fusion feature and the second text fusion feature can be concatenated to obtain a target fusion feature after the concatenation process, and the target fusion feature, the end character representation feature, and the start character representation feature are input into a linear layer in the first consecutive line break recognition model for processing to obtain the output result of the classifier of the first consecutive line break recognition model as the first determination result. The first determination result is the judgment result of whether the first text line and the second text line are consecutive line breaks. The maximum text length that can be modeled by the first consecutive line break recognition model is set to N, and the scale of the attention mechanism matrix is N. 2 .

[0151] Optionally, in some embodiments, the first text fusion feature, the second text fusion feature, the end character representation feature, and the start character representation feature may be input into a second consecutive line break recognition model to obtain a second determination result output by the second consecutive line break recognition model. The second consecutive line break recognition model includes: a second neural network structure for processing the first text fusion feature and a third neural network structure for processing the second text fusion feature. Thus, the first text fusion feature and the second text fusion feature can be processed respectively by the second neural network structure and the third neural network structure according to the shared parameters, reducing the parameter scale of the consecutive line break recognition model, and effectively improving the line break determination processing efficiency of the consecutive line break recognition model.

[0152] Among them, the second consecutive line break recognition model is used to determine whether the first text line and the second text line are consecutive line breaks. The second consecutive line break recognition model includes a second neural network structure for processing the first text fusion feature and a third neural network structure for processing the second text fusion feature.

[0153] In the embodiments of the present disclosure, when determining whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the end character representation feature, and the start character representation feature, the first text fusion feature, the second text fusion feature, the end character representation feature, and the start character representation feature can be input into the second consecutive line break recognition model. The second neural network structure in the second consecutive line break recognition model is used to process the first text fusion feature, the third neural network structure in the second consecutive line break recognition model is used to process the second text fusion feature, and the end character representation feature and the start character representation feature are jointly used for judgment processing to obtain the second judgment result output by the second consecutive line break recognition model.

[0154] For example, as Figure 8 shown, Figure 8 is a schematic diagram of the structure of the second consecutive line break recognition model in the embodiments of the present disclosure. The second consecutive line break recognition model is a two-tower structure, which uses shared parameters to model the text features of the input first text line and second text line. The second neural network structure can be used to process the first text fusion feature respectively, the third neural network structure can be used to process the second text fusion feature, and the end character representation feature and the start character representation feature are jointly used for judgment to obtain the second judgment result of the second consecutive line break recognition model. The second judgment result is the judgment result of whether the first text line and the second text line are consecutive line breaks. The parameter scale of the second consecutive line break recognition model has a linear relationship with the length of the input text, which greatly reduces the number of parameters, improves the inference speed of the second consecutive line break recognition model, and facilitates better storage and deployment of the second consecutive line break recognition model in the production environment.

[0155] In this embodiment, by performing fusion processing on the first semantic information, the first one-dimensional position information, and the first spatial position information, the first text fusion feature of the first text line is obtained. By performing fusion processing on the second semantic information, the second one-dimensional position information, and the second spatial position information, the second text fusion feature of the second text line is obtained. The tail character representation feature in the first text line is obtained, and the head character representation feature in the second text line is obtained. Whether the first text line and the second text line are consecutive line breaks is determined according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature. Thus, the features of the first text line and the second text line can be modeled and extracted, and whether the first text line and the second text line are consecutive line breaks is determined by combining the tail character representation feature and the head character representation feature, and the deep combination processing of text semantic information and two-dimensional spatial position information can be realized, so that the accuracy of the line break determination process for the text lines in the document can be effectively improved.

[0156] In this embodiment, by determining multiple first characters of the first text line and multiple second characters of the second text line, obtaining multiple first character spatial information respectively corresponding to the multiple first characters, and obtaining multiple second character spatial information respectively corresponding to the multiple second characters, the first spatial position information is determined according to the multiple first character spatial information, and the second spatial position information is determined according to the multiple second character spatial information. Thus, the spatial position information of the text line can be determined by using the character spatial information of multiple characters in the text line, and at the same time, the comprehensiveness of the spatial position information obtained by processing the text from the spatial dimension is also enriched. The spatial position information of the text line can be used for the extraction of the text features of the text line to determine whether the text line in the document is a consecutive line break, so as to assist in improving the accuracy of the line break determination process for the text lines in the document. By performing fusion processing on the first semantic information, the first one-dimensional position information, and the first spatial position information, the first text fusion feature of the first text line is obtained. By performing fusion processing on the second semantic information, the second one-dimensional position information, and the second spatial position information, the second text fusion feature of the second text line is obtained. The tail character representation feature in the first text line is obtained, and the head character representation feature in the second text line is obtained. Whether the first text line and the second text line are consecutive line breaks is determined according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature. Thus, the features of the first text line and the second text line can be modeled and extracted, and whether the first text line and the second text line are consecutive line breaks is determined by combining the tail character representation feature and the head character representation feature, and the deep combination processing of text semantic information and two-dimensional spatial position information can be realized, so that the accuracy of the line break determination process for the document can be effectively improved.

[0157] Figure 9It is a schematic diagram according to the fifth embodiment of the present disclosure.

[0158] As Figure 9 shown, the document processing apparatus 90 includes:

[0159] An acquisition module 901 for acquiring a document;

[0160] An analysis module 902 for parsing a first text line and a second text line from the document;

[0161] A first determination module 903 for determining first semantic information of the first text line and determining second semantic information of the second text line;

[0162] A second determination module 904 for determining first layout information of the first text line and determining second layout information of the second text line; and

[0163] A third determination module 905 for determining whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

[0164] In some embodiments of the present disclosure, as Figure 10 shown, Figure 10 is a schematic diagram according to the sixth embodiment of the present disclosure. The document processing apparatus 100 includes: an acquisition module 1001, an analysis module 1002, a first determination module 1003, a second determination module 1004, and a third determination module 1005. Among them, the first determination module 1003 is specifically used for:

[0165] Determine the overall semantics of the document;

[0166] Determine the first local semantics of the first text line and determine the second local semantics of the second text line;

[0167] Obtain the first context semantics of the first local semantics relative to the overall semantics, and obtain the second context semantics of the second local semantics relative to the overall semantics; and

[0168] Take the first local semantics and the first context semantics as the first semantic information, and take the second local semantics and the second context semantics as the second semantic information.

[0169] In some embodiments of the present disclosure, among them, the second determination module 1004 includes:

[0170] A first determination sub-module 10041 for determining the first one-dimensional position information of the first text line in the document and determining the second one-dimensional position information of the second text line in the document;

[0171] The second determination sub-module 10042 is configured to determine the first spatial position information of the first text line in the document and determine the second spatial position information of the second text line in the document; and

[0172] The third determination sub-module 10043 is configured to use the first one-dimensional position information and the first spatial position information as the first layout information, and use the second one-dimensional position information and the second spatial position information as the second layout information.

[0173] In some embodiments of the present disclosure, the second determination sub-module 10042 is specifically configured to:

[0174] Determine a plurality of first characters of the first text line and determine a plurality of second characters of the second text line;

[0175] Obtain a plurality of first character spatial information respectively corresponding to the plurality of first characters, and obtain a plurality of second character spatial information respectively corresponding to the plurality of second characters; and

[0176] Determine the first spatial position information according to the plurality of first character spatial information, and determine the second spatial position information according to the plurality of second character spatial information.

[0177] In some embodiments of the present disclosure, the second determination sub-module 10042 is further configured to:

[0178] Obtain a plurality of first character positions respectively corresponding to the plurality of first characters, and obtain a plurality of second character positions respectively corresponding to the plurality of second characters;

[0179] Obtain a plurality of first character sizes respectively corresponding to the plurality of first characters, and obtain a plurality of second character sizes respectively corresponding to the plurality of second characters; and

[0180] Use the first character position and the first character size as the first character spatial information of the corresponding first character, and use the second character position and the second character size as the second character spatial information of the corresponding second character.

[0181] In some embodiments of the present disclosure, the second determination sub-module 10042 is further configured to:

[0182] Perform a splicing process on the plurality of first character positions to obtain a first text position, and perform a splicing process on the plurality of second character positions to obtain a second text position;

[0183] Perform a splicing process on the plurality of first character sizes to obtain a first text size, and perform a splicing process on the plurality of second character sizes to obtain a second text size; and

[0184] Take the first text position and the first text size as the first spatial position information, and take the second text position and the second text size as the second spatial position information.

[0185] In some embodiments of the present disclosure, the third determination module 1005 includes:

[0186] The first fusion sub-module 10051 is configured to perform fusion processing on the first semantic information, the first one-dimensional position information, and the first spatial position information to obtain the first text fusion feature of the first text line;

[0187] The second fusion sub-module 10052 is configured to perform fusion processing on the second semantic information, the second one-dimensional position information, and the second spatial position information to obtain the second text fusion feature of the second text line;

[0188] The acquisition sub-module 10053 is configured to acquire the tail character representation feature in the first text line and acquire the head character representation feature in the second text line; and

[0189] The fourth determination sub-module 10054 determines whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature.

[0190] In some embodiments of the present disclosure, the fourth determination sub-module 10054 is specifically configured to:

[0191] Perform splicing processing on the first text fusion feature and the second text fusion feature to obtain a target fusion feature;

[0192] Input the target fusion feature, the tail character representation feature, and the head character representation feature into the first consecutive line break recognition model to obtain a first determination result output by the first consecutive line break recognition model;

[0193] Wherein, the first consecutive line break recognition model includes: a first neural network structure for processing the target fusion feature.

[0194] In some embodiments of the present disclosure, the fourth determination sub-module 10054 is further configured to:

[0195] Input the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature into the second consecutive line break recognition model to obtain a second determination result output by the second consecutive line break recognition model;

[0196] Wherein, the second consecutive line break recognition model includes: a second neural network structure for processing the first text fusion feature and a third neural network structure for processing the second text fusion feature.

[0197] It can be understood that the document processing device 100 in this embodiment Figure 10 and the document processing device 90 in the above embodiment, the acquisition module 1001 and the acquisition module 901 in the above embodiment, the parsing module 1002 and the parsing module 902 in the above embodiment, the first determination module 1003 and the first determination module 903 in the above embodiment, the second determination module 1004 and the second determination module 904 in the above embodiment, and the third determination module 1005 and the third determination module 905 in the above embodiment may have the same functions and structures.

[0198] It should be noted that the foregoing explanations of the document processing method also apply to the document processing device in this embodiment, and will not be elaborated herein.

[0199] In this embodiment, by acquiring a document, parsing the first text line and the second text line from the document, determining the first semantic information of the first text line, and determining the second semantic information of the second text line, determining the first layout information of the first text line, and determining the second layout information of the second text line, and determining whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information, it is possible to fully combine the semantic information of the text lines with the layout information of the text lines in the document to determine whether different text lines are consecutive line breaks, which can effectively improve the determination accuracy of whether different text lines in the document are consecutive line breaks and effectively improve the determination efficiency of consecutive line breaks for the document.

[0200] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0201] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0202] As Figure 11As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 1102 or computer programs loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0203] Multiple components in device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0204] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the document processing method. For example, in some embodiments, the document processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the document processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the document processing method in any other appropriate manner (e.g., by means of firmware).

[0205] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0206] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0207] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0208] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0209] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0210] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0211] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0212] The above specific embodiments do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A document processing method, comprising: Obtaining a document; Parsing a first text line and a second text line from the document; Determining first semantic information of the first text line and determining second semantic information of the second text line; Determining first one-dimensional position information of the first text line in the document and determining second one-dimensional position information of the second text line in the document; the one-dimensional position information is used to represent the number information of the text line in the text line sequence of the document or the layout module position of the text line in the document; Determining first spatial position information of the first text line in the document and determining second spatial position information of the second text line in the document, the spatial position information being determined by the spatial position information of each character in the text line; And Taking the first one-dimensional position information and the first spatial position information as first layout information, and taking the second one-dimensional position information and the second spatial position information as second layout information; And Determining whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information.

2. The method according to claim 1, wherein, The determining the first semantic information of the first text line and determining the second semantic information of the second text line includes: Determining the overall semantics of the document; Determining first local semantics of the first text line and determining second local semantics of the second text line; Obtaining first context semantics of the first local semantics relative to the overall semantics and obtaining second context semantics of the second local semantics relative to the overall semantics; and Taking the first local semantics and the first context semantics as the first semantic information, and taking the second local semantics and the second context semantics as the second semantic information.

3. The method according to claim 1, wherein The determining the first spatial position information of the first text line in the document and determining the second spatial position information of the second text line in the document includes: Determining a plurality of first characters of the first text line and determining a plurality of second characters of the second text line; Obtaining a plurality of first character spatial information respectively corresponding to the plurality of first characters and obtaining a plurality of second character spatial information respectively corresponding to the plurality of second characters; and Determining the first spatial position information according to the plurality of first character spatial information and determining the second spatial position information according to the plurality of second character spatial information.

4. The method according to claim 3, wherein, The obtaining the plurality of first character spatial information respectively corresponding to the plurality of first characters and obtaining the plurality of second character spatial information respectively corresponding to the plurality of second characters includes: Obtaining a plurality of first character positions respectively corresponding to the plurality of first characters and obtaining a plurality of second character positions respectively corresponding to the plurality of second characters; Obtaining a plurality of first character sizes respectively corresponding to the plurality of first characters and obtaining a plurality of second character sizes respectively corresponding to the plurality of second characters; and Take the first character position and the first character size as the first character spatial information corresponding to the first character, and take the second character position and the second character size as the second character spatial information corresponding to the second character.

5. The method according to claim 4, wherein The determining the first spatial position information according to the multiple first character spatial information and determining the second spatial position information according to the multiple second character spatial information includes: Performing a splicing process on the multiple first character positions to obtain a first text position, and performing a splicing process on the multiple second character positions to obtain a second text position; Performing a splicing process on the multiple first character sizes to obtain a first text size, and performing a splicing process on the multiple second character sizes to obtain a second text size; and Taking the first text position and the first text size as the first spatial position information, and taking the second text position and the second text size as the second spatial position information.

6. The method according to claim 1, wherein The determining whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information includes: Performing a fusion process on the first semantic information, the first one-dimensional position information, and the first spatial position information to obtain a first text fusion feature of the first text line; Performing a fusion process on the second semantic information, the second one-dimensional position information, and the second spatial position information to obtain a second text fusion feature of the second text line; Obtaining a tail character representation feature in the first text line, and obtaining a head character representation feature in the second text line; and Determining whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature.

7. The method according to claim 6, wherein, The determining whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature includes: Performing a splicing process on the first text fusion feature and the second text fusion feature to obtain a target fusion feature; Inputting the target fusion feature, the tail character representation feature, and the head character representation feature into a first consecutive line break recognition model to obtain a first determination result output by the first consecutive line break recognition model; Wherein, the first consecutive line break recognition model includes: a first neural network structure for processing the target fusion feature.

8. The method according to claim 6, wherein The determining whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature includes: Input the first text fusion feature, the second text fusion feature, the end character representation feature, and the start character representation feature into a second consecutive line break recognition model to obtain a second determination result output by the second consecutive line break recognition model; Wherein, the second consecutive line break recognition model includes: a second neural network structure for processing the first text fusion feature, and a third neural network structure for processing the second text fusion feature.

9. A document processing device, comprising: An acquisition module for acquiring a document; An analysis module for analyzing a first text line and a second text line from the document; A first determination module for determining first semantic information of the first text line and determining second semantic information of the second text line; A second determination module for determining first layout information of the first text line and determining second layout information of the second text line; And A third determination module for determining whether the first text line and the second text line are consecutive line breaks according to the first semantic information, the first layout information, the second semantic information, and the second layout information; Wherein, the second determination module includes: A first determination sub-module for determining first one-dimensional position information of the first text line in the document and determining second one-dimensional position information of the second text line in the document; A second determination sub-module for determining first spatial position information of the first text line in the document and determining second spatial position information of the second text line in the document; and A third determination sub-module for using the first one-dimensional position information and the first spatial position information as the first layout information, and using the second one-dimensional position information and the second spatial position information as the second layout information.

10. The apparatus according to claim 9, wherein, The first determination module is specifically configured to: Determine the overall semantics of the document; Determine first local semantics of the first text line and determine second local semantics of the second text line; Obtain first context semantics of the first local semantics relative to the overall semantics and obtain second context semantics of the second local semantics relative to the overall semantics; and Use the first local semantics and the first context semantics as the first semantic information, and use the second local semantics and the second context semantics as the second semantic information.

11. The device according to claim 9, wherein The second determination sub-module is specifically configured to: Determine a plurality of first characters of the first text line and determine a plurality of second characters of the second text line; Obtain a plurality of first character space information respectively corresponding to the plurality of first characters and obtain a plurality of second character space information respectively corresponding to the plurality of second characters; And Determine the first spatial position information according to the plurality of first character space information and determine the second spatial position information according to the plurality of second character space information.

12. The apparatus according to claim 11, wherein, The second determination sub-module is further configured to: Obtain a plurality of first character positions respectively corresponding to the plurality of first characters and obtain a plurality of second character positions respectively corresponding to the plurality of second characters; Obtain a plurality of first character sizes respectively corresponding to the plurality of first characters, and obtain a plurality of second character sizes respectively corresponding to the plurality of second characters; And Take the first character position and the first character size as the first character spatial information corresponding to the corresponding first character, and take the second character position and the second character size as the second character spatial information corresponding to the corresponding second character.

13. The apparatus according to claim 12, wherein, The second determination sub-module is further configured to: Perform splicing processing on the plurality of first character positions to obtain a first text position, and perform splicing processing on the plurality of second character positions to obtain a second text position; Perform splicing processing on the plurality of first character sizes to obtain a first text size, and perform splicing processing on the plurality of second character sizes to obtain a second text size; and Take the first text position and the first text size as the first spatial position information, and take the second text position and the second text size as the second spatial position information.

14. The device according to claim 9, wherein, The third determination module includes: A first fusion sub-module, configured to perform fusion processing on the first semantic information, the first one-dimensional position information, and the first spatial position information to obtain a first text fusion feature of the first text line; A second fusion sub-module, configured to perform fusion processing on the second semantic information, the second one-dimensional position information, and the second spatial position information to obtain a second text fusion feature of the second text line; An acquisition sub-module, configured to acquire a tail character representation feature in the first text line and acquire a head character representation feature in the second text line; and A fourth determination sub-module, which determines whether the first text line and the second text line are consecutive line breaks according to the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature.

15. The apparatus according to claim 14, wherein The fourth determination sub-module is specifically configured to: Perform splicing processing on the first text fusion feature and the second text fusion feature to obtain a target fusion feature; Input the target fusion feature, the tail character representation feature, and the head character representation feature into a first consecutive line break recognition model to obtain a first determination result output by the first consecutive line break recognition model; Wherein, the first consecutive line break recognition model includes: a first neural network structure for processing the target fusion feature.

16. The device according to claim 14, wherein, The fourth determination sub-module is further configured to: Input the first text fusion feature, the second text fusion feature, the tail character representation feature, and the head character representation feature into a second consecutive line break recognition model to obtain a second determination result output by the second consecutive line break recognition model; Wherein, the second consecutive line break recognition model includes: a second neural network structure for processing the first text fusion feature and a third neural network structure for processing the second text fusion feature.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product, comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Text detection method, device and equipment and storage medium

    CN110929647A