Text processing method, apparatus and system
By recognizing and sorting the text in text images to generate text data, the problem that OCR technology cannot directly input text data is solved, realizing the automated serialization of text images and improving input efficiency.
Patent Information
- Application Number
- CN202110163230.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-05
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Existing OCR technology can only recognize single characters and cannot directly input them, resulting in low text processing efficiency and requiring manual input of single characters.
By recognizing the positional and content information of multiple characters in a text image, the sorting relationship between the characters is determined, and text data is generated based on this, thus achieving automated text serialization and input.
It eliminates the need for manual input of individual characters, enabling automated sequential input of text and images, thus improving input efficiency.
Smart Images

Figure CN114882517B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image recognition, in particular, to a text processing method, device and system. BACKGROUND
[0002] In the process of "text electrization", it is necessary to convert a text image into a whole paragraph of text data entry system. The text can be a paper media text such as a poster, a newspaper, a magazine, and the like, a promotional single text such as a product promotional single page in an e-commerce platform, a book text such as a textbook and a classic, but is not limited to these. At present, the text image can be recognized by using OCR (Optical Character Recognition). However, the OCR can only detect the position of a single character and recognize the content of the character, and cannot directly enter the text, so it is necessary to manually enter the single character, which results in low entry efficiency.
[0003] At present, there is no effective solution to the above problems. SUMMARY
[0004] Embodiments of the present application provide a text processing method, device and system to at least solve the technical problem of the related art text processing method that can only recognize a single character and manually enter the single character, resulting in low entry efficiency.
[0005] According to an aspect of an embodiment of the present application, a text processing method is provided, including: obtaining a text image; recognizing the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and generating text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0006] According to another aspect of an embodiment of the present application, a text processing method is provided, including: displaying a text image; marking recognition results of a plurality of characters in the text image in the text image, wherein the recognition results of the plurality of characters are obtained by recognizing the text image, and the recognition results include content information corresponding to each character and position information of each character in the text image; and displaying text data corresponding to the text image, wherein the text data is generated based on a sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters, and the sorting relationship between the plurality of characters is determined based on the position information of the plurality of characters.
[0007] According to another aspect of the embodiments of the present application, a text processing method is also provided, including: receiving a text image; identifying the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; generating text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters; and outputting the text data.
[0008] According to another aspect of the embodiments of the present application, a text processing method is also provided, including: receiving a text image; identifying the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; generating text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters; and outputting the text data.
[0009] According to another aspect of the embodiments of the present application, a text processing apparatus is also provided, including: an acquisition module configured to acquire a text image; an identification module configured to identify the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; a determination module configured to determine a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and a generation module configured to generate text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0010] According to another aspect of the embodiments of the present application, a text processing apparatus is also provided, including: a first display module configured to display a text image; a marking module configured to mark recognition results of a plurality of characters in the text image in the text image, wherein the recognition results of the plurality of characters are obtained by identifying the text image, and the recognition results include content information corresponding to each character and position information of each character in the text image; and a second display module configured to display text data corresponding to the text image, wherein the text data is generated based on a sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters, and the sorting relationship between the plurality of characters is determined based on the position information of the plurality of characters.
[0011] According to another aspect of the embodiments of the present application, a text processing apparatus is also provided, comprising: a receiving module configured to receive a text image; an identifying module configured to identify the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results comprise content information corresponding to each character and position information of each character in the text image; a determining module configured to determine a sorting relationship between the plurality of characters based on the position information of the plurality of characters; a generating module configured to generate text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters; and an outputting module configured to output the text data.
[0012] According to another aspect of the embodiments of the present application, a text processing apparatus is also provided, comprising: an obtaining module configured to obtain an ancient book image; an identifying module configured to identify the ancient book image to obtain recognition results of a plurality of characters in the ancient book image, wherein the recognition results comprise content information corresponding to each character and position information of each character in the ancient book image; a determining module configured to determine a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and a generating module configured to generate ancient book text corresponding to the ancient book image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0013] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, comprising a stored program, wherein the program, when executed, controls a device in which the computer readable storage medium is located to perform the text processing method described above.
[0014] According to another aspect of the embodiments of the present application, a computer terminal is also provided, comprising a memory and a processor, wherein the processor is configured to execute a program stored in the memory, and the program, when executed, performs the text processing method described above.
[0015] According to another aspect of the embodiments of the present application, a text processing system is also provided, comprising: a processor; and a memory connected to the processor and configured to provide the processor with instructions for processing the following processing steps: obtaining a text image; identifying the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results comprise content information corresponding to each character and position information of each character in the text image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and generating text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0016] In the embodiments of the present application, after the text image is acquired, the text image can be recognized to obtain the recognition results of multiple characters in the text image, and based on the position information in the recognition results, the sorting relationship between the multiple characters is determined, and further based on the sorting relationship between the multiple characters and the content information in the recognition results, the text data corresponding to the text image is generated, so as to realize the purpose of single character serialization of the text image. It is easy to note that after the content information and the position information of each character are recognized, the recognized single characters can be sorted to generate the corresponding text data, so as to realize the automatic input of the serialized text, without manually inputting the single characters, so as to improve the text serialization effect and the input efficiency, and further solve the technical problem that the related art text processing method can only recognize single characters and manually input the single characters, resulting in low input efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a text processing method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of a first text processing method according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of an optional interactive interface according to an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of an optional "big character" in an ancient book according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of an optional adjacent region according to an embodiment of the present application;
[0023] Figure 6 is a schematic diagram of an optional generated serialized text according to an embodiment of the present application;
[0024] Figure 7 is a flowchart of an optional text processing method according to an embodiment of the present application;
[0025] Figure 8 is a flowchart of a second text processing method according to an embodiment of the present application;
[0026] Figure 9is a flow chart of a third text processing method according to an embodiment of the application;
[0027] Figure 10 is a schematic diagram of a first text processing device according to an embodiment of the application;
[0028] Figure 11 is a schematic diagram of a second text processing device according to an embodiment of the application;
[0029] Figure 12 is a schematic diagram of a third text processing device according to an embodiment of the application;
[0030] Figure 13 is a flow chart of a fourth text processing method according to an embodiment of the application;
[0031] Figure 14 is a schematic diagram of a fourth text processing device according to an embodiment of the application;
[0032] Figure 15 is a structural block diagram of a computer terminal according to an embodiment of the application. DETAILED DESCRIPTION
[0033] In order to make the personnel in the art better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0036] Ancient books: can refer to books that are not printed by modern printing technology, and the layout mode is different from that of modern books, newspapers and periodicals. The layout mode of modern books, newspapers and periodicals is "from top to bottom, from left to right", and the layout mode of ancient books is "from top to bottom, from right to left".
[0037] OCR: can refer to the process in which an electronic device (such as a scanner or digital camera) checks characters printed on paper and then translates shapes into computer text by character recognition method. However, for character recognition, only single-character recognition can be performed.
[0038] Embodiment 1
[0039] According to the embodiment of the present application, a text processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0040] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar operation device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the text processing method is shown. As shown in Figure 1 , the computer terminal 10 (or mobile device 10) can include one or more (in the figure, 102a, 102b, …, 102n are used to show) processors 102 (the processor 102 can include but not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or less components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0041] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry." The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any of the other elements of the computer terminal 10 (or mobile device). As referred to in the embodiments of the present application, the data processing circuitry functions as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.
[0042] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the text processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the text processing method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0043] The transmission device 106 is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet in a wireless manner.
[0044] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0045] It should be noted that in some optional embodiments, the above-mentioned Figure 1 The computer device (or mobile device) shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that Figure 1is merely one instance of a specific, concrete example, and is intended to show the types of components that can be present in the above-described computer device (or mobile device).
[0046] In the above operating environment, the present application provides a text processing method as shown in Figure 2 Figure 2 is a flowchart of a first text processing method according to an embodiment of the present application. As shown in Figure 2 the method comprises the following steps:
[0047] Step S202, obtaining a text image.
[0048] The text image in the above steps can be an image of different types of text, which can be directly obtained by photographing different types of text, or can also be obtained by intercepting video frames. Here, the video is a video taken during reading of different types of text. For example, the text image can be an image obtained by photographing a single page of text such as a poster or a leaflet, or can be an image obtained by photographing each page of text in a multi-page text such as a newspaper, a magazine, or a book. It can also be a video taken during reading of a multi-page text such as a newspaper, a magazine, or a book, which contains the content of each page of text. In the present embodiment, an ancient book image is taken as an example for illustration.
[0049] The text image contains all the text in the text, and the text in different types of text is arranged according to different layout methods. It should be noted that in a text image, multiple different layout methods can be used, for example, for a poster image, both "from top to bottom, from left to right" and "from top to bottom, from right to left" layout methods can be used. In addition, for multi-page text, different layout methods can be used in different text images.
[0050] In an optional embodiment, in order to realize the electronicization of the text, the text can be photographed, and the photographed text image can be transmitted to a corresponding processing device for processing, for example, directly transmitted to a user's computer terminal (such as a notebook computer, a personal computer, etc.) for processing, or transmitted to a cloud server through the user's computer terminal for processing. It should be noted that since the processing of the text image requires a large amount of computing resources, in the present embodiment, the processing device is taken as a cloud server for illustration.
[0051] For example, in order to facilitate the user to upload the text image, an interactive interface can be provided to the user, as shown in Figure 3 As shown, the user can select the text image to be processed from a large number of stored images in turn by clicking the "Select Image" button, or select multiple text images in batch, and upload the selected text image to the cloud server for processing by clicking the "Upload" button. In addition, in order to facilitate the user to confirm whether the selected text image is the text image to be processed, the user-selected text image can be displayed in the "Image Display" area, and after the user confirms that it is correct, the data is uploaded by clicking the "Upload" button.
[0052] In step S204, the text image is recognized to obtain the recognition result of the plurality of characters in the text image, wherein the recognition result includes: content information corresponding to each character, and position information of each character in the text image.
[0053] Since different types of text images adopt different typesetting methods, and a text image may adopt multiple different typesetting methods, in order to accurately recognize each character in the text image, in an optional embodiment, an OCR technology can be used to recognize each character in the text image to recognize the specific character content and the specific position of each character in the text image.
[0054] In the embodiment of the present application, the specific position of each character can be represented by the starting point coordinates (including the horizontal coordinate and the vertical coordinate, i.e. the x coordinate and the y coordinate), the width and the height, and therefore, the position information in the above step can include the starting point x coordinate, the starting point y coordinate, the width and the height. The starting point here can refer to the lower left corner pixel point of the character, and the coordinate origin can be the lower left corner pixel point of the text image, but is not limited to this, and can be determined according to the calculation needs.
[0055] In step S206, the sorting relationship between the plurality of characters is determined based on the position information of the plurality of characters.
[0056] The sorting relationship in the above step can refer to the order of different characters in the reading process, for example, assuming that the sorting relationship between character a and character b is that character a is arranged in front of character b, then in the reading process, the user first reads character a and then reads character b.
[0057] It should be noted that, in order to accurately enter all the recognized characters, it is necessary to first determine the reading order between different characters, and then enter different characters according to the reading order to obtain the electronic data corresponding to the text image. In an optional embodiment, the position relationship between two characters can be analyzed based on the position information of all the characters, and then the sorting relationship between all the characters can be determined by synthesizing all the two-by-two position relationships. In another optional embodiment, in order to simplify the analysis process, a text order prediction scheme based on deep learning can be used to input the position information and content information of each character into a neural network model for prediction, and a partial order relationship matrix of all characters can be obtained, that is, the sorting relationship between all characters.
[0058] In step S208, text data corresponding to the text image is generated based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0059] The text data in the above steps can refer to a serialized text obtained by sorting the content information of all characters according to the sorting relationship.
[0060] In an optional embodiment, in order to realize text entry, after identifying the content information of all characters in the text image, a single character needs to be composed into a complete serialized text. Therefore, after analyzing the sorting relationship between all characters, the content information can be sorted according to the sorting relationship to obtain a serialized text (i.e., the above-mentioned text data). In addition, in order to facilitate users to read the electronic text, the currently common reading order can be used to layout all the characters, and the order of all the characters is still determined based on the sorting relationship between the characters, or the user's preferred reading order can be determined according to the user's reading habits, and the characters can be laid out using the user's preferred reading order, and the order of all the characters is still determined based on the sorting relationship between the characters.
[0061] For example, for paper news media text, taking a news paper as an example, the user can take a photo of each page of the news paper, and upload the paper image of each page to the cloud server for processing in turn. In order to realize the electrization of the news paper, the cloud server can first utilize the OCR technology to recognize each character in the paper image, obtain the recognition result of all characters, further combine the characteristics of irregular text of the news paper, that is, the same page of text can be laid out in multiple different layout ways, can divide all characters in the paper image into multiple character blocks according to different layout ways, for each character block, the sorting relationship of all characters in the character block can be determined according to the corresponding layout way, and then the text data corresponding to each character block is obtained, and finally the text data corresponding to all character blocks is summarized, and the text data corresponding to the news paper can be obtained, that is, the final electrization data is obtained.
[0062] For example, for single-page text such as posters and flyers, taking a flyer page as an example, the user can directly take a photo of the flyer page, and upload the flyer image to the cloud server for processing in turn. In order to realize the electrization of the flyer page, the cloud server can first utilize the OCR technology to recognize each character in the flyer image, obtain the recognition result of all characters, further combine the layout characteristics of the flyer page, and determine the sorting relationship of all characters in the flyer image, and then obtain the text data corresponding to the flyer image, that is, the final electrization data is obtained.
[0063] For example, for book text, taking an ancient book as an example, the user can take a photo of each page of the ancient book, and upload the ancient book image to the cloud server for processing in turn. In order to realize the electrization of the ancient book, the cloud server can first utilize the OCR technology to recognize each character in the ancient book image, obtain the recognition result of all characters, further combine the layout characteristics of the ancient book, that is, the layout way of “from top to bottom, from right to left”, and determine the sorting relationship of all characters in the ancient book image, and then obtain the text data corresponding to the ancient book image, that is, the final electrization data is obtained.
[0064] By the scheme provided in the above embodiments of the present application, after obtaining the text image, the text image can be recognized to obtain the recognition result of the plurality of characters in the text image, and based on the position information in the recognition result, the sorting relationship between the plurality of characters is determined, and further based on the sorting relationship between the plurality of characters and the content information in the plurality of recognition results, the text data corresponding to the text image is generated, realizing the purpose of single character serialization of the text image. It is easy to note that after recognizing the content information and the position information of each character, the recognized single character can be sorted to generate the corresponding text data, realizing the automatic input of the serialized text, without manually inputting the single character, achieving the technical effects of improving the text serialization effect and improving the input efficiency, and further solving the technical problems that the related art text processing method can only recognize single characters, and manual input of single characters is required, resulting in low input efficiency.
[0065] In the above embodiments of the present application, determining the sorting relationship between the plurality of characters based on the position information of the plurality of characters comprises: determining the position relationship of any two characters based on the position information of the any two characters; determining the sorting relationship of the any two characters based on the position relationship of the any two characters and the first typesetting mode corresponding to the text image; and summarizing the sorting relationship of the any two characters to obtain the sorting relationship between the plurality of characters.
[0066] The position relationship in the above step can be "top left, top, top right, left, right, bottom left, bottom, bottom right" and the like.
[0067] The first typesetting mode in the above step can be the typesetting mode corresponding to different types of text, for example, taking an ancient book as an example, the first typesetting mode can be "from top to bottom, from right to left".
[0068] In an alternative embodiment, all the characters can be analyzed for positional relationship two by two. Assuming that the starting point is the lower left corner and the coordinate origin is the lower left corner of the text image, a larger x-coordinate indicates that the character is more to the right, and a larger y-coordinate indicates that the character is more to the top. Therefore, the x-coordinate and y-coordinate of the starting point of each character can be used to determine the positional relationship. The specific determination conditions are as follows: taking the position information of character a (x1, y1, w1, h1) and the position information of character b (x2, y2, w2, h2) as an example, the condition for determining that character b is to the left of character a can be (x2 + 0.5*w2) < (x1 + 0.5*w1) and the vertical overlap is less than a certain threshold (e.g. 0.1); the condition for determining that character b is to the right of character a can be (x1 + 0.5*w1) < (x2 + 0.5*w2) and the vertical overlap is less than a certain threshold (e.g. 0.1); the condition for determining that character b is above character a can be (y1 + 0.5*h1) < (y2 + 0.5*h2) and the vertical overlap is less than a certain threshold (e.g. 0.1); the condition for determining that character b is below character a can be (y1 + 0.5*h1) < (y2 + 0.5*h2) and the vertical overlap is less than a certain threshold (e.g. 0.1); and if character b is not to the left of character a and not to the right of character b, it is determined that character b is in the middle. The positional relationship "above" can mean "middle above", the positional relationship "below" can mean "middle below", the positional relationship "left" can mean "middle left", and the positional relationship "right" can mean "middle right". Therefore, the positional relationship between two characters can be determined based on the above conditions.
[0069] Further, the sorting relationship between two characters can be determined according to the positional relationship, and the sorting relationship between different characters can be determined in combination with the layout of different texts. For example, for the layout of ancient books, if the positional relationship between character a and character b is any one of "above, right above, right, right below", etc., it is determined that character a is in front of character b; if the positional relationship between character a and character b is any one of "left above, left, left below, below", etc., it is determined that character b is in front of character a.
[0070] Finally, the sorting relationship between all characters can be determined by synthesizing the sorting relationship between two characters.
[0071] For example, still taking ancient books as an example, for characters a, b and c, assuming that the positional relationship between character b and character a is "above", the positional relationship between character c and character a is "left above", and the positional relationship between character b and character c is "right above", it can be determined that character b is in front of character a, character a is in front of character c, and character b is in front of character c. Therefore, by summarizing, the sorting relationship of all characters can be obtained as: character b, character a, character c.
[0072] It should be noted that after the sorting relationship between the plurality of characters is determined, the sorting relationship can be saved, so that it can be used again for the recognition of a text image with similar sorting relationship. In addition, for a text with damage or loss, the damaged or lost character can be given by a user, and the text can be restored according to the determined sorting relationship between the plurality of characters.
[0073] In the above embodiments of the present application, before the sorting relationship between any two characters is determined based on the position relationship between the two characters and the first typesetting manner corresponding to the text image, the method further includes: determining whether the plurality of characters contains a target character based on the position information of the plurality of characters, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; in the case where the plurality of characters contains the target character, determining the sorting relationship between any two characters based on the position relationship between any two first characters, the position relationship between the first character and the target character, the position relationship between any two target characters, and the first typesetting manner, wherein the first character is a character in the plurality of characters located in the vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; in the case where the plurality of characters does not contain the target character, determining the sorting relationship between any two characters based on the position relationship between the two characters and the first typesetting manner.
[0074] The target character in the above step can refer to a "big character" in the text image, for example, as shown in Figure 4 The width of the big character is often twice that of the ordinary character, and the reading order is to read the two columns of characters above the big character first, then read the big character, and finally read the two columns of characters below the big character. Therefore, when the text image contains a big character, the sorting relationship between the characters cannot be determined based on the position relationship alone, and the position information of the big character needs to be combined to determine the sorting relationship.
[0075] In an optional embodiment, when analyzing the position relationship between each two characters, it is first necessary to determine whether the text image contains a "big character". If no big character is contained, the sorting relationship between each two characters is determined directly based on the position relationship between the two characters. If a big character is contained, for other characters unrelated to the big character, the sorting relationship between each two characters can be determined directly based on the position relationship between the two characters, and for characters related to the big character, i.e., characters located above or below the big character, the sorting relationship between each two characters needs to be determined based on the position relationship between the two characters and the position relationship between each character and the big character.
[0076] For example, still taking the ancient book as an example, for the characters a, b, c, d, e, f and g, it is assumed that the character a is a large character, the character b has a positional relationship of “above” with the character a, the character c has a positional relationship of “above” with the character a, the character b has a positional relationship of “right” with the character c, at this time, the character b is arranged in front of the character c, and the character c is arranged in front of the character a; the character d has a positional relationship of “below” with the character a, the character e has a positional relationship of “below” with the character a, the character d has a positional relationship of “left” with the character e, at this time, the character a is arranged in front of the character e, and the character e is arranged in front of the character d; the character f has a positional relationship of “left above” with the character a, and the character g has a positional relationship of “left below” with the character a, at this time, the character f is arranged in front of the character g, and the character a is arranged in front of the character f. Based on the above analysis, it can be determined that the sorting relationship of the above seven characters is: the character b, the character c, the character a, the character e, the character d, the character f and the character g.
[0077] In the above embodiment of the present application, determining whether the target character is included in the plurality of characters based on the position information of the plurality of characters comprises: obtaining the number of characters adjacent to each character in the vertical direction based on the position information of the plurality of characters, to obtain the character number corresponding to each character; in a case where the character number corresponding to the second character is greater than or equal to the preset number, determining that the second character is the target character, wherein the second character is any one of the plurality of characters; in a case where the character numbers corresponding to the plurality of characters are all less than the preset number, determining that the target character is not included in the plurality of characters.
[0078] The preset number in the above step can be a number determined according to the characteristics of large characters. For example, for an ancient book, since the width of a large character is usually twice the width of other ordinary characters, that is, there are a total of four columns of characters in the vertical direction of the large character, therefore, the value of the preset number can be 4, but is not limited thereto, and can be determined according to the actual recognized text image.
[0079] In an optional embodiment, in a case where any one character meets the following rule, it can be determined that the character is a large character: the character has two upper neighbors, and the character has two lower neighbors, that is, the number of characters adjacent to the character in the vertical direction is four. Based on the rule, for each character recognized in the text image, the number of characters adjacent to each character in the vertical direction can be determined. If there are two adjacent characters above any one character, and there are two adjacent characters below the character, that is, the character number corresponding to the character is four, it is determined that the character is a large character, that is, it is determined that the target character is included in the plurality of characters. If there are not two adjacent characters above the plurality of characters, or there are not two adjacent characters below the plurality of characters, that is, the character numbers corresponding to the plurality of characters are all less than four, it is determined that the target character is not included in the plurality of characters.
[0080] For example, still taking the ancient book as an example, for the character a, if there are adjacent characters b and c above the character a, and there are adjacent characters d and e below the character a, it can be determined that the character a is a large character.
[0081] In the above embodiment of the present application, in the case where the number of characters corresponding to the second character is greater than or equal to the preset number, the method further comprises: obtaining a character adjacent to the second character in the vertical direction to obtain an adjacent character corresponding to the second character; determining a first coincidence degree of the second character and the adjacent character based on the position information of the second character and the adjacent character; in the case where the first coincidence degree is greater than a first preset coincidence degree, determining that the second character is a target character; in the case where the first coincidence degree is less than the first preset coincidence degree, determining that the second character is not a target character.
[0082] The first preset coincidence degree in the above step can be a threshold value determined according to experiments, but the threshold value can be modified as needed, for example, in the embodiment of the present application, 0.6 is taken as an example for illustration.
[0083] In an optional embodiment, in the case where any one character meets the following rules, it can be determined that the character is a large character: the character has two upper neighbors, and the coincidence degree with all neighbors is greater than a first preset coincidence degree; the character has two lower neighbors, and the coincidence degree with all neighbors is greater than a first preset coincidence degree. Based on the rules, for each character recognized in the text image, the adjacent characters above and below each character can be determined, if there are two adjacent characters above any one character, and there are two adjacent characters below the character, and the first coincidence degree of the character with all adjacent characters is greater than the first preset coincidence degree, it is determined that the character is a large character, if there are not two adjacent characters above the character, or there are not two adjacent characters below the character, or the first coincidence degree of the character with any one adjacent character is less than the first preset coincidence degree, it is determined that the character is not a large character.
[0084] For example, still taking the ancient book as an example, for the character a, if there are adjacent characters b and c above the character a, and there are adjacent characters d and e below the character a, and the first coincidence degree of the character a with the character b is greater than 0.6, the first coincidence degree of the character a with the character c is greater than 0.6, the first coincidence degree of the character a with the character d is greater than 0.6, and the first coincidence degree of the character a with the character e is greater than 0.6, it can be determined that the character a is a large character.
[0085] In the above embodiments of the present application, the number of characters adjacent to each character in the vertical direction is obtained based on the position information of the plurality of characters, and the number of characters corresponding to each character is obtained by: obtaining candidate characters in the vertical direction of each character based on the position information of the plurality of characters; determining a first target region of each character and the candidate characters in the vertical direction based on the position information of each character and the candidate characters; determining whether there are other characters in the first target region based on the position information of the plurality of characters; and if there are no other characters in the first target region, obtaining the number of candidate characters to obtain the number of characters corresponding to each character.
[0086] It should be noted that the condition for determining the vertical direction neighbors of each character is to determine whether there are other characters in the adjacent region of the two characters, and if not, it is determined that the two characters are neighbors, that is, the two characters are adjacent; if there are, it is determined that the two characters are not neighbors, that is, the two characters are not adjacent. Therefore, the first target region described above can refer to the adjacent region between each character and the candidate character, such as the region shown by the dashed box in Figure 5
[0087] In an optional embodiment, the above condition can be used for judgment. First, based on the position information of each character, the characters above and below each character are determined to obtain candidate characters, which may or may not be adjacent to each character and need to be further confirmed. Further, based on the position information of the character and the candidate character, the adjacent region between the character and the candidate character can be determined. Assuming that the position information of the character 1 above is (x1, y1, w1, h1), and the position information of the character 2 below is (x2, y2, w2, h2), the position information of the adjacent region is (max(x1, x2), (y2+h2), a, b), where a = min(x1+w1, x2+w2)-max(x1, x2), and b = y1-(y2+h2). After determining the position information of the adjacent region, it can be determined based on the position information of other characters whether there are other characters in the adjacent region, and if not, the candidate character is determined to be the character adjacent to the character in the vertical direction.
[0088] In the above embodiments of the present application, the position information at least includes a first coordinate in the horizontal direction and a width, and the first coincidence degree of the second character and the adjacent character is determined based on the position information of the second character and the adjacent character, which includes: obtaining the minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determining a second target region of the second character and the adjacent character in the vertical direction; and obtaining the ratio of the width of the second target region to the first width to obtain the first coincidence degree.
[0089] In an optional embodiment, the first overlap degree can be obtained by using the following formula: a / min(w1, w2), as shown in the following formula: Figure 5 wherein a represents the width of the adjacent area. Thus, for the second character and the adjacent character, the minimum width of the widths of the two characters can be obtained to obtain the first width, and the adjacent area of the two characters is determined, and then the ratio of the width of the adjacent area to the first width can be obtained to obtain the vertical direction overlap degree of the two characters (i.e., the first overlap degree described above).
[0090] In the above embodiments of the present application, based on the positional relationship between any two first characters, the positional relationship between the first character and the target character, the positional relationship between any two target characters, and the first layout mode, the sorting relationship between any two characters includes: based on the positional relationship between any two first characters, the positional relationship between the first character and the target character, and the first layout mode, the sorting relationship between any two first characters is determined; based on the positional relationship between any two target characters and the first layout mode, the sorting relationship between any two target characters is determined; and the sorting relationship between any two first characters and the sorting relationship between any two target characters are summarized to obtain the sorting relationship between any two characters.
[0091] In an optional embodiment, for a large character, the sorting relationship between any two large characters can be determined based on the positional relationship between the two large characters; for the characters above the large character, the sorting relationship between the two characters can be determined based on the positional relationship between the two characters; and for the characters below the large character, the sorting relationship between the two characters can also be determined based on the positional relationship between the two characters. Furthermore, all the sorting relationships described above can be summarized to obtain the sorting relationship between any two characters related to the large character.
[0092] For example, still taking the ancient book as an example, for characters a, b, c, d and e, it is assumed that the character a is a large character, the positional relationship between the character b and the character a is "above", the positional relationship between the character c and the character a is "above", the positional relationship between the character b and the character c is "right", at this time, the character b is arranged in front of the character c, and the character c is arranged in front of the character a; the positional relationship between the character d and the character a is "below", the positional relationship between the character e and the character a is "below", and the positional relationship between the character d and the character e is "left", at this time, the character a is arranged in front of the character e, and the character e is arranged in front of the character d.
[0093] In the above embodiments of the present application, based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters, the text data corresponding to the text image is generated by: splicing the content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate the text data, wherein, in the case that the plurality of characters include a target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0094] The preset identification information in the above steps may refer to a line break character, but is not limited thereto, and may also be implemented by other means that can distinguish large characters from other ordinary texts.
[0095] In an optional embodiment, the texts may be concatenated in the order of all texts. When a large character is concatenated, line break characters are added before and after the large character, so as to obtain the final serialized text. It should be noted that if there are multiple consecutive large characters, a line break character is added before the first large character and a line break character is added after the last large character. For example, taking ancient books as an example, as Figure 6 shown, "○ Zheng Huangqi Gaha" and "○ Zheng Hongqi Kerqi" are large characters, so line break characters are added.
[0096] In the above embodiments of the present application, after identifying a text image and obtaining the recognition results of multiple texts in the text image, the method further includes: based on the position information of the multiple texts, determining whether any two texts meet a second preset condition; if any two texts meet the second preset condition, merging any two texts to obtain multiple merged text blocks; based on the position information of the multiple text blocks, determining the sorting relationship between the multiple text blocks; and generating text data based on the sorting relationship between the multiple text blocks.
[0097] It should be noted that since the number of texts included in the text image is large, if the sorting relationship between texts is directly determined based on the position information of all texts, the calculation amount is large, which may lead to an increase in processing time and affect the processing efficiency. To solve this problem, the texts can be merged, and the texts in the vertical direction are merged, so as to achieve the effect of improving the processing efficiency.
[0098] The second preset condition in the above steps may refer to the condition for merging two single characters, which specifically includes the following conditions: the two single characters are vertical neighbors; since the size of large characters in the text image is different from that of ordinary texts, and the reading order for large characters is different from that of ordinary texts, in order to avoid the influence of large characters on the merged text blocks, ordinary texts and large characters are not merged. Therefore, it is necessary to determine that the relative size of the two single characters is less than a certain threshold; the vertical overlap degree of the two single characters is greater than a certain threshold; and the relative distance in the vertical direction of the two single characters is less than a certain threshold.
[0099] In an optional embodiment, based on the position relationship between any two texts, it can be determined whether the two texts meet the second preset condition. If so, the two texts are merged, so that multiple text blocks can be obtained for all texts in the text image, as Figure 6As shown, adjacent large characters can be merged, characters in the same column above and below the large character can be merged, and characters in the same column without large characters can be merged, so that a smaller number of character blocks can be obtained. At this time, the positional relationship between most character blocks is a "left, right" relationship, so that the sorting relationship between multiple character blocks can be determined based on the positional information of the character blocks, and further, the sorting relationship of all characters in the same character block can be combined to generate a serialized text.
[0100] In the above embodiment of the present application, determining whether any two characters satisfy the second preset condition based on the positional information of the multiple characters includes: determining the size relative ratio, the second coincidence degree and the relative distance of any two characters based on the positional information of the any two characters; determining whether the any two characters are adjacent in the vertical direction, whether the size relative ratio is less than the preset relative ratio, whether the second coincidence degree is greater than the second preset coincidence degree, and whether the relative distance is less than the preset distance; if the any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance, it is determined that the any two characters satisfy the second preset condition; if the any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance, it is determined that the any two characters do not satisfy the second preset condition.
[0101] The preset relative ratio, the second preset coincidence degree and the preset distance in the above step can be threshold values determined according to experiments, but the threshold values can be modified as needed. For example, in the embodiment of the present application, the preset relative ratio is 0.8, the second preset coincidence degree is 0.75, and the preset distance is 0.5.
[0102] In an alternative embodiment, in combination with the conditions for single character merging, it can be first determined whether two characters are vertical neighbors, that is, whether there are other characters in the adjacent regions of the two characters. If the two characters are vertical neighbors, it is determined whether the character size relative ratio (i.e. the size relative ratio described above) of the two characters is less than the preset relative ratio. If so, it is determined whether the vertical direction coincidence degree (i.e. the second coincidence degree described above) of the two single characters is greater than the second preset coincidence degree. If so, it is determined whether the vertical direction relative distance (i.e. the relative distance described above) of the two single characters is less than the preset distance. If so, it is determined that the two characters can be merged. If any one of the conditions is not met, it is determined that the two characters cannot be merged.
[0103] It should be noted that the second coincidence degree described above can be implemented in the same way as the first coincidence degree described above, and the present application will not be repeated here.
[0104] In the above embodiments of the present application, the position information at least includes a width, and the determining of the size relative ratio of the two arbitrary characters based on the position information of the two arbitrary characters includes: obtaining a ratio of the widths of the two arbitrary characters to obtain two ratios; and obtaining a minimum ratio of the two ratios to obtain the size relative ratio.
[0105] In an optional embodiment, the size relative ratio can be calculated by the following formula: min(w1 / w2, w2 / w1). Thus, for the character a and the character b, a ratio of the width of the character a to the width of the character b and a ratio of the width of the character b to the width of the character a can be obtained to obtain two ratios, and then the minimum ratio is taken as the size relative ratio.
[0106] In the above embodiments of the present application, the position information at least includes a second coordinate in a vertical direction and a width, and the determining of the relative distance of the two arbitrary characters based on the position information of the two arbitrary characters includes: obtaining a minimum width of the widths of the two arbitrary characters to obtain a second width; determining a third target region of the two arbitrary characters in the vertical direction; and obtaining a ratio of a height of the third target region to the second width to obtain the relative distance.
[0107] In an optional embodiment, the relative distance can be obtained by the following formula: b / min(w1, w2), as shown in Figure 5 Thus, for the two arbitrary characters, a minimum width of the widths of the two characters can be obtained to obtain a second width, and a joint region of the two characters is determined, and then a ratio of a height of the joint region to the second width can be obtained to obtain the relative distance of the two characters in the vertical direction (i.e., the relative distance described above).
[0108] In the above embodiments of the present application, after the text data corresponding to the text image is generated, the method further includes: outputting the text data; and obtaining response data corresponding to the text data, wherein the response data is obtained by modifying the text data.
[0109] In order to facilitate the user to read the electronic text, the cloud server can download the text data to the computer terminal of the user through the network after obtaining the text data, and display the text data in the "image display" region of the interactive interface as shown in Figure 3
[0110] Further, after the user views the text data, the user can directly edit or modify the text data in the interactive interface, modify the content of the text recognized by the OCR technology, or modify the order of the text in the wrong order, to obtain modified text data (i.e., the above-mentioned response data). After the user finishes modifying or editing, the modified text data can be uploaded to the cloud server through the computer terminal, and the cloud server can further improve based on the modified text data, for example, upgrade the OCR technology, or adjust the first layout mode corresponding to the text image, which is not limited in the present application.
[0111] In the above embodiments of the present application, the output text data includes: obtaining a second layout mode; laying out the text data according to the second layout mode to obtain laid-out text data; and outputting the laid-out text data.
[0112] The second layout mode in the above steps can be a common layout mode, for example, a layout mode of "from top to bottom, from left to right", but is not limited thereto, and can also be other different types of layout modes.
[0113] In an optional embodiment, since the layout mode of the text image can be different from the common layout mode, in order to facilitate the user to view, the serialized text can be laid out according to the common layout mode to obtain laid-out text data, and the laid-out text data is displayed to the user for viewing.
[0114] In another optional embodiment, since different users have different reading habits, in order to facilitate different users to view, the second layout mode can be determined according to the user's reading habits, and the serialized text can be laid out according to the second layout mode to obtain laid-out text data, which is further displayed to the user for viewing.
[0115] It should be noted that if the text data contains a target text, the target text can be displayed in a special way. The specific display method can be pre-set or set by the user.
[0116] In the above embodiments of the present application, obtaining the second layout mode includes one of the following: obtaining a layout mode corresponding to a target object to obtain the second layout mode; and obtaining a selected layout mode from a plurality of output layout modes to obtain the second layout mode.
[0117] The target object in the above steps can refer to a user who uploads the text image, or a user who checks and corrects the text data, which is not limited in the present application.
[0118] Since reading habits of different users are different, the second layout mode can be determined in combination with the reading habits of different users, the second layout mode can be pre-set by the user according to the reading habits of the user, or the user is provided with a plurality of different layout modes, the user selects the layout mode, and the layout mode selected by the user is taken as the second layout mode.
[0119] The preferred embodiment of the application will be described in detail below Figure 7 The method can be executed by a mobile terminal or a server, and in the embodiment of the application, the method is executed by the server as an example. As shown in the figure, Figure 7 The method can include the following steps:
[0120] In step S71, the text image is recognized by using OCR to obtain a single character result, wherein the single character result is the character content and the position, and the position is identified by (start point x coordinate, start point y coordinate, width, height).
[0121] In step S72, the single characters are merged based on the single character result to obtain a character string.
[0122] Optionally, two single characters can be determined based on a single character merging rule, and the two single characters meeting the single character merging rule are merged.
[0123] In step S73, the single character result is used for large character determination.
[0124] Optionally, the text blocks can be determined based on a large character determination rule, and the text blocks meeting the large character determination rule are determined as large characters.
[0125] In step S74, the reading order of the character string and the large character is analyzed to determine the order of the text blocks.
[0126] Optionally, the position relationship of all text blocks can be analyzed two by two to determine the sorting relationship of two text blocks, the position relationship related to the large character is checked to determine the sorting relationship of the text blocks, and the sorting of the text blocks is determined by comprehensively considering the two-by-two sorting relationship.
[0127] In step S75, the text data is composed according to the order of the text blocks.
[0128] Optionally, the text is spliced according to the order of the text blocks, and a line feed character is added when a large character is encountered.
[0129] Through the above steps, the data characteristics of ancient books can be adapted to the large character detection unique to ancient books and the unique character arrangement order of ancient books, so that better serialization effect is achieved.
[0130] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0131] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a necessary general hardware platform, and of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or the part that contributes to the prior art, and the computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for causing an end device (which can be a mobile phone, a computer, a server, or a network device) to execute the method described in each embodiment of the present application.
[0132] Embodiment 2
[0133] According to the embodiments of the present application, a text processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0134] Figure 8 is a flowchart of a second text processing method according to an embodiment of the present application. As shown in Figure 8 , the method comprises the following steps:
[0135] Step S802, displaying a text image.
[0136] In an optional embodiment, the user-selected text image can be displayed on an interactive interface as shown in Figure 3 .
[0137] Step S804, marking the recognition results of a plurality of characters in the text image in the text image, wherein the recognition results of the plurality of characters are obtained by recognizing the text image, and the recognition results include content information corresponding to each character and position information of each character in the text image.
[0138] In an alternative embodiment, each character in the text image can be framed and the corresponding content information can be displayed based on the position information of each character in the text image, so as to achieve the purpose of marking the recognition result in the text image.
[0139] In step S806, the text data corresponding to the text image is displayed, wherein the text data is generated based on the ordering relationship between the plurality of characters and the content information corresponding to the plurality of characters, and the ordering relationship between the plurality of characters is determined based on the position information of the plurality of characters.
[0140] In an alternative embodiment, the generated serialized text, i.e., the text data described above, can be displayed on an interactive interface as shown in FIG. 8B. Figure 3
[0141] In the above embodiments of the present application, before marking the recognition result of the plurality of characters in the text image, the method further comprises: determining the position relationship between any two characters based on the position information of the any two characters; determining the ordering relationship between the any two characters based on the position relationship between the any two characters and the first typesetting mode corresponding to the text image; and aggregating the ordering relationship between the any two characters to obtain the ordering relationship between the plurality of characters.
[0142] In the above embodiments of the present application, before determining the ordering relationship between the any two characters based on the position relationship between the any two characters and the first typesetting mode corresponding to the text image, the method further comprises: determining whether the plurality of characters contain a target character based on the position information of the plurality of characters, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; in the case where the plurality of characters contain the target character, determining the ordering relationship between the any two characters based on the position relationship between the any two first characters, the position relationship between the first character and the target character, the position relationship between the any two target characters, and the first typesetting mode, wherein the first character is a character in the plurality of characters located in the vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; in the case where the plurality of characters do not contain the target character, determining the ordering relationship between the any two characters based on the position relationship between the any two characters and the first typesetting mode.
[0143] In the above embodiments of the present application, determining whether the plurality of characters contain a target character based on the position information of the plurality of characters comprises: obtaining the number of characters adjacent to each character in the vertical direction based on the position information of the plurality of characters to obtain the character number corresponding to each character; in the case where the character number corresponding to the second character is greater than or equal to a preset number, determining that the second character is the target character, wherein the second character is any one of the plurality of characters; in the case where the character number corresponding to the plurality of characters is less than the preset number, determining that the plurality of characters do not contain the target character.
[0144] In the above embodiments of the present application, in the case where the number of characters corresponding to the second character is greater than or equal to the preset number, the method further comprises: obtaining a character adjacent to the second character in the vertical direction to obtain an adjacent character corresponding to the second character; determining a first coincidence degree of the second character and the adjacent character based on the position information of the second character and the adjacent character; in the case where the first coincidence degree is greater than a first preset coincidence degree, determining that the second character is the target character; in the case where the first coincidence degree is less than the first preset coincidence degree, determining that the second character is not the target character.
[0145] In the above embodiments of the present application, the obtaining of the number of characters adjacent to each character in the vertical direction comprises: obtaining candidate characters located in the vertical direction of each character based on the position information of the plurality of characters; determining a first target region of each character and the candidate characters in the vertical direction based on the position information of each character and the candidate characters; judging whether there are other characters in the first target region based on the position information of the plurality of characters; and if there are no other characters in the first target region, obtaining the number of candidate characters to obtain the number of characters corresponding to each character.
[0146] In the above embodiments of the present application, the position information at least comprises a first coordinate in the horizontal direction and a width, and the determining of the first coincidence degree of the second character and the adjacent character based on the position information of the second character and the adjacent character comprises: obtaining a minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determining a second target region of the second character and the adjacent character in the vertical direction; and obtaining a ratio of the width of the second target region to the first width to obtain the first coincidence degree.
[0147] In the above embodiments of the present application, the determining of the sorting relationship of any two characters based on the positional relationship of any two first characters, the positional relationship of the first character and the target character, the positional relationship of any two target characters, and the first typesetting mode comprises: determining the sorting relationship of any two first characters based on the positional relationship of any two first characters, the positional relationship of the first character and the target character, and the first typesetting mode; determining the sorting relationship of any two target characters based on the positional relationship of any two target characters and the first typesetting mode; and summarizing the sorting relationship of any two first characters and the sorting relationship of any two target characters to obtain the sorting relationship of any two characters.
[0148] In the above embodiments of the present application, before displaying the text data corresponding to the text image, the method further comprises: splicing the content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate the text data, wherein in the case where the plurality of characters include the target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0149] In the above embodiments of the present application, after marking the recognition results of the plurality of characters in the text image in the text image, the method further includes: marking the plurality of character blocks in the text image, wherein the plurality of character blocks are obtained by merging any two characters that satisfy the second preset condition based on the position information of the plurality of characters; and displaying text data corresponding to the text image, wherein the text data is generated based on the ordering relationship between the plurality of character blocks, and the ordering relationship between the plurality of character blocks is generated based on the position information of the plurality of character blocks.
[0150] In the above embodiments of the present application, determining whether any two characters satisfy the second preset condition based on the position information of the plurality of characters includes: determining the size relative ratio, the second coincidence degree and the relative distance of any two characters based on the position information of the two characters; determining whether any two characters are adjacent in the vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; if any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance, it is determined that any two characters satisfy the second preset condition; if any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance, it is determined that any two characters do not satisfy the second preset condition.
[0151] In the above embodiments of the present application, the position information at least includes a width, and determining the size relative ratio of any two characters based on the position information of the two characters includes: obtaining the ratio of the widths of any two characters to obtain two ratios; and obtaining the minimum ratio of the two ratios to obtain the size relative ratio.
[0152] In the above embodiments of the present application, the position information at least includes a second coordinate in the vertical direction and a width, and determining the relative distance of any two characters based on the position information of the two characters includes: obtaining the minimum width of the widths of any two characters to obtain a second width; determining a third target area of the two characters in the vertical direction; and obtaining the ratio of the height of the third target area to the second width to obtain the relative distance.
[0153] In the above embodiments of the present application, after displaying the text data corresponding to the text image, response data corresponding to the text data is obtained, wherein the response data is obtained by modifying the text data.
[0154] In the above embodiments of the present application, displaying the text data corresponding to the text image includes: obtaining a second typesetting manner; typesetting the text data according to the second typesetting manner to obtain typeset text data; and displaying the typeset text data.
[0155] In the above embodiments of the present application, the second layout mode is obtained in one of the following ways: obtaining the layout mode corresponding to the target object to obtain the second layout mode; and selecting a layout mode from the output multiple layout modes to obtain the second layout mode.
[0156] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same application scenarios and implementation processes as the schemes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0157] Embodiment 3
[0158] According to the embodiments of the present application, a text processing method is also provided. It should be noted that the steps shown in the flowchart can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0159] Figure 9 is a flowchart of a third text processing method according to an embodiment of the present application. The method can be executed by a cloud server, as shown in Figure 9 The method comprises the following steps:
[0160] Step S902, receiving a text image.
[0161] In an optional embodiment, the user can take a text image by a shooting device and transmit it to the cloud server through the user's computer terminal, so that the cloud server can receive the text image to be processed.
[0162] Step S904, recognizing the text image to obtain the recognition results of multiple characters in the text image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the text image.
[0163] Step S906, determining the sorting relationship between the multiple characters based on the position information of the multiple characters.
[0164] Step S908, generating text data corresponding to the text image based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters.
[0165] Step S910, outputting the text data.
[0166] In an optional embodiment, the cloud server can transmit the generated serialized text, i.e., the text data described above, to the computer terminal, which displays it on, for example, Figure 3The interaction interface shown is for a user to view. In addition, the text data can be transmitted to a computer terminal for storage, or directly stored in a local cloud server.
[0167] In the above embodiment of the present application, determining the ordering relationship between the plurality of characters based on the position information of the plurality of characters includes: determining the position relationship between any two characters based on the position information of the two characters; determining the ordering relationship between the two characters based on the position relationship between the two characters and the first typesetting mode corresponding to the text image; and aggregating the ordering relationship between the two characters to obtain the ordering relationship between the plurality of characters.
[0168] In the above embodiment of the present application, before determining the ordering relationship between the two characters based on the position relationship between the two characters and the first typesetting mode corresponding to the text image, the method further includes: determining whether the plurality of characters contains a target character based on the position information of the plurality of characters, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; in the case where the plurality of characters contains the target character, determining the ordering relationship between the two characters based on the position relationship between the two characters, the position relationship between the first character and the target character, the position relationship between the two target characters, and the first typesetting mode, wherein the first character is a character in the plurality of characters located in the vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and in the case where the plurality of characters does not contain the target character, determining the ordering relationship between the two characters based on the position relationship between the two characters and the first typesetting mode.
[0169] In the above embodiment of the present application, determining whether the plurality of characters contains a target character based on the position information of the plurality of characters includes: obtaining the number of characters adjacent to each character in the vertical direction based on the position information of the plurality of characters, to obtain the number of characters corresponding to each character; in the case where the number of characters corresponding to the second character is greater than or equal to a preset number, determining that the second character is the target character, wherein the second character is any one of the plurality of characters; and in the case where the number of characters corresponding to the plurality of characters is all less than the preset number, determining that the plurality of characters does not contain the target character.
[0170] In the above embodiment of the present application, in the case where the number of characters corresponding to the second character is greater than or equal to the preset number, the method further includes: obtaining a character adjacent to the second character in the vertical direction to obtain the adjacent character corresponding to the second character; determining a first coincidence degree between the second character and the adjacent character based on the position information of the second character and the adjacent character; in the case where the first coincidence degree is greater than a first preset coincidence degree, determining that the second character is the target character; and in the case where the first coincidence degree is less than the first preset coincidence degree, determining that the second character is not the target character.
[0171] In the above embodiments of the present application, the number of characters corresponding to each character is obtained by: obtaining candidate characters in the vertical direction of each character based on the position information of the plurality of characters; determining a first target region of each character and the candidate characters in the vertical direction based on the position information of each character and the candidate characters; determining whether there are other characters in the first target region based on the position information of the plurality of characters; and if there are no other characters in the first target region, obtaining the number of candidate characters to obtain the number of characters corresponding to each character.
[0172] In the above embodiments of the present application, the position information at least includes a first coordinate in the horizontal direction and a width, and the first coincidence degree of the second character and the adjacent character is determined based on the position information of the second character and the adjacent character, which includes: obtaining a minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determining a second target region of the second character and the adjacent character in the vertical direction; and obtaining a ratio of the width of the second target region to the first width to obtain the first coincidence degree.
[0173] In the above embodiments of the present application, the sorting relationship of any two characters is determined based on the positional relationship of any two first characters, the positional relationship of the first character and the target character, the positional relationship of any two target characters, and the first layout mode, which includes: determining the sorting relationship of any two first characters based on the positional relationship of any two first characters, the positional relationship of the first character and the target character, and the first layout mode; determining the sorting relationship of any two target characters based on the positional relationship of any two target characters and the first layout mode; and summarizing the sorting relationship of any two first characters and the sorting relationship of any two target characters to obtain the sorting relationship of any two characters.
[0174] In the above embodiments of the present application, the text data corresponding to the text image is generated based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters, which includes: splicing the content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate the text data, wherein in the case that the plurality of characters include the target character, the preset identification information is added before and / or after the content information corresponding to the target character.
[0175] In the above embodiments of the present application, after identifying the text image to obtain the recognition result of the plurality of characters in the text image, the method further includes: determining whether any two characters satisfy a second preset condition based on the position information of the plurality of characters; if any two characters satisfy the second preset condition, merging any two characters to obtain a plurality of merged character blocks; determining the sorting relationship between the plurality of character blocks based on the position information of the plurality of character blocks; and generating text data based on the sorting relationship between the plurality of character blocks.
[0176] In the above embodiments of the present application, determining whether the two arbitrary characters satisfy the second preset condition based on the position information of the plurality of characters includes: determining the size relative ratio, the second coincidence degree and the relative distance of the two arbitrary characters based on the position information of the two arbitrary characters; determining whether the two arbitrary characters are adjacent in the vertical direction, whether the size relative ratio is less than the preset relative ratio, whether the second coincidence degree is greater than the second preset coincidence degree, and whether the relative distance is less than the preset distance; if the two arbitrary characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance, it is determined that the two arbitrary characters satisfy the second preset condition; if the two arbitrary characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance, it is determined that the two arbitrary characters do not satisfy the second preset condition.
[0177] In the above embodiments of the present application, the position information at least includes a width, and determining the size relative ratio of the two arbitrary characters based on the position information of the two arbitrary characters includes: obtaining two ratio values of the widths of the two arbitrary characters; obtaining the minimum ratio value of the two ratio values to obtain the size relative ratio.
[0178] In the above embodiments of the present application, the position information at least includes a second coordinate in the vertical direction and a width, and determining the relative distance of the two arbitrary characters based on the position information of the two arbitrary characters includes: obtaining the minimum width of the widths of the two arbitrary characters to obtain a second width; determining a third target area of the two arbitrary characters in the vertical direction; obtaining the ratio of the height of the third target area to the second width to obtain the relative distance.
[0179] In the above embodiments of the present application, after the text data is output, the response data corresponding to the text data is obtained, wherein the response data is obtained by modifying the text data.
[0180] In the above embodiments of the present application, outputting the text data includes: obtaining a second layout mode; performing layout on the text data according to the second layout mode to obtain the laid-out text data; and outputting the laid-out text data.
[0181] In the above embodiments of the present application, obtaining the second layout mode includes one of the following: obtaining a layout mode corresponding to the target object to obtain the second layout mode; and obtaining a selected layout mode from a plurality of output layout modes to obtain the second layout mode.
[0182] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same application scenarios and implementation processes as the schemes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0183] Embodiment 4
[0184] According to the embodiments of the present application, a text processing apparatus for implementing the above text processing method is further provided, as shown in the apparatus 1000 comprises: an acquisition module 1002, an identification module 1004, a determination module 1006 and a generation module 1008. Figure 10
[0185] The acquisition module 1002 is configured to acquire a text image; the identification module 1004 is configured to identify the text image to obtain identification results of a plurality of characters in the text image, wherein the identification results comprise content information corresponding to each character and position information of each character in the text image; the determination module 1006 is configured to determine a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and the generation module 1008 is configured to generate text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0186] It should be noted that the acquisition module 1002, the identification module 1004, the determination module 1006 and the generation module 1008 correspond to steps S202 to S208 in the embodiment 1, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above embodiment 1. It should be noted that the above modules as part of the apparatus can run in the computer terminal 10 provided in the embodiment 1.
[0187] In the above embodiments of the present application, the determination module comprises a first determination unit, a second determination unit and a summary unit.
[0188] The first determination unit is configured to determine a position relationship between any two characters based on the position information of the any two characters; the second determination unit is configured to determine a sorting relationship between the any two characters based on the position relationship between the any two characters and a first layout mode corresponding to the text image; and the summary unit is configured to summarize the sorting relationship between the any two characters to obtain the sorting relationship between the plurality of characters.
[0189] In the above embodiments of the present application, the determination module further comprises a third determination unit and a fourth determination unit.
[0190] The third determining unit is configured to determine, based on the position information of the plurality of characters, whether the plurality of characters contain a target character, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; the fourth determining unit is configured to, in a case where the plurality of characters contain the target character, determine, based on a positional relationship between any two first characters, a positional relationship between the first character and the target character, a positional relationship between any two target characters, and a first layout mode, a sorting relationship between any two characters, wherein the first character is a character in the plurality of characters that is located in a vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and the second determining unit is further configured to, in a case where the plurality of characters do not contain the target character, determine, based on the positional relationship between any two characters and the first layout mode, the sorting relationship between any two characters.
[0191] In the above embodiments of the present application, the third determining unit comprises a first obtaining subunit, a first determining subunit, and a second determining subunit.
[0192] The first obtaining subunit is configured to obtain, based on the position information of the plurality of characters, a number of characters adjacent to each character in a vertical direction, to obtain a character number corresponding to each character; the first determining subunit is configured to, in a case where the character number corresponding to the second character is greater than or equal to a preset number, determine that the second character is the target character, wherein the second character is any one character in the plurality of characters; and the second determining subunit is configured to, in a case where the character numbers corresponding to the plurality of characters are all less than the preset number, determine that the plurality of characters do not contain the target character.
[0193] In the above embodiments of the present application, the third determining unit further comprises a second obtaining subunit, a third determining subunit, and a fourth determining subunit.
[0194] The second obtaining subunit is configured to obtain, based on the position information of the second character and the adjacent character, a first coincidence degree between the second character and the adjacent character; the second determining subunit is further configured to, in a case where the first coincidence degree is greater than a first preset coincidence degree, determine that the second character is the target character; and the fourth determining subunit is configured to, in a case where the first coincidence degree is less than the first preset coincidence degree, determine that the second character is not the target character.
[0195] In the above embodiments of the present application, the first obtaining subunit is further configured to obtain a candidate character located in a vertical direction of each character based on the position information of the plurality of characters; determine a first target region of each character and the candidate character in the vertical direction based on the position information of each character and the candidate character; determine whether there is another character in the first target region based on the position information of the plurality of characters; and if there is no other character in the first target region, obtain the number of candidate characters to obtain the number of characters corresponding to each character.
[0196] In the above embodiments of the present application, the position information at least includes a first coordinate and a width in a horizontal direction, wherein the third determining subunit is further configured to obtain a minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determine a second target region of the second character and the adjacent character in a vertical direction; and obtain a ratio of the width of the second target region to the first width to obtain a first coincidence degree.
[0197] In the above embodiments of the present application, the fourth determining unit includes a fifth determining subunit, a sixth determining subunit, and a summarizing subunit.
[0198] The fifth determining subunit is configured to determine a sorting relationship of any two first characters based on the positional relationship of the any two first characters, the positional relationship of the first character and the target character, and the first layout mode; the sixth determining subunit is configured to determine a sorting relationship of any two target characters based on the positional relationship of the any two target characters and the first layout mode; and the summarizing subunit is configured to summarize the sorting relationship of the any two first characters and the sorting relationship of the any two target characters to obtain a sorting relationship of any two characters.
[0199] In the above embodiments of the present application, the generation module includes a generation unit.
[0200] The generation unit is configured to splice content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate text data, wherein in the case where the plurality of characters include the target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0201] In the above embodiments of the present application, the device further includes a judgment module and a merging module.
[0202] The judgment module is configured to determine whether any two characters satisfy a second preset condition based on the position information of the plurality of characters; the merging module is configured to merge any two characters to obtain a plurality of merged character blocks if the any two characters satisfy the second preset condition; the determination module is further configured to determine a sorting relationship between the plurality of character blocks based on the position information of the plurality of character blocks; and the generation module is configured to generate text data based on the sorting relationship between the plurality of character blocks.
[0203] In the above embodiments of the present application, the determining module comprises: a fifth determining unit, a judging unit, a sixth determining unit and a seventh determining unit.
[0204] The fifth determining unit is configured to determine the size relative ratio, the second coincidence degree and the relative distance of any two characters based on the position information of the two characters. The judging unit is configured to judge whether the two characters are adjacent in the vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance. The sixth determining unit is configured to determine that the two characters satisfy the second preset condition if the two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance. The seventh determining unit is configured to determine that the two characters do not satisfy the second preset condition if the two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance.
[0205] In the above embodiments of the present application, the position information at least comprises a width, and the fifth determining unit comprises: a third obtaining subunit and a fourth obtaining subunit.
[0206] The third obtaining subunit is configured to obtain the ratio of the widths of any two characters to obtain two ratios. The fourth obtaining subunit is configured to obtain the minimum ratio of the two ratios to obtain the size relative ratio.
[0207] In the above embodiments of the present application, the position information at least comprises a second coordinate in the vertical direction and a width, and the fifth determining unit comprises: a fifth obtaining subunit, a seventh determining subunit and a sixth obtaining subunit.
[0208] The fifth obtaining subunit is configured to obtain the minimum width of the widths of any two characters to obtain a second width. The seventh determining subunit is configured to determine a third target region of the two characters in the vertical direction. The sixth obtaining subunit is configured to obtain the ratio of the height of the third target region to the second width to obtain the relative distance.
[0209] In the above embodiments of the present application, the device further comprises: an output module.
[0210] The output module is configured to output the text data. The obtaining module is further configured to obtain response data corresponding to the text data, wherein the response data is obtained by modifying the text data.
[0211] In the above embodiments of the present application, the output module comprises: an obtaining unit, a layout unit and an output unit.
[0212] The obtaining unit is configured to obtain a second layout mode; the layout unit is configured to layout the text data according to the second layout mode to obtain laid-out text data; and the output unit is configured to output the laid-out text data.
[0213] In the above embodiments, the obtaining unit includes one of a seventh obtaining subunit and an eighth obtaining subunit.
[0214] The seventh obtaining subunit is configured to obtain a layout mode corresponding to a target object to obtain the second layout mode; and the eighth obtaining subunit is configured to obtain a selected layout mode from a plurality of output layout modes to obtain the second layout mode.
[0215] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same application scenarios and implementation processes as the scheme provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0216] Embodiment 5
[0217] According to the embodiments of the present application, a text processing device for implementing the above text processing method is further provided, as shown in the figure, the device 1100 includes a first display module 1102, a marking module 1104, and a second display module 1106. Figure 11
[0218] The first display module 1102 is configured to display a text image; the marking module 1104 is configured to mark recognition results of a plurality of characters in the text image in the text image, wherein the recognition results of the plurality of characters are obtained by recognizing the text image, and the recognition results include content information corresponding to each character and position information of each character in the text image; and the second display module 1106 is configured to display text data corresponding to the text image, wherein the text data is generated based on an ordering relationship between the plurality of characters and content information corresponding to the plurality of characters, and the ordering relationship between the plurality of characters is determined based on the position information of the plurality of characters.
[0219] It should be noted that the first display module 1102, the marking module 1104, and the second display module 1106 correspond to steps S802 to S806 in Embodiment 2, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules as part of the device can run in the computer terminal 10 provided in Embodiment 1.
[0220] In the above embodiments, the device further includes a first determination module, a second determination module, and a summary module.
[0221] The first determining module is configured to determine a position relationship between any two characters based on position information of the any two characters; the second determining module is configured to determine a sorting relationship between the any two characters based on the position relationship between the any two characters and the first layout mode corresponding to the text image; and the collecting module is configured to collect the sorting relationship between the any two characters to obtain the sorting relationship between the plurality of characters.
[0222] In the above embodiments, the device further includes a third determining module and a fourth determining module.
[0223] The third determining module is configured to determine whether the plurality of characters contain a target character based on position information of the plurality of characters, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; the fourth determining module is configured to, in a case where the plurality of characters contain the target character, determine a sorting relationship between any two characters based on a position relationship between the any two first characters, a position relationship between the first character and the target character, a position relationship between any two target characters, and the first layout mode, wherein the first character is a character in the plurality of characters that is located in a vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and the second determining module is further configured to, in a case where the plurality of characters do not contain the target character, determine the sorting relationship between the any two characters based on the position relationship between the any two characters and the first layout mode.
[0224] In the above embodiments, the third determining module includes a first obtaining unit, a first determining unit, and a second determining unit.
[0225] The first obtaining unit is configured to obtain a number of characters adjacent to each character in a vertical direction based on position information of the plurality of characters to obtain a character number corresponding to each character; the first determining unit is configured to determine that a second character is a target character in a case where the character number corresponding to the second character is greater than or equal to a preset number, wherein the second character is any one character in the plurality of characters; and the second determining unit is configured to determine that the plurality of characters do not contain the target character in a case where the character numbers corresponding to the plurality of characters are all less than the preset number.
[0226] In the above embodiments, the third determining module further includes a second obtaining unit, a third determining unit, and a fourth determining unit.
[0227] The second obtaining unit is configured to obtain a character adjacent to the second character in a vertical direction, to obtain an adjacent character corresponding to the second character; the third determining unit is configured to determine a first coincidence degree of the second character and the adjacent character based on position information of the second character and the adjacent character; the second determining unit is further configured to determine that the second character is the target character in a case where the first coincidence degree is greater than a first preset coincidence degree; and the fourth determining unit is configured to determine that the second character is not the target character in a case where the first coincidence degree is less than the first preset coincidence degree.
[0228] In the above embodiment, the first obtaining unit is further configured to obtain, based on the position information of the plurality of characters, a candidate character located in a vertical direction of each character; determine, based on the position information of each character and the candidate character, a first target region of each character and the candidate character in the vertical direction; determine, based on the position information of the plurality of characters, whether there is another character in the first target region; and if there is no other character in the first target region, obtain a number of the candidate characters to obtain a character number corresponding to each character.
[0229] In the above embodiment, the position information at least includes a first coordinate and a width in a horizontal direction, wherein the third determining unit is further configured to obtain a minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determine a second target region of the second character and the adjacent character in the vertical direction; and obtain a ratio of the width of the second target region to the first width to obtain the first coincidence degree.
[0230] In the above embodiment, the fourth determining module includes a fifth determining unit, a sixth determining unit, and a summarizing unit.
[0231] The fifth determining unit is configured to determine a sorting relationship of any two first characters based on a positional relationship of the any two first characters, a positional relationship of the first character and the target character, and the first layout mode; the sixth determining unit is configured to determine a sorting relationship of any two target characters based on a positional relationship of the any two target characters and the first layout mode; and the summarizing unit is configured to summarize the sorting relationship of the any two first characters and the sorting relationship of the any two target characters to obtain a sorting relationship of any two characters.
[0232] In the above embodiment, the device further includes a generating module.
[0233] The generating module is configured to splice content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate text data, wherein in a case where the plurality of characters include the target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0234] In the above embodiment, the device further includes a judging module, a merging module, and a fifth determining module.
[0235] The judgment module is configured to judge whether any two characters satisfy a second preset condition based on the position information of the plurality of characters; the merging module is configured to merge any two characters to obtain a plurality of merged character blocks if the any two characters satisfy the second preset condition; and the fifth determination module is configured to determine a sorting relationship between the plurality of character blocks based on the position information of the plurality of character blocks; and the generation module is further configured to generate the text data based on the sorting relationship between the plurality of character blocks.
[0236] In the above embodiments of the present application, the judgment module comprises a seventh determination unit, a judgment unit, an eighth determination unit and a ninth determination unit.
[0237] The seventh determination unit is configured to determine a size relative ratio, a second coincidence degree and a relative distance of any two characters based on the position information of the any two characters; the judgment unit is configured to judge whether the any two characters are adjacent in a vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; the eighth determination unit is configured to determine that the any two characters satisfy the second preset condition if the any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance; and the ninth determination unit is configured to determine that the any two characters do not satisfy the second preset condition if the any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance.
[0238] In the above embodiments of the present application, the position information at least comprises a width, and the seventh determination unit comprises a first obtaining subunit and a second obtaining subunit.
[0239] The first obtaining subunit is configured to obtain a ratio of the width of any two characters to obtain two ratios; and the second obtaining subunit is configured to obtain a minimum ratio of the two ratios to obtain the size relative ratio.
[0240] In the above embodiments of the present application, the position information at least comprises a second coordinate in a vertical direction and a width, and the seventh determination unit comprises a third obtaining subunit, a determination subunit and a fourth obtaining subunit.
[0241] The third obtaining subunit is configured to obtain a minimum width of the width of any two characters to obtain a second width; the determination subunit is configured to determine a third target region of the any two characters in the vertical direction; and the fourth obtaining subunit is configured to obtain a ratio of a height of the third target region to the second width to obtain the relative distance.
[0242] In the above embodiments of the present application, the device further comprises an obtaining module.
[0243] The acquisition module is used to acquire the response data corresponding to the text data, which is obtained by modifying the text data.
[0244] In the above embodiments of this application, the second display module includes: a third acquisition unit, a layout unit, and a display unit.
[0245] The third acquisition unit is used to acquire the second layout method; the layout unit is used to layout the text data according to the second layout method to obtain the layout text data; and the display unit is used to display the layout text data.
[0246] In the above embodiments of this application, the third acquisition unit includes one of the following: a fifth acquisition subunit and a sixth acquisition subunit.
[0247] The fifth acquisition subunit is used to acquire the layout method corresponding to the target object and obtain the second layout method; the sixth acquisition subunit is used to acquire the selected layout method among the multiple output layout methods and obtain the second layout method.
[0248] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0249] Example 6
[0250] According to embodiments of this application, a text processing apparatus for implementing the above-described text processing method is also provided, such as... Figure 12 As shown, the device 1200 includes: a receiving module 1202, an identification module 1204, a determining module 1206, a generating module 1208, and an output module 1210.
[0251] The receiving module 1202 is used to receive a text image; the recognition module 1204 is used to recognize the text image and obtain the recognition results of multiple characters in the text image, wherein the recognition results include: the content information corresponding to each character and the position information of each character in the text image; the determining module 1206 is used to determine the sorting relationship between multiple characters based on the position information of multiple characters; the generating module 1208 is used to generate text data corresponding to the text image based on the sorting relationship between multiple characters and the content information corresponding to multiple characters; and the output module 1210 is used to output the text data.
[0252] It should be noted that the receiving module 1202, the identifying module 1204, the determining module 1206, the generating module 1208, and the output module 1210 correspond to steps S902 to S910 in Embodiment 3, and the five modules have the same instances and application scenarios as the corresponding steps, but are not limited to the above-mentioned embodiment 1. It should be noted that the above-mentioned modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0253] In the above embodiments of the present application, the determining module comprises a first determining unit, a second determining unit, and a summarizing unit.
[0254] The first determining unit is configured to determine the positional relationship between any two characters based on the positional information of the any two characters; the second determining unit is configured to determine the sorting relationship between any two characters based on the positional relationship between the any two characters and the first layout mode corresponding to the text image; and the summarizing unit is configured to summarize the sorting relationship between any two characters to obtain the sorting relationship between the plurality of characters.
[0255] In the above embodiments of the present application, the determining module further comprises a third determining unit and a fourth determining unit.
[0256] The third determining unit is configured to determine whether the plurality of characters contain a target character based on the positional information of the plurality of characters, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; the fourth determining unit is configured to determine the sorting relationship between any two characters based on the positional relationship between any two first characters, the positional relationship between the first character and the target character, the positional relationship between any two target characters, and the first layout mode, in a case where the plurality of characters contain the target character, wherein the first character is a character in the plurality of characters located in the vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and the second determining unit is further configured to determine the sorting relationship between any two characters based on the positional relationship between the any two characters and the first layout mode, in a case where the plurality of characters do not contain the target character.
[0257] In the above embodiments of the present application, the third determining unit comprises a first obtaining subunit, a first determining subunit, and a second determining subunit.
[0258] The first obtaining subunit is configured to obtain the number of characters adjacent to each character in the vertical direction based on the positional information of the plurality of characters to obtain the number of characters corresponding to each character; the first determining subunit is configured to determine the second character as the target character in a case where the number of characters corresponding to the second character is greater than or equal to a preset number, wherein the second character is any one character in the plurality of characters; and the second determining subunit is configured to determine that the plurality of characters do not contain the target character in a case where the number of characters corresponding to the plurality of characters are all less than the preset number.
[0259] In the above embodiments of the present application, the third determining unit further comprises a second obtaining subunit, a third determining subunit and a fourth determining subunit.
[0260] The second obtaining subunit is configured to obtain a character adjacent to the second character in the vertical direction, to obtain an adjacent character corresponding to the second character; the third determining subunit is configured to determine a first coincidence degree of the second character and the adjacent character based on position information of the second character and the adjacent character; the second determining subunit is further configured to determine that the second character is the target character if the first coincidence degree is greater than a first preset coincidence degree; and the fourth determining subunit is configured to determine that the second character is not the target character if the first coincidence degree is less than the first preset coincidence degree.
[0261] In the above embodiments of the present application, the first obtaining subunit is further configured to obtain a candidate character located in the vertical direction of each character based on the position information of the plurality of characters; determine a first target region of each character and the candidate character in the vertical direction based on position information of each character and the candidate character; determine whether there is another character in the first target region based on the position information of the plurality of characters; and obtain the number of candidate characters to obtain the number of characters corresponding to each character if there is no other character in the first target region.
[0262] In the above embodiments of the present application, the position information at least comprises a first coordinate in the horizontal direction and a width, wherein the third determining subunit is further configured to obtain a minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determine a second target region of the second character and the adjacent character in the vertical direction; and obtain a ratio of the width of the second target region to the first width to obtain the first coincidence degree.
[0263] In the above embodiments of the present application, the fourth determining unit comprises a fifth determining subunit, a sixth determining subunit and a summary subunit.
[0264] The fifth determining subunit is configured to determine a sorting relationship of any two first characters based on the positional relationship of the any two first characters, the positional relationship of the first character and the target character, and the first layout mode; the sixth determining subunit is configured to determine a sorting relationship of any two target characters based on the positional relationship of the any two target characters and the first layout mode; and the summary subunit is configured to summarize the sorting relationship of the any two first characters and the sorting relationship of the any two target characters to obtain the sorting relationship of any two characters.
[0265] In the above embodiments of the present application, the generation module comprises a generation unit.
[0266] The generating unit is configured to splice the content information corresponding to the plurality of characters according to the ordering relationship among the plurality of characters to generate the text data, and in a case where the plurality of characters include a target character, add preset identification information before and / or after the content information corresponding to the target character.
[0267] In the above embodiments of the present application, the device further includes a judging module and a merging module.
[0268] The judging module is configured to judge whether any two characters satisfy a second preset condition based on the position information of the plurality of characters; the merging module is configured to merge any two characters if the any two characters satisfy the second preset condition to obtain a plurality of merged character blocks; and the determining module is further configured to determine an ordering relationship among the plurality of character blocks based on the position information of the plurality of character blocks; and the generating module is configured to generate text data based on the ordering relationship among the plurality of character blocks.
[0269] In the above embodiments of the present application, the judging module includes a fifth determining unit, a judging unit, a sixth determining unit and a seventh determining unit.
[0270] The fifth determining unit is configured to determine a size relative ratio, a second coincidence degree and a relative distance of any two characters based on the position information of the any two characters; the judging unit is configured to judge whether the any two characters are adjacent in a vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; the sixth determining unit is configured to determine that the any two characters satisfy the second preset condition if the any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance; and the seventh determining unit is configured to determine that the any two characters do not satisfy the second preset condition if the any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance.
[0271] In the above embodiments of the present application, the position information at least includes a width, and the fifth determining unit includes a third obtaining subunit and a fourth obtaining subunit.
[0272] The third obtaining subunit is configured to obtain a ratio of the width of any two characters to obtain two ratios; and the fourth obtaining subunit is configured to obtain a minimum ratio of the two ratios to obtain the size relative ratio.
[0273] In the above embodiments of the present application, the position information at least includes a second coordinate in a vertical direction and a width, and the fifth determining unit includes a fifth obtaining subunit, a seventh determining subunit and a sixth obtaining subunit.
[0274] The fifth acquisition subunit is configured to acquire a minimum width in widths of any two characters, to obtain a second width; the seventh determination subunit is configured to determine a third target region of the any two characters in a vertical direction; and the sixth acquisition subunit is configured to acquire a ratio of a height of the third target region to the second width, to obtain a relative distance.
[0275] In the above embodiments of the present application, the device further includes an acquisition module.
[0276] The acquisition module is configured to acquire response data corresponding to the text data, wherein the response data is obtained by modifying the text data.
[0277] In the above embodiments of the present application, the output module includes an acquisition unit, a layout unit and an output unit.
[0278] The acquisition unit is configured to acquire a second layout mode; the layout unit is configured to layout the text data according to the second layout mode, to obtain laid-out text data; and the output unit is configured to output the laid-out text data.
[0279] In the above embodiments of the present application, the acquisition unit includes one of a seventh acquisition subunit and an eighth acquisition subunit.
[0280] The seventh acquisition subunit is configured to acquire a layout mode corresponding to a target object, to obtain the second layout mode; and the eighth acquisition subunit is configured to acquire a selected layout mode from a plurality of output layout modes, to obtain the second layout mode.
[0281] It should be noted that the preferred implementation schemes and embodiments involved in the above embodiments of the present application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0282] Embodiment 7
[0283] According to the embodiments of the present application, a text processing system is further provided, including:
[0284] a processor; and
[0285] a memory connected with the processor, configured to provide the processor with instructions for processing the following processing steps: acquiring a text image; recognizing the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and generating text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0286] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same scheme, application scenario and implementation process as provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0287] Embodiment 8
[0288] According to the embodiments of the present application, a text processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0289] Figure 13 is a flowchart of a fourth text processing method according to an embodiment of the present application, as shown in Figure 13 The method comprises the following steps:
[0290] Step S1302, obtaining an ancient book image.
[0291] The ancient book image in the above steps can be an image of each page of the ancient book, which can be directly obtained by photographing the pages of the ancient book, or can also be obtained by intercepting video frames. The video here is a video taken during the reading of the ancient book, which contains the content of each page.
[0292] The ancient book image contains all the text in the page, and the text in the page is typeset according to the typesetting mode of the ancient book, which can be "from top to bottom, from right to left".
[0293] Step S1304, identifying the ancient book image to obtain the recognition result of the plurality of texts in the ancient book image, wherein the recognition result comprises: content information corresponding to each text, and position information of each text in the ancient book image.
[0294] Due to the difference in typesetting mode between the ancient book and the modern book, newspaper and the like, the existing text serialization scheme cannot be directly used to predict the order of a single text by using spatial rules, text rules or deep learning network. In an optional embodiment, the OCR technology can be used to identify each text in the ancient book image to identify the specific text content and the specific position of each text in the ancient book image.
[0295] In the embodiments of the present application, the specific position of each character can be represented by the starting point coordinates (including the horizontal coordinates and the vertical coordinates, i.e., the x coordinates and the y coordinates), the width, and the height, and therefore, the position information in the above step can include the starting point x coordinates, the starting point y coordinates, the width, and the height. The starting point here can refer to the pixel point at the lower left corner of the character, and the coordinate origin can be the pixel point at the lower left corner of the ancient book image, but is not limited thereto, and can be determined according to the calculation needs.
[0296] In step S1306, the ordering relationship between the plurality of characters is determined based on the position information of the plurality of characters.
[0297] The ordering relationship in the above step can refer to the order of different characters in the reading process. For example, assuming that the ordering relationship between character a and character b is that character a is arranged in front of character b, then in the reading process, the user first reads character a and then reads character b.
[0298] It should be noted that, in order to accurately enter all the recognized characters, it is necessary to first determine the reading order between different characters, and then enter different characters according to the reading order to obtain the electronic data corresponding to the ancient book image. In an optional embodiment, the position relationship between two characters can be analyzed based on the position information of all the characters, and then all the ordering relationships between the characters can be determined by comprehensively analyzing all the position relationships between two characters. In another optional embodiment, in order to simplify the analysis process, a text order prediction scheme based on deep learning can be used to input the position information and the content information of each character into a neural network model for prediction, and a partial order relationship matrix of all the characters can be obtained, i.e., the ordering relationship between all the characters.
[0299] In step S1308, the ancient book text corresponding to the ancient book image is generated based on the ordering relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0300] In the above embodiments of the present application, determining the ordering relationship between the plurality of characters based on the position information of the plurality of characters includes: determining the position relationship between any two characters based on the position information of the any two characters; determining the ordering relationship between the any two characters based on the position relationship between the any two characters and the first typesetting mode corresponding to the ancient book image; and summarizing the ordering relationship between the any two characters to obtain the ordering relationship between the plurality of characters.
[0301] The position relationship in the above step can refer to eight position relationships, such as "upper left, upper, upper right, left, right, lower left, lower, and lower right".
[0302] The first typesetting mode in the above step can refer to the typesetting mode of the ancient book, and specifically can be "from top to bottom and from right to left".
[0303] In the above embodiments of the present application, before determining the ordering relationship of any two characters based on the position relationship of any two characters and the first typesetting mode corresponding to the ancient book image, the method further comprises: determining whether the target character is contained in the plurality of characters based on the position information of the plurality of characters, wherein the target character is a character in the plurality of characters that meets the first preset condition; in the case that the plurality of characters contains the target character, determining the ordering relationship of any two characters based on the position relationship of any two first characters, the position relationship of the first character and the target character, the position relationship of any two target characters and the first typesetting mode, wherein the first character is a character in the plurality of characters located in the vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; in the case that the plurality of characters does not contain the target character, determining the ordering relationship of any two characters based on the position relationship of any two characters and the first typesetting mode.
[0304] The target character in the above step can refer to "big characters" in the ancient book image, such as Figure 4 As shown, the width of the big character is often twice that of the ordinary character, and the reading order is to read the two columns of characters above the big character first, then read the big character, and finally read the two columns of characters below the big character. Therefore, when the ancient book image contains a big character, the ordering relationship between the characters cannot be determined based on the position relationship alone, and the position information of the big character needs to be combined to determine the ordering relationship.
[0305] In the above embodiments of the present application, determining whether the target character is contained in the plurality of characters based on the position information of the plurality of characters comprises: obtaining the number of characters adjacent to each character in the vertical direction based on the position information of the plurality of characters to obtain the character number corresponding to each character; in the case that the character number corresponding to the second character is greater than or equal to the preset number, determining that the second character is the target character, wherein the second character is any one character in the plurality of characters; in the case that the character numbers corresponding to the plurality of characters are all less than the preset number, determining that the plurality of characters do not contain the target character.
[0306] The preset number in the above step can be a number determined according to the characteristics of the big character, for example, for an ancient book, since the width of the big character is often twice the width of other ordinary characters, i.e., there are a total of four columns of characters in the vertical direction of the big character, therefore, the value of the preset number can be 4, but is not limited thereto, and can be determined according to the actual recognized ancient book image.
[0307] In the above embodiments of the present application, in a case where the number of characters corresponding to the second character is greater than or equal to the preset number, the method further comprises: obtaining a character adjacent to the second character in the vertical direction to obtain an adjacent character corresponding to the second character; determining a first coincidence degree of the second character and the adjacent character based on the position information of the second character and the adjacent character; in a case where the first coincidence degree is greater than a first preset coincidence degree, determining that the second character is the target character; and in a case where the first coincidence degree is less than the first preset coincidence degree, determining that the second character is not the target character.
[0308] The first preset coincidence degree in the above step can be a threshold value determined according to experiments, but the threshold value can be modified as needed, for example, in the embodiments of the present application, 0.6 is taken as an example for illustration.
[0309] In the above embodiments of the present application, based on the position information of the plurality of characters, the number of characters adjacent to each character in the vertical direction is obtained to obtain the number of characters corresponding to each character, which comprises: based on the position information of the plurality of characters, candidate characters located in the vertical direction of each character are obtained; based on the position information of each character and the candidate characters, a first target area of each character and the candidate characters in the vertical direction is determined; based on the position information of the plurality of characters, it is judged whether there are other characters in the first target area; if there are no other characters in the first target area, the number of candidate characters is obtained to obtain the number of characters corresponding to each character.
[0310] It should be noted that the condition for determining the vertical neighbors of each character is to judge whether there are other characters between the two characters in the adjacent area, if not, it is determined that the two characters are neighbors, that is, the two characters are adjacent; if there are, it is determined that the two characters are not neighbors, that is, the two characters are not adjacent. Therefore, the first target area described above can refer to the adjacent area between each character and the candidate character, such as the area shown by the dashed line box in Figure 5 .
[0311] In the above embodiments of the present application, the position information at least includes a first coordinate in the horizontal direction and a width, wherein based on the position information of the second character and the adjacent character, the first coincidence degree of the second character and the adjacent character is determined, which comprises: obtaining the minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determining a second target area of the second character and the adjacent character in the vertical direction; obtaining the ratio of the width of the second target area to the first width to obtain the first coincidence degree.
[0312] In the above embodiments of the present application, the determining of the sorting relationship between any two characters based on the position relationship between any two first characters, the position relationship between the first character and the target character, the position relationship between any two target characters, and the first typesetting manner comprises: determining the sorting relationship between any two first characters based on the position relationship between any two first characters, the position relationship between the first character and the target character, and the first typesetting manner; determining the sorting relationship between any two target characters based on the position relationship between any two target characters and the first typesetting manner; and summarizing the sorting relationship between any two first characters and the sorting relationship between any two target characters to obtain the sorting relationship between any two characters.
[0313] In the above embodiments of the present application, the generating of the ancient book text corresponding to the ancient book image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters comprises: splicing the content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate the ancient book text, wherein, in the case that the plurality of characters include a target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0314] The preset identification information in the above step can be a line break, but is not limited thereto, and can also be other ways that can distinguish large characters from other ordinary characters.
[0315] In the above embodiments of the present application, after the ancient book image is recognized and the recognition result of the plurality of characters in the ancient book image is obtained, the method further comprises: determining whether any two characters satisfy a second preset condition based on the position information of the plurality of characters; if any two characters satisfy the second preset condition, merging the any two characters to obtain a plurality of merged character blocks; determining the sorting relationship between the plurality of character blocks based on the position information of the plurality of character blocks; and generating the ancient book text based on the sorting relationship between the plurality of character blocks.
[0316] The second preset condition in the above step can refer to a condition under which two single characters can be merged, and specifically includes the following conditions: the two single characters are vertically adjacent; because the size of the large character in the ancient book image is different from the size of the ordinary character, and the reading order of the large character is different from that of the ordinary character, in order to avoid the influence of the large character on the merged character block, the ordinary character and the large character are not merged, and therefore, it is necessary to determine whether the relative size of the two single characters is less than a certain threshold; the vertical coincidence degree of the two single characters is greater than a certain threshold; and the relative distance of the two single characters in the vertical direction is less than a certain threshold.
[0317] In the above embodiments of the present application, the determining whether the two arbitrary characters satisfy the second preset condition based on the position information of the characters includes: determining the size relative ratio, the second coincidence degree and the relative distance of the two arbitrary characters based on the position information of the two arbitrary characters; determining whether the two arbitrary characters are adjacent in the vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; if the two arbitrary characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance, it is determined that the two arbitrary characters satisfy the second preset condition; if the two arbitrary characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance, it is determined that the two arbitrary characters do not satisfy the second preset condition.
[0318] The preset relative ratio, the second preset coincidence degree and the preset distance in the above steps can be threshold values determined according to experiments, but the threshold values can be modified as needed, for example, in the embodiments of the present application, the preset relative ratio is 0.8, the second preset coincidence degree is 0.75, and the preset distance is 0.5.
[0319] In the above embodiments of the present application, the position information at least includes a width, and the determining the size relative ratio of the two arbitrary characters based on the position information of the characters includes: obtaining the ratio of the widths of the two arbitrary characters to obtain two ratios; obtaining the minimum ratio of the two ratios to obtain the size relative ratio.
[0320] In the above embodiments of the present application, the position information at least includes a second coordinate in the vertical direction and a width, and the determining the relative distance of the two arbitrary characters based on the position information of the characters includes: obtaining the minimum width of the widths of the two arbitrary characters to obtain a second width; determining a third target area of the two arbitrary characters in the vertical direction; obtaining the ratio of the height of the third target area to the second width to obtain the relative distance.
[0321] In the above embodiments of the present application, after the ancient book text corresponding to the ancient book image is generated, the method further includes: outputting the ancient book text; obtaining response data corresponding to the ancient book text, wherein the response data is obtained by modifying the ancient book text.
[0322] In the above embodiments of the present application, the outputting the ancient book text includes: obtaining a second typesetting mode; typesetting the ancient book text according to the second typesetting mode to obtain a typeset ancient book text; and outputting the typeset ancient book text.
[0323] The second layout mode in the above step can be a common layout mode, for example, a layout mode of "from top to bottom and from left to right", but is not limited thereto, and can be other different types of layout modes.
[0324] In the above embodiment of the present application, the second layout mode is obtained in one of the following ways: obtaining the layout mode corresponding to the target object to obtain the second layout mode; and obtaining the selected layout mode from the multiple layout modes output to obtain the second layout mode.
[0325] The target object in the above step can refer to a user uploading the ancient book image, or a user checking and proofreading the ancient book text, which is not limited in the present application.
[0326] It should be noted that the preferred embodiments involved in the above embodiment of the present application have the same application scenarios and implementation processes as the scheme provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.
[0327] Embodiment 9
[0328] According to the embodiments of the present application, a text processing device for implementing the above text processing method is also provided, as shown in the figure, the device 1400 includes an acquisition module 1402, an identification module 1404, a determination module 1406 and a generation module 1408. Figure 14
[0329] The acquisition module 1402 is configured to acquire an ancient book image; the identification module 1404 is configured to identify the ancient book image to obtain the recognition results of multiple characters in the ancient book image, wherein the recognition results include the content information corresponding to each character and the position information of each character in the ancient book image; the determination module 1406 is configured to determine the sorting relationship between the multiple characters based on the position information of the multiple characters; and the generation module 1408 is configured to generate an ancient book text corresponding to the ancient book image based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters.
[0330] It should be noted that the above acquisition module 1402, identification module 1404, determination module 1406 and generation module 1408 correspond to steps S1302 to S1308 in Embodiment 8, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in the above Embodiment 1. It should be noted that the above modules as part of the device can run in the computer terminal 10 provided in Embodiment 1.
[0331] In the above embodiment of the present application, the determination module includes a first determination unit, a second determination unit and a summary unit.
[0332] The first determining unit is configured to determine a position relationship between any two characters based on position information of the any two characters; the second determining unit is configured to determine a sorting relationship between the any two characters based on the position relationship between the any two characters and the first typesetting mode corresponding to the ancient book image; and the collecting unit is configured to collect the sorting relationship between the any two characters to obtain the sorting relationship between the plurality of characters.
[0333] In the above embodiments, the determining module further includes a third determining unit and a fourth determining unit.
[0334] The third determining unit is configured to determine whether the plurality of characters contain a target character based on position information of the plurality of characters, where the target character is a character in the plurality of characters that satisfies a first preset condition; the fourth determining unit is configured to, in a case where the plurality of characters contain the target character, determine a sorting relationship between any two characters based on a position relationship between the any two characters, a position relationship between a first character and the target character, a position relationship between any two target characters, and the first typesetting mode, where the first character is a character in the plurality of characters that is located in a vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and the second determining unit is further configured to, in a case where the plurality of characters do not contain the target character, determine the sorting relationship between the any two characters based on the position relationship between the any two characters and the first typesetting mode.
[0335] In the above embodiments, the third determining unit includes a first obtaining subunit, a first determining subunit, and a second determining subunit.
[0336] The first obtaining subunit is configured to obtain a number of characters adjacent to each character in a vertical direction based on position information of the plurality of characters to obtain a character number corresponding to each character; the first determining subunit is configured to determine that a second character is a target character in a case where the character number corresponding to the second character is greater than or equal to a preset number, where the second character is any one of the plurality of characters; and the second determining subunit is configured to determine that the plurality of characters do not contain the target character in a case where the character numbers corresponding to the plurality of characters are all less than the preset number.
[0337] In the above embodiments, the third determining unit further includes a second obtaining subunit, a third determining subunit, and a fourth determining subunit.
[0338] The second obtaining subunit is configured to obtain a character adjacent to the second character in a vertical direction, to obtain an adjacent character corresponding to the second character.
[0339] In the above embodiments, the first obtaining subunit is further configured to obtain, based on the position information of the plurality of characters, a candidate character located in the vertical direction of each character; determine, based on the position information of each character and the candidate character, a first target region of each character and the candidate character in the vertical direction; determine, based on the position information of the plurality of characters, whether there is another character in the first target region; if there is no other character in the first target region, obtain the number of candidate characters to obtain the number of characters corresponding to each character.
[0340] In the above embodiments, the position information at least includes a first coordinate and a width in the horizontal direction, wherein the third determining subunit is further configured to obtain a minimum width of the width of the second character and the width of the adjacent character to obtain a first width; determine a second target region of the second character and the adjacent character in the vertical direction; obtain a ratio of the width of the second target region to the first width to obtain the first coincidence degree.
[0341] In the above embodiments, the fourth determining unit includes a fifth determining subunit, a sixth determining subunit, and a summarizing subunit.
[0342] The fifth determining subunit is configured to determine, based on the positional relationship between any two first characters, the positional relationship between the first character and the target character, and the first typesetting mode, the sorting relationship between any two first characters; the sixth determining subunit is configured to determine, based on the positional relationship between any two target characters and the first typesetting mode, the sorting relationship between any two target characters; and the summarizing subunit is configured to summarize the sorting relationship between any two first characters and the sorting relationship between any two target characters to obtain the sorting relationship between any two characters.
[0343] In the above embodiments, the generating module includes a generating unit.
[0344] The generating unit is configured to splice the content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate the ancient book text, wherein in a case where the plurality of characters include the target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0345] In the above embodiments of the present application, the device further comprises a judging module and a merging module.
[0346] The judging module is configured to determine whether any two characters satisfy a second preset condition based on the position information of the plurality of characters; the merging module is configured to merge any two characters to obtain a plurality of merged character blocks if the any two characters satisfy the second preset condition; the determining module is further configured to determine a sorting relationship between the plurality of character blocks based on the position information of the plurality of character blocks; and the generating module is configured to generate the ancient text based on the sorting relationship between the plurality of character blocks.
[0347] In the above embodiments of the present application, the judging module comprises a fifth determining unit, a judging unit, a sixth determining unit and a seventh determining unit.
[0348] The fifth determining unit is configured to determine a size relative ratio, a second coincidence degree and a relative distance of any two characters based on the position information of the any two characters; the judging unit is configured to determine whether the any two characters are adjacent in a vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; the sixth determining unit is configured to determine that the any two characters satisfy the second preset condition if the any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance; and the seventh determining unit is configured to determine that the any two characters do not satisfy the second preset condition if the any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance.
[0349] In the above embodiments of the present application, the position information at least comprises a width, and the fifth determining unit comprises a third obtaining subunit and a fourth obtaining subunit.
[0350] The third obtaining subunit is configured to obtain a ratio of the width of any two characters to obtain two ratios; and the fourth obtaining subunit is configured to obtain a minimum ratio of the two ratios to obtain the size relative ratio.
[0351] In the above embodiments of the present application, the position information at least comprises a second coordinate in a vertical direction and a width, and the fifth determining unit comprises a fifth obtaining subunit, a seventh determining subunit and a sixth obtaining subunit.
[0352] The fifth obtaining subunit is configured to obtain a minimum width of the width of any two characters to obtain a second width; the seventh determining subunit is configured to determine a third target region of the any two characters in the vertical direction; and the sixth obtaining subunit is configured to obtain a ratio of a height of the third target region to the second width to obtain the relative distance.
[0353] In the above embodiments of the present application, the device further comprises an output module.
[0354] The output module is configured to output the ancient book text. The acquisition module is further configured to acquire response data corresponding to the ancient book text, wherein the response data is obtained by modifying the ancient book text.
[0355] In the above embodiments of the present application, the output module comprises an acquisition unit, a typesetting unit and an output unit.
[0356] The acquisition unit is configured to acquire the second typesetting mode. The typesetting unit is configured to typeset the ancient book text according to the second typesetting mode to obtain the typeset ancient book text. The output unit is configured to output the typeset ancient book text.
[0357] In the above embodiments of the present application, the acquisition unit comprises one of a seventh acquisition subunit and an eighth acquisition subunit.
[0358] The seventh acquisition subunit is configured to acquire the typesetting mode corresponding to the target object to obtain the second typesetting mode. The eighth acquisition subunit is configured to acquire a selected typesetting mode from the output multiple typesetting modes to obtain the second typesetting mode.
[0359] It should be noted that the preferred embodiments involved in the above embodiments of the present application have the same application scenarios and implementation processes as the schemes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0360] Embodiment 10
[0361] The embodiments of the present application can provide a computer terminal, which can be any one of the computer terminal devices in the computer terminal group. Alternatively, in the present embodiment, the above-mentioned computer terminal can be replaced by a terminal device such as a mobile terminal.
[0362] Alternatively, in the present embodiment, the above-mentioned computer terminal can be located in at least one network device of the multiple network devices of the computer network.
[0363] In the present embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the text processing method: acquiring a text image; recognizing the text image to obtain the recognition result of multiple characters in the text image, wherein the recognition result comprises the content information corresponding to each character and the position information of each character in the text image; determining the sorting relationship between the multiple characters based on the position information of the multiple characters; and generating the text data corresponding to the text image based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters.
[0364] Alternatively, Figure 15is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in Figure 15 The computer terminal A can include one or more (only one is shown in the figure) processors 1502 and a memory 1504.
[0365] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the text processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the text processing method described above. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal A through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0366] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining a text image; identifying the text image to obtain recognition results of a plurality of characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and generating text data corresponding to the text image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0367] Optionally, the processor can further execute program codes of the following steps: determining a position relationship between any two characters based on the position information of the any two characters; determining the sorting relationship between the any two characters based on the position relationship between the any two characters and a first layout mode corresponding to the text image; and summarizing the sorting relationship between the any two characters to obtain the sorting relationship between the plurality of characters.
[0368] Optionally, the processor can further execute program codes of the following steps: determining whether the target character is included in the plurality of characters based on the position information of the plurality of characters, wherein the target character is a character in the plurality of characters satisfying a first preset condition; determining the ordering relationship between any two characters based on the positional relationship between any two first characters, the positional relationship between the first character and the target character, the positional relationship between any two target characters, and the first layout mode, in a case where the plurality of characters includes the target character, wherein the first character is a character in the plurality of characters located in a vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and determining the ordering relationship between any two characters based on the positional relationship between any two characters and the first layout mode, in a case where the plurality of characters does not include the target character.
[0369] Optionally, the processor can further execute program codes of the following steps: obtaining the number of characters adjacent to each character in a vertical direction based on the position information of the plurality of characters, to obtain the character number corresponding to each character; determining the second character as the target character in a case where the character number corresponding to the second character is greater than or equal to a preset number, wherein the second character is any one of the plurality of characters; and determining that the plurality of characters does not include the target character in a case where the character number corresponding to each of the plurality of characters is less than the preset number.
[0370] Optionally, the processor can further execute program codes of the following steps: obtaining the adjacent character corresponding to the second character based on the position information of the second character and the adjacent character in a case where the character number corresponding to the second character is greater than or equal to the preset number; determining a first coincidence degree between the second character and the adjacent character based on the position information of the second character and the adjacent character; determining the second character as the target character in a case where the first coincidence degree is greater than a first preset coincidence degree; and determining that the second character is not the target character in a case where the first coincidence degree is less than the first preset coincidence degree.
[0371] Optionally, the processor can further execute program codes of the following steps: obtaining the candidate character located in a vertical direction of each character based on the position information of the plurality of characters; determining a first target area in the vertical direction between each character and the candidate character based on the position information of each character and the candidate character; determining whether there is another character in the first target area based on the position information of the plurality of characters; and obtaining the number of candidate characters to obtain the character number corresponding to each character if there is no other character in the first target area.
[0372] Optionally, the processor can further execute program codes of the following steps: obtaining a minimum width of the width of the second character and the width of the adjacent character, to obtain a first width; determining a second target area of the second character and the adjacent character in a vertical direction; obtaining a ratio of the width of the second target area to the first width, to obtain a first coincidence degree.
[0373] Optionally, the processor can further execute program codes of the following steps: determining a sorting relationship of any two first characters based on the positional relationship of the any two first characters, the positional relationship of the first character and the target character, and the first layout mode; determining a sorting relationship of any two target characters based on the positional relationship of the any two target characters and the first layout mode; and summarizing the sorting relationship of the any two first characters and the sorting relationship of the any two target characters, to obtain a sorting relationship of any two characters.
[0374] Optionally, the processor can further execute program codes of the following steps: splicing content information corresponding to the plurality of characters according to the sorting relationship among the plurality of characters, to generate the text data, wherein, in a case where the plurality of characters include the target character, preset identification information is added before and / or after the content information corresponding to the target character.
[0375] Optionally, the processor can further execute program codes of the following steps: determining whether any two characters satisfy a second preset condition based on the positional information of the plurality of characters; if the any two characters satisfy the second preset condition, merging the any two characters to obtain a plurality of merged character blocks; determining a sorting relationship among the plurality of character blocks based on the positional information of the plurality of character blocks; and generating the text data based on the sorting relationship among the plurality of character blocks.
[0376] Optionally, the processor can further execute program codes of the following steps: determining a size relative ratio, a second coincidence degree, and a relative distance of any two characters based on the positional information of the any two characters; determining whether the any two characters are adjacent in a vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; if the any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance, it is determined that the any two characters satisfy the second preset condition; and if the any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance, it is determined that the any two characters do not satisfy the second preset condition.
[0377] Optionally, the processor can further execute program codes of the following steps: obtaining a ratio of the width of any two characters, to obtain two ratios; obtaining the minimum ratio of the two ratios, to obtain the size relative ratio.
[0378] Optionally, the processor can further execute program codes of the following steps: obtaining the minimum width of the width of any two characters, to obtain a second width; determining a third target area of the any two characters in the vertical direction; obtaining a ratio of the height of the third target area to the second width, to obtain a relative distance.
[0379] Optionally, the processor can further execute program codes of the following steps: outputting the text data; obtaining response data corresponding to the text data, wherein the response data is obtained by modifying the text data.
[0380] Optionally, the processor can further execute program codes of the following steps: obtaining a second layout mode; performing layout on the text data according to the second layout mode, to obtain the laid-out text data; and outputting the laid-out text data.
[0381] Optionally, the processor can further execute program codes of the following steps: obtaining a second layout mode corresponding to the target object, to obtain the second layout mode; or, obtaining a selected layout mode from the outputted multiple layout modes, to obtain the second layout mode.
[0382] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: displaying a text image; marking recognition results of multiple characters in the text image in the text image, wherein the recognition results of the multiple characters are obtained by recognizing the text image, and the recognition results include content information corresponding to each character and position information of each character in the text image; and displaying text data corresponding to the text image, wherein the text data is generated based on an ordering relationship between the multiple characters and content information corresponding to the multiple characters, and the ordering relationship between the multiple characters is determined based on the position information of the multiple characters.
[0383] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: receiving a text image; recognizing the text image to obtain recognition results of multiple characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining an ordering relationship between the multiple characters based on the position information of the multiple characters; generating text data corresponding to the text image based on the ordering relationship between the multiple characters and the content information corresponding to the multiple characters; and outputting the text data.
[0384] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining an ancient book image; identifying the ancient book image to obtain identification results of a plurality of characters in the ancient book image, wherein the identification results include content information corresponding to each character and position information of each character in the ancient book image; determining a sorting relationship between the plurality of characters based on the position information of the plurality of characters; and generating an ancient book text corresponding to the ancient book image based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0385] By adopting the embodiment of the present application, a text image processing scheme is provided. After the content information and the position information of each character are identified, the identified single character can be sorted to generate corresponding text data, so as to realize automatic input of serialized text without manual input of single character, thereby achieving the technical effects of improving text serialization effect and improving input efficiency, and further solving the technical problems of the related art that the text processing method can only recognize single character and needs manual input of single character, which leads to low input efficiency.
[0386] Those skilled in the art can understand that Figure 15 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, and the like. Figure 15 It does not limit the structure of the electronic device. For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than Figure 15 or have a different configuration from Figure 15 .
[0387] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by programs instructing the related hardware of the terminal device, and the programs can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.
[0388] Embodiment 11
[0389] The embodiment of the present application further provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the text processing method provided by the above embodiment.
[0390] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0391] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a text image; identifying the text image to obtain identification results of a plurality of characters in the text image, wherein the identification results comprise content information corresponding to each character and position information of each character in the text image; determining an ordering relationship between the plurality of characters based on the position information of the plurality of characters; and generating text data corresponding to the text image based on the ordering relationship between the plurality of characters and the content information corresponding to the plurality of characters.
[0392] Optionally, the storage medium is further configured to store program code for performing the following steps: determining a positional relationship between any two characters based on the position information of the any two characters; determining the ordering relationship between the any two characters based on the positional relationship between the any two characters and the first typesetting manner corresponding to the text image; and summarizing the ordering relationship between the any two characters to obtain the ordering relationship between the plurality of characters.
[0393] Optionally, the storage medium is further configured to store program code for performing the following steps: determining whether the plurality of characters contain a target character based on the position information of the plurality of characters, wherein the target character is a character in the plurality of characters that satisfies a first preset condition; in a case where the plurality of characters contain the target character, determining the ordering relationship between the any two characters based on the positional relationship between the any two first characters, the positional relationship between the first character and the target character, the positional relationship between the any two target characters, and the first typesetting manner, wherein the first character is a character in the plurality of characters located in a vertical direction of the target character, and the second character is a character in the plurality of characters other than the first character and the target character; and in a case where the plurality of characters do not contain the target character, determining the ordering relationship between the any two characters based on the positional relationship between the any two characters and the first typesetting manner.
[0394] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a number of characters adjacent to each character in a vertical direction based on the position information of the plurality of characters to obtain a character number corresponding to each character; in a case where the character number corresponding to the second character is greater than or equal to a preset number, determining that the second character is the target character, wherein the second character is any one of the plurality of characters; and in a case where the character numbers corresponding to the plurality of characters are all less than the preset number, determining that the plurality of characters do not contain the target character.
[0395] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining adjacent characters of the second character in the vertical direction, when the number of characters corresponding to the second character is greater than or equal to the preset number; determining the first coincidence degree of the second character and the adjacent character based on the position information of the second character and the adjacent character; determining the second character as the target character when the first coincidence degree is greater than the first preset coincidence degree; determining the second character as not the target character when the first coincidence degree is less than the first preset coincidence degree.
[0396] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining candidate characters in the vertical direction of each character based on the position information of the plurality of characters; determining a first target area of each character and the candidate character in the vertical direction based on the position information of each character and the candidate character; determining whether there are other characters in the first target area based on the position information of the plurality of characters; and obtaining the number of characters corresponding to each character if there are no other characters in the first target area.
[0397] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining the minimum width of the width of the second character and the width of the adjacent character, to obtain the first width; determining a second target area of the second character and the adjacent character in the vertical direction; and obtaining the ratio of the width of the second target area to the first width, to obtain the first coincidence degree.
[0398] Optionally, the storage medium is further configured to store program code for performing the following steps: determining the sorting relationship of any two first characters based on the positional relationship of the two first characters, the positional relationship of the first character and the target character, and the first layout mode; determining the sorting relationship of any two target characters based on the positional relationship of the two target characters and the first layout mode; and summarizing the sorting relationship of any two first characters and the sorting relationship of any two target characters to obtain the sorting relationship of any two characters.
[0399] Optionally, the storage medium is further configured to store program code for performing the following steps: splicing the content information corresponding to the plurality of characters according to the sorting relationship between the plurality of characters to generate text data, wherein preset identification information is added before and / or after the content information corresponding to the target character when the plurality of characters include the target character.
[0400] Optionally, the storage medium is further configured to store program code for performing the following steps: determining, based on the position information of the plurality of characters, whether any two characters satisfy a second preset condition; if any two characters satisfy the second preset condition, merging the any two characters to obtain a plurality of merged character blocks; determining, based on the position information of the plurality of character blocks, a sorting relationship between the plurality of character blocks; and generating the text data based on the sorting relationship between the plurality of character blocks.
[0401] Optionally, the storage medium is further configured to store program code for performing the following steps: determining, based on the position information of any two characters, a size relative ratio, a second coincidence degree and a relative distance of the any two characters; determining whether the any two characters are adjacent in a vertical direction, whether the size relative ratio is less than a preset relative ratio, whether the second coincidence degree is greater than a second preset coincidence degree, and whether the relative distance is less than a preset distance; if the any two characters are adjacent in the vertical direction, the size relative ratio is less than the preset relative ratio, the second coincidence degree is greater than the second preset coincidence degree, and the relative distance is less than the preset distance, determining that the any two characters satisfy the second preset condition; and if the any two characters are not adjacent in the vertical direction, the size relative ratio is greater than the preset relative ratio, the second coincidence degree is less than the second preset coincidence degree, or the relative distance is greater than the preset distance, determining that the any two characters do not satisfy the second preset condition.
[0402] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a ratio of the widths of any two characters to obtain two ratios; and obtaining a minimum ratio of the two ratios to obtain the size relative ratio.
[0403] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a minimum width of the widths of any two characters to obtain a second width; determining a third target area of the any two characters in a vertical direction; and obtaining a ratio of a height of the third target area to the second width to obtain the relative distance.
[0404] Optionally, the storage medium is further configured to store program code for performing the following steps: outputting the text data; and obtaining response data corresponding to the text data, wherein the response data is obtained by modifying the text data.
[0405] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a second layout mode; performing layout on the text data according to the second layout mode to obtain laid-out text data; and outputting the laid-out text data.
[0406] Optionally, the storage medium is further configured to store program code for obtaining a layout mode corresponding to the target object to obtain the second layout mode; or, obtaining a selected layout mode from the output multiple layout modes to obtain the second layout mode.
[0407] Optionally, in the embodiment, the storage medium is configured to store program code for displaying the text image; marking recognition results of multiple characters in the text image, wherein the recognition results of the multiple characters are obtained by recognizing the text image, and the recognition results include content information corresponding to each character and position information of each character in the text image; and displaying text data corresponding to the text image, wherein the text data is generated based on an ordering relationship between the multiple characters and the content information corresponding to the multiple characters, and the ordering relationship between the multiple characters is determined based on the position information of the multiple characters.
[0408] Optionally, in the embodiment, the storage medium is configured to store program code for receiving the text image; recognizing the text image to obtain recognition results of multiple characters in the text image, wherein the recognition results include content information corresponding to each character and position information of each character in the text image; determining an ordering relationship between the multiple characters based on the position information of the multiple characters; generating text data corresponding to the text image based on the ordering relationship between the multiple characters and the content information corresponding to the multiple characters; and outputting the text data.
[0409] Optionally, in the embodiment, the storage medium is configured to store program code for obtaining an ancient book image; recognizing the ancient book image to obtain recognition results of multiple characters in the ancient book image, wherein the recognition results include content information corresponding to each character and position information of each character in the ancient book image; determining an ordering relationship between the multiple characters based on the position information of the multiple characters; and generating ancient book text corresponding to the ancient book image based on the ordering relationship between the multiple characters and the content information corresponding to the multiple characters.
[0410] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0411] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0412] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.
[0413] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0414] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0415] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0416] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A text processing method, characterized in that, include: Get text image; The text image is identified to obtain the recognition results of multiple characters in the text image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the text image; Based on the position information of the multiple characters, determine the sorting relationship between the multiple characters; Based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters, text data corresponding to the text image is generated; The method for determining the sorting relationship between the multiple characters based on their positional information includes: obtaining the number of characters vertically adjacent to each character based on their positional information, thus obtaining the number of characters corresponding to each character; determining whether the multiple characters contain a target character based on the number of characters and a preset number, wherein the target character is a character among the multiple characters that satisfies a first preset condition; determining the sorting relationship between any two characters based on whether the multiple characters contain the target character; and determining the sorting relationship between the multiple characters based on the sorting relationship between any two characters.
2. The method according to claim 1, characterized in that, Determining the ordering relationship among the plurality of characters based on the ordering relationship between any two characters includes: Summarize the sorting relationships of any two characters to obtain the sorting relationships among the multiple characters.
3. The method according to claim 1, characterized in that, Determining the ordering relationship between any two characters based on whether the target character is contained among the plurality of characters includes: When the plurality of characters contain the target character, the ordering relationship of the two characters is determined based on the positional relationship between any two first characters, the positional relationship between the first characters and the target character, the positional relationship between any two target characters, and the first layout method, wherein the first characters are other characters among the plurality of characters besides the target character; In the case where the target text is not contained in any of the multiple texts, the positional relationship between the two texts is determined based on their positional information, and the sorting relationship between the two texts is determined based on their positional relationship and the first layout method.
4. The method according to claim 1, characterized in that, Determining whether the target text is included among the plurality of texts based on the number of texts and the preset number includes: If the number of characters corresponding to the second character is greater than or equal to the preset number, the second character is determined to be the target character, wherein the second character is any one of the plurality of characters; If the number of characters corresponding to the plurality of characters is less than the preset number, it is determined that the plurality of characters does not contain the target character.
5. The method according to claim 4, characterized in that, When the number of characters corresponding to the second character is greater than or equal to a preset number, the method further includes: Obtain the characters that are vertically adjacent to the second character, and thus obtain the adjacent characters corresponding to the second character; Based on the positional information of the second character and the adjacent characters, a first degree of overlap between the second character and the adjacent characters is determined; If the first overlap degree is greater than the first preset overlap degree, the second character is determined to be the target character; If the first overlap is less than the first preset overlap, it is determined that the second text is not the target text.
6. The method according to claim 1, characterized in that, Based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters, the text data corresponding to the text image is generated as follows: According to the sorting relationship between the multiple characters, the content information corresponding to the multiple characters is concatenated to generate the text data. In the case where the multiple characters contain the target character, preset identification information is added before and / or after the content information corresponding to the target character.
7. The method according to any one of claims 1 to 6, characterized in that, After recognizing the text image and obtaining the recognition results of the multiple characters, the method further includes: Based on the position information of the multiple characters, determine whether any two characters satisfy the second preset condition; If any two characters satisfy the second preset condition, then the two characters are merged to obtain multiple merged character blocks; Based on the position information of the multiple text blocks, the sorting relationship between the multiple text blocks is determined; The text data is generated based on the sorting relationship between the multiple text blocks.
8. The method according to claim 1, characterized in that, After generating the text data corresponding to the text image, the method further includes: Output the text data; Obtain the response data corresponding to the text data, wherein the response data is obtained by modifying the text data.
9. The method according to claim 8, characterized in that, The output of the text data includes: Obtain the second typesetting method; The text data is formatted according to the second formatting method to obtain the formatted text data; Output the formatted text data.
10. The method according to claim 9, characterized in that, One of the following methods can be used to obtain the second layout: Obtain the layout method corresponding to the target object to get the second layout method; The selected layout method from the multiple output layout options is the second layout method.
11. A text processing method, characterized in that, include: Display text images; The recognition results of multiple characters in the text image are marked in the text image, wherein the recognition results of the multiple characters are obtained by recognizing the text image, and the recognition results include: content information corresponding to each character, and position information of each character in the text image; The text data corresponding to the text image is displayed, wherein the text data is generated based on the sorting relationship between the plurality of characters and the content information corresponding to the plurality of characters, the sorting relationship between the plurality of characters is determined based on the position information of the plurality of characters, and the text data corresponding to the text image is obtained by processing the text processing method according to any one of claims 1 to 10.
12. A text processing method, characterized in that, include: Receive text images; The text image is identified to obtain the recognition results of multiple characters in the text image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the text image; Based on the position information of the multiple characters, determine the sorting relationship between the multiple characters; Based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters, text data corresponding to the text image is generated; The text data is output, wherein the text data is obtained by processing the text processing method according to any one of claims 1 to 10.
13. A text processing method, characterized in that, include: Acquire images of ancient books; The ancient book image is identified to obtain the recognition results of multiple characters in the ancient book image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the ancient book image; Based on the position information of the multiple characters, determine the sorting relationship between the multiple characters; Based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters, the ancient text corresponding to the ancient book image is generated, wherein the ancient text corresponding to the ancient book image is obtained by processing the text processing method according to any one of claims 1 to 10.
14. The method according to claim 13, characterized in that, Determining the sorting relationship between the multiple characters based on their position information includes: Determine the positional relationship between any two characters based on their positional information; Based on the positional relationship between any two characters and the preset layout, determine the sorting relationship between any two characters; Summarize the sorting relationships of any two characters to obtain the sorting relationships among the multiple characters.
15. The method according to claim 14, characterized in that, Before determining the sorting relationship of any two characters based on their positional relationship and a preset layout, the method further includes: Based on the location information of the plurality of characters, it is determined whether the plurality of characters contain a target character, wherein the target character is a character among the plurality of characters that satisfies a first preset condition; When the plurality of characters contain the target character, the sorting relationship of the two characters is determined based on the positional relationship between any two first characters, the positional relationship between the first characters and the target character, the positional relationship between any two target characters, and the preset layout method, wherein the first characters are other characters among the plurality of characters besides the target character; If the target text is not included in any of the multiple texts, the sorting relationship between the two texts is determined based on their positional relationship and the preset layout.
16. A text processing device, characterized in that, The text processing method applied to any one of claims 1 to 10 includes: The acquisition module is used to acquire text images; The recognition module is used to recognize the text image and obtain the recognition results of multiple characters in the text image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the text image; The determining module is used to determine the sorting relationship between the multiple characters based on their position information; The generation module is used to generate text data corresponding to the text image based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters.
17. A text processing device, characterized in that, The text processing method applied to claim 11 includes: The first display module is used to display text images; A tagging module is used to tag the recognition results of multiple characters in the text image, wherein the recognition results of the multiple characters are obtained by recognizing the text image, and the recognition results include: content information corresponding to each character, and position information of each character in the text image; The second display module is used to display the text data corresponding to the text image, wherein the text data is generated based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters, and the sorting relationship between the multiple characters is determined based on the position information of the multiple characters.
18. A text processing device, characterized in that, The text processing method applied to claim 12 includes: The receiving module is used to receive text images; The recognition module is used to recognize the text image and obtain the recognition results of multiple characters in the text image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the text image; The determining module is used to determine the sorting relationship between the multiple characters based on their position information; The generation module is used to generate text data corresponding to the text image based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters; The output module is used to output the text data.
19. A text processing device, characterized in that, The text processing method applied to claim 13 includes: The acquisition module is used to acquire images of ancient books; The recognition module is used to recognize the ancient book image and obtain the recognition results of multiple characters in the ancient book image, wherein the recognition results include: content information corresponding to each character, and position information of each character in the ancient book image; The determining module is used to determine the sorting relationship between the multiple characters based on their position information; The generation module is used to generate the ancient text corresponding to the ancient text image based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the text processing method according to any one of claims 1 to 15.
21. A computer terminal, characterized in that, The device includes a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, performs the text processing method according to any one of claims 1 to 15.
22. A text processing system, characterized in that, include: processor; as well as A memory, connected to the processor, is used to provide the processor with instructions to perform the following processing steps: acquiring a text image; The text image is identified to obtain recognition results of multiple characters in the text image, wherein the recognition results include: content information corresponding to each character and position information of each character in the text image; based on the position information of the multiple characters, the sorting relationship between the multiple characters is determined; based on the sorting relationship between the multiple characters and the content information corresponding to the multiple characters, text data corresponding to the text image is generated, wherein the text data corresponding to the text image is obtained by processing using the text processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Ancient book document digitization method
CN111507351A