Text processing method, apparatus and system
By identifying and detecting text and layout areas in text images and using deep learning models to generate text data, the problem of low efficiency in single-character recognition in OCR technology is solved, and automatic text entry and efficient processing are achieved.
Patent Information
- Application Number
- CN202110163246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-05
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Existing OCR technology can only recognize single characters and cannot directly input text, resulting in low text input efficiency and requiring manual operation.
By identifying multiple characters and target areas in text images and combining them with layout types, text data is generated. Deep learning models such as CenterNet, FastRCNN, and MaskRCNN are used for layout detection and classification to achieve automated text entry.
It realizes the automatic input of text, improves the input efficiency, reduces manual operations, and improves the layout detection and classification effects.
Smart Images

Figure CN114882209B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition, and more specifically, to a text processing method, device, and system. Background Art
[0002] In "text digitization," text images need to be converted into complete text data and entered into the system. The text here can be, but is not limited to, text from paper media such as posters, newspapers, and magazines, flyers for products on e-commerce platforms, and book texts such as textbooks and ancient books. Currently, OCR (Optical Character Recognition) can be used for image recognition of text images. However, OCR can only detect the position of a single character and recognize the content of the character. It cannot be directly entered, requiring manual entry of individual characters, resulting in low entry efficiency.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a text processing method, device and system to at least solve the technical problem that the text processing method in the related art can only recognize a single character and requires manual input of the single character, resulting in low input efficiency.
[0005] According to one aspect of an embodiment of the present application, a text processing method is provided, including: acquiring a first text image; identifying the first text image to obtain a first recognition result of multiple characters in the first text image; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area; based on the first recognition result and the second recognition result, generating text data corresponding to the first text image.
[0006] According to another aspect of an embodiment of the present application, a text processing method is also provided, including: displaying a first text image; marking a first recognition result of multiple characters in the first text image in the first text image; marking a second recognition result of at least one target area in the first text image in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area; displaying text data corresponding to the first text image, wherein the text data is generated based on the first recognition result and the second recognition result.
[0007] According to another aspect of the embodiments of the present application, a text processing method is also provided, including: receiving a first text image; identifying the first text image to obtain a first recognition result of a plurality of characters in the first text image; detecting the first text image to obtain a second recognition result of at least one target region in the first text image, wherein the second recognition result includes region position information of each target region and a type of a layout corresponding to each target region; generating text data corresponding to the first text image based on the first recognition result and the second recognition result; and outputting the text data corresponding to the first text image.
[0008] According to another aspect of the embodiments of the present application, a text processing method is also provided, including: receiving a first text image; identifying the first text image to obtain a first recognition result of a plurality of characters in the first text image; detecting the first text image to obtain a second recognition result of at least one target region in the first text image, wherein the second recognition result includes region position information of each target region and a type of a layout corresponding to each target region; generating text data corresponding to the first text image based on the first recognition result and the second recognition result; and outputting the text data corresponding to the first text image.
[0009] According to another aspect of the embodiments of the present application, a text processing apparatus is also provided, including: an acquisition module configured to acquire a first text image; an identification module configured to identify the first text image to obtain a first recognition result of a plurality of characters in the first text image; a detection module configured to detect the first text image to obtain a second recognition result of at least one target region in the first text image, wherein the second recognition result includes region position information of each target region and a type of a layout corresponding to each target region; and a generation module configured to generate text data corresponding to the first text image based on the first recognition result and the second recognition result.
[0010] According to another aspect of the embodiments of the present application, a text processing apparatus is also provided, including: a first display module configured to display a first text image; a first marking module configured to mark a first recognition result of a plurality of characters in the first text image in the first text image; a second marking module configured to mark a second recognition result of at least one target region in the first text image in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes region position information of each target region and a type of a layout corresponding to each target region; and a second display module configured to display text data corresponding to the first text image, wherein the text data is generated based on the first recognition result and the second recognition result.
[0011] According to another aspect of an embodiment of the present application, a text processing device is also provided, including: a receiving module for receiving a first text image; a recognition module for recognizing the first text image and obtaining a first recognition result of multiple characters in the first text image; a detection module for detecting the first text image and obtaining a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area; a generation module for generating text data corresponding to the first text image based on the first recognition result and the second recognition result; and an output module for outputting the text data corresponding to the first text image.
[0012] According to another aspect of an embodiment of the present application, a text processing device is also provided, including: an acquisition module for acquiring an ancient book image; an identification module for identifying the ancient book image and obtaining a first recognition result of multiple characters in the ancient book image; a detection module for detecting the ancient book image and obtaining a second recognition result of at least one target area in the ancient book image, wherein the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area; a generation module for generating an ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result.
[0013] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned text processing method.
[0014] According to another aspect of an embodiment of the present application, a computer terminal is further provided, comprising a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the above-mentioned text processing method is executed when the program is run.
[0015] According to another aspect of an embodiment of the present application, a text processing system is also provided, including: a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining a first text image; identifying the first text image to obtain a first recognition result of multiple characters in the first text image; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area; based on the first recognition result and the second recognition result, generating text data corresponding to the first text image.
[0016] In an embodiment of the present application, after acquiring a first text image, the first text image can be identified to obtain a first recognition result of multiple characters, and the first text image can be detected to obtain a second recognition result of at least one target area, and further based on the first recognition result and the second recognition result, text data corresponding to the first text image can be generated to achieve the purpose of digitizing the text. It is easy to notice that after identifying the content information and position information of each character, the layout of the text image can be detected and classified in combination with the layout characteristics of the text image, and the corresponding text data can be generated in combination with the detection and classification results to achieve automatic text entry, without the need for manual entry of individual characters, achieving the technical effect of improving the layout detection and classification effect and improving the entry efficiency, thereby solving the technical problem that the text processing method in the related art can only recognize a single character and requires manual entry of a single character, resulting in low entry efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1a is a schematic diagram of an ancient book cover according to an embodiment of the present application;
[0019] Figure 1b is a schematic diagram of an ancient book catalog according to an embodiment of the present application;
[0020] Figure 1c is a schematic diagram of the main text of an ancient book according to an embodiment of the present application;
[0021] Figure 1d is a schematic diagram of an ancient book illustration according to an embodiment of the present application;
[0022] Figure 1e is a schematic diagram of an ancient book table according to an embodiment of the present application;
[0023] Figure 1f This is a schematic diagram of an ancient book edge according to an embodiment of the present application;
[0024] Figure 1g Schematic diagram of an ancient book eyebrow annotation according to an embodiment of the present application;
[0025] Figure 2 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a text processing method according to an embodiment of the present application;
[0026] Figure 3 is a flowchart of a first text processing method according to an embodiment of the present application;
[0027] Figure 4 is a schematic diagram of an optional interactive interface according to an embodiment of the present application;
[0028] Figure 5 is a flowchart of an optional text processing method according to an embodiment of the present application;
[0029] Figure 6 is a flowchart of a second text processing method according to an embodiment of the present application;
[0030] Figure 7 is a flowchart of a third text processing method according to an embodiment of the present application;
[0031] Figure 8 is a schematic diagram of a first text processing device according to an embodiment of the present application;
[0032] Figure 9 is a schematic diagram of a second text processing device according to an embodiment of the present application;
[0033] Figure 10 is a schematic diagram of a third text processing device according to an embodiment of the present application;
[0034] Figure 11 is a flowchart of a fourth text processing method according to an embodiment of the present application;
[0035] Figure 12 is a schematic diagram of a fourth text processing device according to an embodiment of the present application;
[0036] Figure 13 This is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0039] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0040] Ancient books: may refer to books that are not printed using modern printing technology. They are different from the typesetting methods of modern books, newspapers, etc. The typesetting method of modern books and newspapers is "from top to bottom, from left to right", while the typesetting method of ancient books is "from top to bottom, from right to left".
[0041] OCR: This refers to the process by which an electronic device (such as a scanner or digital camera) examines characters printed on paper and then uses character recognition to translate the shapes into computer text. However, character recognition is limited to single-word recognition.
[0042] Layout: It can refer to the layout format of graphic design publications (books, magazines, newspapers or electronic publications). For ancient books, the layout types mainly include: cover (such as Figure 1a As shown), directory (as Figure 1b As shown), the text (as shown Figure 1c As shown), diagram (as shown Figure 1d As shown), table (as Figure 1e As shown), book edge (as shown Figure 1f As shown in the rectangular box in Figure 1g as shown in the rectangular box in ), etc.
[0043] CenterNet: It is a center-point-based detection network that can locate the target to be detected at a point, that is, the center point of the detection rectangle, and then regress other attributes of the target by finding the center point.
[0044] RCNN: Region-CNN (Convolutional Neural Network), a local convolutional neural network, follows the same principles of traditional object detection, using the same four-step process of bounding box extraction, feature extraction for each bounding box, image classification, and non-maximum value suppression for object detection. However, in the feature extraction step, traditional features are replaced with features extracted by a deep convolutional network.
[0045] FastRCNN: Compared with RCNN, the biggest difference is the unification of target classification and detection box regression fine-tuning in the pooling layer and the fully connected layer.
[0046] MaskRCNN: It is an instance segmentation algorithm that can be used for "target detection", "target instance segmentation", and "target key point detection".
[0047] FCN: Fully Convolutional Networks, which can classify images at the pixel level.
[0048] Example 1
[0049] According to an embodiment of the present application, a text processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0050] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 2 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a text processing method. Figure 2 As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 2 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 2 More or fewer components than shown, or with Figure 2Different configurations shown.
[0051] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0052] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the text processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned text processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0053] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0054] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0055] It should be noted that, in some optional embodiments, the above Figure 2 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 2This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the aforementioned computer device (or mobile device).
[0056] Under the above operating environment, this application provides Figure 3 The text processing method shown. Figure 3 This is a flow chart of the first text processing method according to an embodiment of the present application. Figure 3 As shown, the method may include the following steps:
[0057] Step S302: Acquire a first text image.
[0058] The first text image in the above steps can be an image of different types of text, which can be obtained directly by photographing different types of text, or by capturing video frames. The video here is a video taken during the process of reading different types of text. For example, the text image can be an image obtained by photographing a single page of text such as a poster or a leaflet, or an image obtained by photographing each page of a multi-page text such as a newspaper, magazine, or book. It can also be a video taken during the process of reading a multi-page text such as a newspaper, magazine, or book, which includes the content of each page of text. In the embodiments of the present application, an image of an ancient book is used as an example for explanation.
[0059] The first text image in the above steps may include at least one target area, and the multiple characters contained in the first text image are typeset according to the first typesetting method. The target area here may be an area of different layouts in the first text image. The layout types vary for different types of text. For example, in an ancient book, the layout of the first page is often the cover, and the layout of the second page is often the table of contents.
[0060] The text image contains all the characters in the text. Furthermore, characters in different types of text are typeset in different layouts. Therefore, the first layout mentioned above can refer to the layout of the first text image. For example, using the example of an ancient book, the first layout can specifically be "top to bottom, right to left." However, the first text image not only contains characters in different layouts but also may contain watermarks, labels added during post-processing, and other text unrelated to the ancient book itself.
[0061] It should be noted that a single text image can use multiple layout styles. For example, a poster image can use both a top-to-bottom, left-to-right layout and a top-to-bottom, right-to-left layout. Furthermore, for multi-page text, different layout styles can be used on different pages.
[0062] In an optional embodiment, in order to realize the digitization of text, the text can be photographed and the photographed text image can be transmitted to a corresponding processing device for processing, for example, directly transmitted to the user's computer terminal (e.g., a laptop computer, personal computer, etc.) for processing, or transmitted to a cloud server through the user's computer terminal for processing. It should be noted that because the processing of text images requires a large amount of computing resources, the processing device is described as a cloud server in the embodiments of this application.
[0063] For example, in order to facilitate users to upload text images, an interactive interface can be provided to users, such as Figure 4 As shown, users can click the "Select Image" button to select a text image to be processed from a large number of stored images, or batch select multiple text images and click the "Upload" button to upload the selected text images to the cloud server for processing. In addition, to facilitate the user's confirmation of whether the selected text image is the text image to be processed, the user's selected text image can be displayed in the "Image Display" area. After the user confirms that it is correct, the data can be uploaded by clicking the "Upload" button.
[0064] Step S304 : Recognize the first text image to obtain a first recognition result of the plurality of characters in the first text image, wherein the first recognition result includes: content information corresponding to each character and character position information of each character.
[0065] In an optional embodiment, for the first text image, OCR technology can be used to recognize each character in the first text image to identify the specific text content (i.e., the above-mentioned content information) and the specific position of each character in the text image (i.e., the above-mentioned text position information).
[0066] In the embodiment of the present application, the specific position of each character can be represented by starting point coordinates (including horizontal and vertical coordinates, i.e., x-coordinates and y-coordinates), width, and height. Therefore, the character position information in the above steps may include: starting point x-coordinates, starting point y-coordinates, width, and height. The starting point here can refer to the pixel point in the lower left corner of the character, and the coordinate origin can be the pixel point in the lower left corner of the text image, but is not limited to this and can be determined according to calculation needs.
[0067] Step S306 : Detect the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and the type of layout corresponding to each target area.
[0068] Since different types of text images use different types of layouts, and a text image may use multiple different types of layouts, it is impossible to directly use existing layout analysis algorithms to detect and mark elements such as "headers, footers, charts, tables, titles, and texts." In order to achieve the purpose of detecting and classifying the layout of the first text image, a detection algorithm can be designed based on the layout characteristics of the first text image, and the layout of the first text image can be detected and classified through the detection algorithm to determine the specific location of the layout (i.e., the layout box) and the specific type of the layout. In an optional embodiment, a deep learning detection model can be pre-trained based on the layout characteristics of the first text image, such as models such as CenterNet, FastRCNN, MaskRCNN, or FCN based on image segmentation. The deep learning detection model can predict the layout box and the corresponding layout classification in the first text image.
[0069] In this embodiment of the present application, the specific location of each target area can be represented by the four vertices of a quadrilateral as coordinates, and can be arranged in the order of upper left, upper right, lower right, and lower left. For example, the area location information in the above steps is: (1,1), (100,100), (100,100), (1,100). An example of the second recognition result is: [(1,1), (100,100), (100,100), (1,100, Type: Text].
[0070] Step S308 : generating text data corresponding to the first text image based on the first recognition result and the second recognition result.
[0071] The text data in the above steps may refer to serialized text obtained by sorting the content information corresponding to all the words according to the sorting relationship and layout type.
[0072] In an optional embodiment, in order to realize text entry, after identifying the content information of all the text in the text image and detecting the publication frame and layout classification, all the text identified in the text image is screened based on the layout frame and version classification, and watermarks, labels added during post-processing and other texts that are not related to the book itself are removed. Then all the texts in the same layout classification are merged, and finally all the layout frame texts are structurally merged to obtain the final text data, that is, the electronic result of the text image. In addition, in order to facilitate users to read electronic texts, a second typesetting method can be used to typeset the text data, and the order of all texts is still determined based on the sorting relationship between the texts. The second typesetting method here can be a currently common typesetting method, for example, a "from top to bottom, from left to right" typesetting method; or it can be a typesetting method that determines the user's preference based on the user's reading habits.
[0073] For example, for paper news media text, using newspapers as an example, users can take a photo of each page of the newspaper and upload the newspaper image of each page to the cloud server for processing. To digitize the newspaper, the cloud server can first use OCR technology to recognize each character in the newspaper image, obtaining recognition results for all characters. Furthermore, based on the layout characteristics of the newspaper, the newspaper image can be detected to obtain the regional location information and layout type of all layouts in the newspaper image. Finally, the characters within each layout can be merged according to different layout types to obtain structured data. By summarizing the structured data within all layouts, the corresponding text data of the newspaper can be obtained, that is, the final digitized result.
[0074] For example, for single-page texts such as posters and flyers, using flyers as an example, users can directly take a photo of the flyer and upload the flyer image to the cloud server for processing. To digitize the flyer, the cloud server can first use OCR technology to recognize each character in the flyer image, obtaining recognition results for all characters. Furthermore, based on the layout characteristics of the flyer, the flyer image can be detected to obtain the regional location information and layout type of all layouts in the flyer image. Finally, the text within each layout can be merged according to the different layout types to obtain structured data. By summarizing the structured data within all layouts, the text data corresponding to the flyer can be obtained, that is, the final digitized result.
[0075] For example, for book texts, taking ancient books as an example, users can take photos of each page of the ancient book and upload the ancient book images to the cloud server for processing. In order to realize the digitization of ancient books, the cloud server can first use OCR technology to recognize each character in the ancient book image, obtain the recognition results of all characters, and further combine the layout characteristics of the ancient book to detect the ancient book image to obtain the regional location information and layout type of all layouts in the ancient book image. Finally, the characters in each layout can be merged according to different layout types to obtain structured data. Then, the structured data in all layouts can be summarized to obtain the text data corresponding to the ancient book, that is, the final digitization result.
[0076] Through the solution provided by the above embodiment of the present application, after obtaining the first text image, the first text image can be identified to obtain a first recognition result of multiple characters, and the first text image can be detected to obtain a second recognition result of at least one target area, and further based on the first recognition result and the second recognition result, text data corresponding to the first text image can be generated to achieve the purpose of digitization of the text. It is easy to notice that after identifying the content information and position information of each character, the layout of the text image can be detected and classified in combination with the layout characteristics of the text image, and the corresponding text data can be generated in combination with the detection and classification results to achieve automatic text entry, without the need for manual entry of individual characters, achieving the technical effect of improving the layout detection and classification effect and improving the entry efficiency, thereby solving the technical problem that the text processing method in the related art can only recognize a single character and requires manual entry of a single character, resulting in low entry efficiency.
[0077] In the above embodiment of the present application, detecting the first text image to obtain a second recognition result of at least one target area in the first text image includes: detecting the first text image using a target detection model to obtain a second recognition result.
[0078] The target detection model in the above steps can be CenterNet, FastRCNN, MaskRCNN, FCN, etc. In the embodiment of the present application, the CenterNet model is taken as an example for illustration.
[0079] In an optional embodiment, the first text image may be input into the target detection model to obtain the layout frame contained in the text image, that is, to determine the specific position and type of the layout.
[0080] In the above embodiment of the present application, the first text image is detected using a target detection model to obtain a second recognition result, which includes: inputting the first text image into the target detection model for prediction to obtain a prediction result for each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; and determining the second recognition result based on the prediction result of each pixel.
[0081] The prediction results in the above steps may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0082] Optionally, determining the second recognition result based on the prediction result of each pixel may include: determining at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; performing a regression operation based on the offset result of the at least one target pixel to obtain the second recognition result.
[0083] In an optional embodiment, for the CenterNet model, the input of the model is a single text image, and the output is the position and type of all target areas in the image. Specifically, it is possible to predict whether each pixel in the image is located at the center point of any target area, and obtain the classification result in the above steps, and the offset of each pixel to the upper left, upper right, lower right, and lower left vertices of any target area in the horizontal and vertical directions can be determined to obtain the offset result in the above steps, that is, the layout frame prediction is performed through 8 offset predictions after a two-classification judgment. Finally, the four vertices regressed from the pixel determined to be the center point (that is, the above-mentioned target pixel) can be selected as the final recognition result, that is, the specific coordinates of the four vertices can be determined based on the offset of the pixel of the center point, thereby obtaining the specific position of the target area. In addition, after determining the specific position of the target area, the target area can be classified to determine the specific type of the target area.
[0084] In the above embodiment of the present application, before using the target detection model to detect the first text image and obtain the second recognition result, the method also includes: obtaining multiple second text images; annotating each second text image according to the preset layout classification rules to obtain the annotation results of each second text image, wherein the annotation results include: the location information of the preset area of each second text image, and the type of layout corresponding to the preset area; generating training samples based on the multiple second text images and the annotation results of each second text image; and using the training samples to train the target detection model.
[0085] In an optional embodiment, in order to improve the processing accuracy of the target detection model, training samples can be constructed based on the data characteristics of text images, and the target detection model can be trained using the training samples. For example, still taking ancient books as an example, the training samples can be generated in the following way: first, based on the data characteristics of ancient books, the layout classification of ancient books can be defined, which is divided into seven categories: "cover, table of contents, text, illustrations, tables, book edges, and marginal notes." Then, a large number of ancient book images can be collected, and the target areas and corresponding types in the ancient book images can be annotated according to the defined layout classification of ancient books to obtain the annotation results, thereby generating the corresponding training samples.
[0086] In the above embodiment of the present application, based on the first recognition result and the second recognition result, generating text data corresponding to the first text image includes: matching each character with each target area based on the character position information of each character and the area position information of each target area, and determining the target character contained in each target area; merging the content information corresponding to the target character contained in each target area to generate text data corresponding to each target area; combining the text data corresponding to at least one target area to generate text data corresponding to the first text image.
[0087] In an optional embodiment, during the text digitization process, in order to group the characters in different layouts such as headers, footers, book edges, main text, and marginal notes, all the characters in the first text image can be matched with each layout to determine the area to which each character belongs, that is, all the characters are grouped according to different layouts, and all the characters belonging to the same layout are determined. Then, the text contents of all the characters in the same layout are merged to obtain the serialized text of the same layout, that is, to obtain the text data of the same layout. The specific implementation scheme can determine the reading order between different characters and merge the text contents of different characters according to the reading order. Finally, the text data of all layouts are structurally combined. For example, still taking ancient books as an example, the final digitization result can be "main text area: XXXXXXX; marginal notes area: XXXXXXX; book edge area: XXXXXXX".
[0088] In the above embodiment of the present application, each character is matched with each target area based on the character position information of each character and the area position information of each target area, and the target characters contained in each target area are determined, including: determining the intersection area ratio of each character and each target area based on the character position information of each character and the area position information of each target area; for each character, determining that the target area corresponding to the target intersection area ratio is the target area to which each character belongs, wherein the target intersection area ratio is greater than a preset threshold value, and the target intersection area ratio is the maximum intersection area ratio; based on the target area to which each character belongs, determining the target characters contained in each target area.
[0089] The preset threshold in the above steps may be a threshold for determining whether each character matches each target area, and may be a preset fixed value, such as 0.7, but is not limited thereto, and may be adjusted and determined according to actual needs.
[0090] In an optional embodiment, the purpose of matching characters with target areas can be achieved by calculating the intersection area ratio between each character and each target area, that is, the ratio of the area of the intersection area between each character and each target area to the area of each character. After calculating the intersection area ratio, if the intersection area ratio between a certain character and a certain target area is the largest, and the intersection area ratio is greater than a preset threshold, it can be determined that the character matches the target area, that is, it is determined that the target area contains the character, and the target area is the area to which the character belongs.
[0091] It should be noted that the maximum cross-area ratio here may refer to the maximum cross-area ratio obtained after calculating the cross-area ratios between the text and all target areas.
[0092] In the above embodiment of the present application, determining the intersection area ratio of each character and each target area based on the character position information of each character and the area position information of each target area includes: determining the intersection area of each character and each target area based on the character position information of each character and the area position information of each target area; obtaining the ratio of the area of the intersection area to the area of each character to obtain the intersection area ratio, wherein the area of each character is determined based on the character position information of each character.
[0093] In an optional embodiment, the intersection area between the text and the target area can be determined by coordinates based on the specific position of the text and the specific position of the target area, and then the area of the intersection area can be determined based on the coordinates of the four vertices of the intersection area. Similarly, the area of the text can be determined based on the width and height of the text. Finally, the ratio of the area of the intersection area to the area of the text is calculated to obtain the intersection area ratio.
[0094] In the above embodiment of the present application, the content information corresponding to the target text contained in each target area is merged to generate text data corresponding to each target area, including: determining the sorting relationship between the target texts based on the position information of the target texts; and generating text data corresponding to each target area based on the sorting relationship between the target texts and the content information corresponding to the target texts.
[0095] The sorting relationship in the above steps may refer to the order of different characters during the reading process. For example, assuming that the sorting relationship between character a and character b is that character a comes before character b, the user reads character a first and then character b during the reading process.
[0096] It should be noted that in order to accurately enter all the recognized text, it is first necessary to determine the reading order between different texts, and then enter the different texts according to the reading order to obtain the electronic data corresponding to the text image. In an optional embodiment, the position relationship between the two texts can be directly analyzed based on the position information of all texts, and then all the two-to-two position relationships can be combined to determine the sorting relationship between all texts. In another optional embodiment, in order to simplify the analysis process, a text order prediction scheme based on deep learning can be used, and the position information and content information of each text can be input into the neural network model for prediction, so as to obtain the partial order relationship matrix of all texts, that is, the sorting relationship between all texts.
[0097] In an optional embodiment, in order to realize the electronic entry of text, after identifying the content information of all characters in the same target area, it is necessary to organize the individual characters into a complete serialized text. Therefore, after analyzing and obtaining the sorting relationship between all characters in the same target area, the content information can be sorted according to the sorting relationship to obtain the serialized text (that is, the above-mentioned text data).
[0098] Optionally, the positional relationship between any two target characters can be determined based on their positional information. Furthermore, based on this positional relationship, the ordering relationship between any two target characters can be determined. Finally, the ordering relationships between any two target characters can be summarized to obtain the ordering relationship between all characters within the target area. The positional relationship here can include eight types of positional relationships, such as "upper left, upper, upper right, left, right, lower left, lower, and lower right."
[0099] Furthermore, some texts may contain large characters, which are often twice as wide as regular text. Furthermore, the reading order is to first read the two columns of text above the large characters, then the large characters, and finally the two columns of text below them. Therefore, when a text image contains large characters, the ordering of the characters cannot be determined solely based on their positional relationship; it must be combined with the positional information of the large characters.
[0100] When analyzing the positional relationship between two characters, it is first necessary to determine whether the text image contains "large characters". If it does not contain large characters, the sorting relationship between the two characters is determined directly based on the positional relationship between the two characters. If it contains large characters, then for other characters that are not related to the large characters, the sorting relationship between the two characters can be determined directly based on the positional relationship between the two characters. For characters related to the large characters, that is, characters located above or below the large characters, the sorting relationship between the two characters needs to be determined based on the positional relationship between the two characters, as well as the positional relationship between each character and the large characters.
[0101] Finally, the text can be concatenated in the order of all the characters. When concatenating to a large character, line breaks are added before and after the large character to obtain the final serialized text. It should be noted that if there are multiple large characters in a row, a line break is added before the first large character and after the last large character.
[0102] It should be noted that after determining the ordering relationship between multiple characters, the ordering relationship can be saved so that it can be reused to recognize text images with similar ordering relationships. In addition, for damaged or missing text, the user can provide the damaged or missing characters, and the text can be restored based on the determined ordering relationship between the multiple characters.
[0103] In the above embodiment of the present application, after generating text data corresponding to each target area, the method also includes: outputting the text data corresponding to each target area; receiving response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; combining the text data and / or response data corresponding to at least one target area to generate text data corresponding to the first text image.
[0104] Due to the limited recognition accuracy of OCR technology and the recognition accuracy of the target detection model, the first recognition result or the second recognition result may be incorrect, which in turn may cause errors in the text data corresponding to each target area.
[0105] In order to avoid the above problems, in an optional embodiment, after the cloud server generates the text data corresponding to each layout, it can be sent to the user's computer terminal through the network and displayed on a computer such as Figure 4 The interactive interface shown is for users to review. If the user determines that the text data corresponding to a certain layout contains errors, the user can modify it directly in the interactive interface, obtain response data, and return it to the cloud server for processing. The cloud server can then adjust the OCR technology or object detection model based on the response data fed back by the user, improving recognition accuracy and thus improving the processing performance of the cloud server.
[0106] In another alternative embodiment, if the cloud server does not receive response data, the generated text data corresponding to each layout can be directly combined to generate the final electronic result, that is, the text data corresponding to the first text image is generated by combining the text data corresponding to at least one target region. If the cloud server receives response data corresponding to part of the layout, the received response data and the text data corresponding to the remaining layout can be combined to generate the final electronic result, that is, the text data corresponding to the first text image is generated by combining the text data corresponding to at least one target region and the response data. If the cloud server receives response data corresponding to all layouts, the received response data can be combined to generate the final electronic result, that is, the text data corresponding to the first text image is generated by combining the response data.
[0107] In the above embodiments of the present application, combining the text data corresponding to at least one target region to generate the text data corresponding to the first text image includes: combining the type of each target region and the text data corresponding to each target region to generate structured data of each target region; and combining the structured data of at least one target region to obtain the text data corresponding to the first text image.
[0108] In an alternative embodiment, the name of the target region, that is, the type of the target region, and the serialized text in the target region can be structured to obtain structured data of the target region, and then all the structured data of the target region can be combined to obtain the final text data.
[0109] In the above embodiments of the present application, combining the structured data of at least one target region to obtain the text data corresponding to the first text image includes: obtaining a target structured template; and combining the structured data of at least one target region according to the target structured template to generate the text data corresponding to the first text image.
[0110] The target structured template in the above steps includes the ordering manner of different types of layouts and the layout manner of different characters in each layout, which can be a common structured template preset by the system, a structured template determined in advance according to the reading habits of the user, or a structured template uploaded or selected by the user, but is not limited thereto.
[0111] In an alternative embodiment, since the ordering manner of the layout in the first text image and the layout manner of different characters in each layout are different, in order to facilitate the user to view, all the layouts and the structured data in the layouts can be structured according to the target structured template to obtain text data conforming to the target structured template, and displayed to the user for viewing.
[0112] In the above embodiment of the present application, obtaining the target structured template includes one of the following: receiving the target structured template; receiving a structured template selected from multiple output structured templates to obtain the target structured template.
[0113] In an optional embodiment, the user may upload a structured template, so that the cloud server may generate final text data according to the structured template uploaded by the user.
[0114] In another optional embodiment, for some users who may not understand the specific format of the structured template, the cloud server can provide users with a variety of different structured templates, such as Figure 3 In the interactive interface shown, the user makes selections based on his or her own habits, and the final text data is generated according to the structured template selected by the user.
[0115] The following combination Figure 5 Taking ancient books as an example, a preferred embodiment of the present application is described in detail. The method can be executed by a mobile terminal or a server. In the embodiment of the present application, the method is described by taking the server as an example. Figure 5 As shown in the figure, this method can be divided into two parts: model training and application reasoning, which can specifically include the following steps:
[0116] Step S51, obtaining an ancient book image;
[0117] Step S52: Marking the ancient book image with a layout frame and classification according to the defined ancient book layout classification;
[0118] Step S53, using the annotated ancient book images to train a deep learning detection model;
[0119] Optionally, the deep learning detection module used in this application may be CenterNet, which may be implemented through existing processing methods.
[0120] The above steps S51 to S53 are the model training part.
[0121] Step S54, obtaining the image of the ancient book to be recorded;
[0122] Step S55: input the ancient book image into the deep learning detection network to obtain the predicted layout frame and the corresponding layout classification;
[0123] Step S56, recognizing the ancient book image through OCR to extract the position and content of all the characters in the ancient book image;
[0124] In step S57, the layout box and the text recognition results are matched in area to obtain the corresponding text inside the layout box and form a text block. The text blocks of the layout boxes in all areas are structurally merged to output the final electronic result.
[0125] Optionally, the purpose of area matching can be achieved by calculating the cross area ratio.
[0126] The above steps S54 to S57 are the application reasoning part.
[0127] Through the above steps, we can define the detection and classification of ancient book layouts based on the data characteristics of ancient books, and structure the OCR results of ancient books. This has better detection and classification effects and is more suitable for downstream use in ancient book scenarios.
[0128] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0129] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0130] Example 2
[0131] According to an embodiment of the present application, a text processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0132] Figure 6 Flowchart of the second text processing method according to the embodiment of the present application. Figure 6 As shown, the method includes the following steps:
[0133] Step S602: display a first text image.
[0134] In an optional embodiment, the Figure 4 The interactive interface shown displays the first text image selected by the user.
[0135] Step S604: marking a first recognition result of a plurality of characters in the first text image in the first text image, wherein the first recognition result is obtained by recognizing the first text image, and the first recognition result includes: content information corresponding to each character, and character position information of each character.
[0136] In an optional embodiment, each character may be framed in the first text image based on the character position information of each character in the first text image, and the corresponding content information may be displayed, thereby achieving the purpose of marking the first recognition result in the first text image.
[0137] Step S606, marking a second recognition result of at least one target area in the first text image in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area.
[0138] In an optional embodiment, each target area can be framed in the first text image based on the area position information of each target area, and the corresponding layout type can be displayed, thereby achieving the purpose of marking the second recognition result in the first text image.
[0139] Step S608 : Displaying text data corresponding to the first text image, wherein the text data is generated based on the first recognition result and the second recognition result.
[0140] In an optional embodiment, the Figure 4 The generated electronic result, ie the above-mentioned text data, is displayed on the interactive interface shown.
[0141] In the above embodiment of the present application, before marking the second recognition result of at least one target area in the first text image in the first text image, the method further includes: detecting the first text image using a target detection model to obtain a second recognition result.
[0142] In the above embodiment of the present application, the first text image is detected using a target detection model to obtain a second recognition result, which includes: inputting the first text image into the target detection model for prediction to obtain a prediction result for each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; and determining the second recognition result based on the prediction result of each pixel.
[0143] The prediction result in the above step can include a classification result and an offset result, the classification result being used to represent whether each pixel is located at the center of any one target region, and the offset result being used to represent an offset of each pixel from a vertex of any one target region in a horizontal direction and a vertical direction.
[0144] Optionally, determining the second recognition result based on the prediction result of each pixel can include: determining at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target region respectively; and performing regression operation based on the offset result of the at least one target pixel to obtain the second recognition result.
[0145] In the above embodiments, before the target detection model is used to detect the first text image to obtain the second recognition result, the method further includes: obtaining a plurality of second text images; labeling each second text image according to a preset layout classification rule to obtain a labeling result of each second text image, wherein the labeling result includes position information of a preset region of each second text image and a type of a layout corresponding to the preset region; generating a training sample based on the plurality of second text images and the labeling result of each second text image; and training the target detection model by using the training sample.
[0146] In the above embodiments, before the text data corresponding to the first text image is displayed, the method further includes: matching each character and each target region based on the character position information of each character and the region position information of each target region to determine target characters contained in each target region; merging content information corresponding to the target characters contained in each target region to generate text data corresponding to each target region; and combining the text data corresponding to the at least one target region to generate the text data corresponding to the first text image.
[0147] In the above embodiments, matching each character and each target region based on the character position information of each character and the region position information of each target region to determine target characters contained in each target region includes: determining a cross-area ratio of each character and each target region based on the character position information of each character and the region position information of each target region; for each character, determining a target region corresponding to a target cross-area ratio as a target region to which the character belongs, wherein the target cross-area ratio is greater than a preset threshold, and the target cross-area ratio is a maximum cross-area ratio; and determining target characters contained in each target region based on the target region to which each character belongs.
[0148] In the above embodiment of the present application, determining the intersection area ratio of each character and each target area based on the character position information of each character and the area position information of each target area includes: determining the intersection area of each character and each target area based on the character position information of each character and the area position information of each target area; obtaining the ratio of the area of the intersection area to the area of each character to obtain the intersection area ratio, wherein the area of each character is determined based on the character position information of each character.
[0149] In the above embodiment of the present application, the content information corresponding to the target text contained in each target area is merged to generate text data corresponding to each target area, including: determining the sorting relationship between the target texts based on the position information of the target texts; and generating text data corresponding to each target area based on the sorting relationship between the target texts and the content information corresponding to the target texts.
[0150] In the above embodiment of the present application, after generating text data corresponding to each target area, the method also includes: outputting the text data corresponding to each target area; receiving response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; combining the text data and / or response data corresponding to at least one target area to generate text data corresponding to the first text image.
[0151] In the above embodiment of the present application, combining the text data corresponding to at least one target area to generate text data corresponding to the first text image includes: combining the type of each target area with the text data corresponding to each target area to generate structured data of each target area; combining the structured data of at least one target area to obtain text data corresponding to the first text image.
[0152] In the above embodiment of the present application, combining the structured data of at least one target area to obtain text data corresponding to the first text image includes: obtaining a target structured template; combining the structured data of at least one target area according to the target structured template to generate text data corresponding to the first text image.
[0153] In the above embodiment of the present application, obtaining the target structured template includes one of the following: receiving the target structured template; receiving a structured template selected from multiple output structured templates to obtain the target structured template.
[0154] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0155] Example 3
[0156] According to an embodiment of the present application, a text processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0157] Figure 7 This is a flowchart of the third text processing method according to an embodiment of the present application. Figure 7 As shown, the method includes the following steps:
[0158] Step S702: Receive a first text image.
[0159] In an optional embodiment, the user may capture a text image using a camera and transmit the image to the cloud server via the user's computer terminal, so that the cloud server may receive the text image to be processed.
[0160] Step S704 : Recognize the first text image to obtain a first recognition result of the plurality of characters in the first text image, wherein the first recognition result includes: content information corresponding to each character and character position information of each character.
[0161] Step S706 : Detect the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and the type of layout corresponding to each target area.
[0162] Step S708 : Generate text data corresponding to the first text image based on the first recognition result and the second recognition result.
[0163] Step S710: output text data corresponding to the first text image.
[0164] In an optional embodiment, the cloud server can transmit the generated electronic result, that is, the above text data, to the computer terminal, which displays it on the computer terminal. Figure 4 In the interactive interface shown, the user can view it. In addition, the text data can be transmitted to a computer terminal for storage, or directly stored in the cloud server.
[0165] In the above embodiment of the present application, detecting the first text image to obtain a second recognition result of at least one target area in the first text image includes: detecting the first text image using a target detection model to obtain a second recognition result.
[0166] In the above embodiment of the present application, the first text image is detected using a target detection model to obtain a second recognition result, which includes: inputting the first text image into the target detection model for prediction to obtain a prediction result for each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; and determining the second recognition result based on the prediction result of each pixel.
[0167] The prediction results in the above steps may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0168] Optionally, determining the second recognition result based on the prediction result of each pixel may include: determining at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; performing a regression operation based on the offset result of the at least one target pixel to obtain the second recognition result.
[0169] In the above embodiment of the present application, before using the target detection model to detect the first text image and obtain the second recognition result, the method also includes: obtaining multiple second text images; annotating each second text image according to the preset layout classification rules to obtain the annotation results of each second text image, wherein the annotation results include: the location information of the preset area of each second text image, and the type of layout corresponding to the preset area; generating training samples based on the multiple second text images and the annotation results of each second text image; and using the training samples to train the target detection model.
[0170] In the above embodiment of the present application, based on the first recognition result and the second recognition result, generating text data corresponding to the first text image includes: matching each character with each target area based on the character position information of each character and the area position information of each target area, and determining the target character contained in each target area; merging the content information corresponding to the target character contained in each target area to generate text data corresponding to each target area; combining the text data corresponding to at least one target area to generate text data corresponding to the first text image.
[0171] In the above embodiments of the present application, the matching of each character and each target region based on the character position information of each character and the region position information of each target region to determine the target character contained in each target region comprises: determining a cross-area ratio of each character and each target region based on the character position information of each character and the region position information of each target region; determining, for each character, a target region corresponding to a target cross-area ratio as a target region to which each character belongs, wherein the target cross-area ratio is greater than a preset threshold, and the target cross-area ratio is a maximum cross-area ratio; and determining the target character contained in each target region based on the target region to which each character belongs.
[0172] In the above embodiments of the present application, the determination of the cross-area ratio of each character and each target region based on the character position information of each character and the region position information of each target region comprises: determining a cross region of each character and each target region based on the character position information of each character and the region position information of each target region; and obtaining a ratio of an area of the cross region to an area of each character to obtain the cross-area ratio, wherein the area of each character is determined based on the character position information of each character.
[0173] In the above embodiments of the present application, the merging of the content information corresponding to the target character contained in each target region to generate the text data corresponding to each target region comprises: determining an ordering relationship between the target characters based on the position information of the target characters; and generating the text data corresponding to each target region based on the ordering relationship between the target characters and the content information corresponding to the target characters.
[0174] In the above embodiments of the present application, after the generation of the text data corresponding to each target region, the method further comprises: outputting the text data corresponding to each target region; receiving response data corresponding to each target region, wherein the response data is obtained by modifying the text data corresponding to each target region; and combining the text data and / or the response data corresponding to at least one target region to generate the text data corresponding to the first text image.
[0175] In the above embodiments of the present application, the combination of the text data corresponding to at least one target region to generate the text data corresponding to the first text image comprises: combining the type of each target region and the text data corresponding to each target region to generate structured data of each target region; and combining the structured data of at least one target region to obtain the text data corresponding to the first text image.
[0176] In the above embodiment of the present application, combining the structured data of at least one target area to obtain text data corresponding to the first text image includes: obtaining a target structured template; combining the structured data of at least one target area according to the target structured template to generate text data corresponding to the first text image.
[0177] In the above embodiment of the present application, obtaining the target structured template includes one of the following: receiving the target structured template; receiving a structured template selected from multiple output structured templates to obtain the target structured template.
[0178] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0179] Example 4
[0180] According to an embodiment of the present application, a text processing device for implementing the above text processing method is also provided. Figure 8 As shown, the apparatus 800 includes: an acquisition module 802 , an identification module 804 , a detection module 806 and a generation module 808 .
[0181] Among them, the acquisition module 802 is used to acquire the first text image; the recognition module 804 is used to recognize the first text image and obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and text position information of each character; the detection module 806 is used to detect the first text image and obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; the generation module 808 is used to generate text data corresponding to the first text image based on the first recognition result and the second recognition result.
[0182] It should be noted that the acquisition module 802, identification module 804, detection module 806, and generation module 808 correspond to steps S202 to S208 in Example 1. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0183] In the above embodiment of the present application, the detection module is further used to detect the first text image using the target detection model to obtain a second recognition result.
[0184] In the above embodiments of the present application, the detection module includes: an input unit and a determination unit.
[0185] Among them, the input unit is used to input the first text image into the target detection model for prediction to obtain the prediction result of each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; the determination unit is used to determine the second recognition result based on the prediction result of each pixel.
[0186] The above prediction results may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0187] Optionally, the determination unit is also used to determine at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; and perform a regression operation based on the offset result of the at least one target pixel to obtain a second recognition result.
[0188] In the above embodiments of the present application, the device further includes: a labeling module, a generation module and a training module.
[0189] Among them, the acquisition module is also used to acquire multiple second text images; the annotation module is used to annotate each second text image according to the preset layout classification rules to obtain the annotation results of each second text image, wherein the annotation results include: the location information of the preset area of each second text image, and the type of layout corresponding to the preset area; the generation module is used to generate training samples based on multiple second text images and the annotation results of each second text image; the training module is used to train the target detection model using the training samples.
[0190] In the above embodiments of the present application, the generation module includes: a matching unit, a merging unit and a combining unit.
[0191] Among them, the matching unit is used to match each character with each target area based on the character position information of each character and the area position information of each target area, and determine the target character contained in each target area; the merging unit is used to merge the content information corresponding to the target character contained in each target area, and generate text data corresponding to each target area; the combining unit is used to combine the text data corresponding to at least one target area, and generate text data corresponding to the first text image.
[0192] In the above embodiment of the present application, the matching unit includes: a first determining subunit, a second determining subunit and a third determining subunit.
[0193] Among them, the first determination subunit is used to determine the intersection area ratio of each character and each target area based on the character position information of each character and the area position information of each target area; the second determination subunit is used to determine, for each character, that the target area corresponding to the target intersection area ratio is the target area to which each character belongs, wherein the target intersection area ratio is greater than a preset threshold and the target intersection area ratio is the maximum intersection area ratio; the third determination subunit is used to determine the target characters contained in each target area based on the target area to which each character belongs.
[0194] In the above embodiment of the present application, the first determination subunit is also used to determine the intersection area between each character and each target area based on the character position information of each character and the area position information of each target area; obtain the ratio of the area of the intersection area to the area of each character to obtain the intersection area ratio, wherein the area of each character is determined based on the character position information of each character.
[0195] In the above embodiment of the present application, the merging unit includes: a fourth determining subunit and a generating subunit.
[0196] The fourth determining subunit is used to determine the sorting relationship between target characters based on the position information of the target characters; the generating subunit is used to generate text data corresponding to each target area based on the sorting relationship between the target characters and the content information corresponding to the target characters.
[0197] In the above embodiments of the present application, the device further includes: an output module and a receiving module.
[0198] Among them, the output module is used to output the text data corresponding to each target area; the receiving module is used to receive the response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; the combination unit is also used to combine the text data and / or response data corresponding to at least one target area to generate text data corresponding to the first text image.
[0199] In the above embodiments of the present application, the combination unit includes: a first combination sub-unit and a second combination sub-unit.
[0200] Among them, the first combination subunit is used to combine the type of each target area with the text data corresponding to each target area to generate structured data of each target area; the second combination subunit is used to combine the structured data of at least one target area to obtain text data corresponding to the first text image.
[0201] In the above embodiment of the present application, the second combining subunit is further used to obtain a target structured template; and according to the target structured template, the structured data of at least one target area are combined to generate text data corresponding to the first text image.
[0202] In the above embodiment of the present application, the second combination subunit is further used to receive a target structured template; or, receive a structured template selected from the output multiple structured templates to obtain the target structured template.
[0203] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0204] Example 5
[0205] According to an embodiment of the present application, a text processing device for implementing the above text processing method is also provided. Figure 9 As shown, the apparatus 900 includes: a first display module 902 , a first marking module 904 , a second marking module 906 and a second display module 908 .
[0206] Among them, the first display module 902 is used to display the first text image; the first marking module 904 is used to mark the first recognition result of multiple characters in the first text image in the first text image, wherein the first recognition result is obtained by recognizing the first text image, and the first recognition result includes: content information corresponding to each character, and text position information of each character; the second marking module 906 is used to mark the second recognition result of at least one target area in the first text image in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; the second display module 908 is used to display text data corresponding to the first text image, wherein the text data is generated based on the first recognition result and the second recognition result.
[0207] It should be noted that the first display module 902, the first marking module 904, the second marking module 906, and the second display module 908 correspond to steps S602 to S608 in Example 2. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0208] In the above embodiments of the present application, the device further includes: a detection module.
[0209] The detection module is used to detect the first text image using the target detection model to obtain a second recognition result.
[0210] In the above embodiments of the present application, the detection module includes: an input unit and a determination unit.
[0211] Among them, the input unit is used to input the first text image into the target detection model for prediction to obtain the prediction result of each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; the determination unit is used to determine the second recognition result based on the prediction result of each pixel.
[0212] The above prediction results may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0213] Optionally, the determination unit includes: a determination subunit and an operation subunit.
[0214] Among them, the determination subunit is used to determine at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; the operation subunit is used to perform a regression operation based on the offset result of at least one target pixel to obtain a second recognition result.
[0215] In the above embodiments of the present application, the device further includes: a labeling module, a generation module and a training module.
[0216] Among them, the acquisition module is also used to acquire multiple second text images; the annotation module is used to annotate each second text image according to the preset layout classification rules to obtain the annotation results of each second text image, wherein the annotation results include: the location information of the preset area of each second text image, and the type of layout corresponding to the preset area; the generation module is used to generate training samples based on multiple second text images and the annotation results of each second text image; the training module is used to train the target detection model using the training samples.
[0217] In the above embodiments of the present application, the device further includes: a matching module, a merging module and a combining module.
[0218] Among them, the matching module is used to match each text and each target area based on the text position information of each text and the area position information of each target area, and determine the target text contained in each target area; the merging module is used to merge the content information corresponding to the target text contained in each target area, and generate text data corresponding to each target area; the combining module is used to combine the text data corresponding to at least one target area, and generate text data corresponding to the first text image.
[0219] In the above embodiment of the present application, the matching module includes: a first determination unit, a second determination unit and a third determination unit.
[0220] The first determining unit is configured to determine, based on the character position information of each character and the region position information of each target region, a cross-area ratio of each character and each target region. The second determining unit is configured to determine, for each character, a target region corresponding to a target cross-area ratio as a target region to which the character belongs, where the target cross-area ratio is greater than a preset threshold and the target cross-area ratio is a maximum cross-area ratio. The third determining unit is configured to determine, based on the target region to which each character belongs, a target character contained in each target region.
[0221] In the above embodiments of the present application, the first determining unit is further configured to determine, based on the character position information of each character and the region position information of each target region, a cross region of each character and each target region; and obtain a ratio of an area of the cross region to an area of each character to obtain the cross-area ratio, where the area of each character is determined based on the character position information of each character.
[0222] In the above embodiments of the present application, the merging module includes a fourth determining unit and a generating unit.
[0223] The fourth determining unit is configured to determine, based on the position information of the target characters, an ordering relationship between the target characters. The generating unit is configured to generate, based on the ordering relationship between the target characters and the content information corresponding to the target characters, text data corresponding to each target region.
[0224] In the above embodiments of the present application, the device further includes a third display module and a receiving module.
[0225] The third display module is configured to display the text data corresponding to each target region. The receiving module is configured to receive response data corresponding to each target region, where the response data is obtained by modifying the text data corresponding to each target region. The combining module is further configured to combine the text data and / or the response data corresponding to at least one target region to generate text data corresponding to the first text image.
[0226] In the above embodiments of the present application, the combining module includes a first combining unit and a second combining unit.
[0227] The first combining unit is configured to combine the type of each target region and the text data corresponding to each target region to generate structured data of each target region. The second combining unit is configured to combine the structured data of at least one target region to obtain the text data corresponding to the first text image.
[0228] In the above embodiments of the present application, the second combining unit includes an obtaining subunit and a combining subunit.
[0229] The acquisition subunit is used to acquire the target structured template; the combination subunit is used to combine the structured data of at least one target area according to the target structured template to generate text data corresponding to the first text image.
[0230] In the above embodiment of the present application, the acquisition subunit is further used to receive a target structured template; or, receive a structured template selected from multiple output structured templates to obtain the target structured template.
[0231] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0232] Example 6
[0233] According to an embodiment of the present application, a text processing device for implementing the above text processing method is also provided. Figure 10 As shown, the apparatus 1000 includes: a receiving module 1002 , an identification module 1004 , a detection module 1006 , a generation module 1008 and an output module 1010 .
[0234] Among them, the receiving module 1002 is used to receive a first text image; the recognition module 1004 is used to recognize the first text image and obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and text position information of each character; the detection module 1006 is used to detect the first text image and obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; the generation module 1008 is used to generate text data corresponding to the first text image based on the first recognition result and the second recognition result; the output module 1010 is used to output the text data corresponding to the first text image.
[0235] It should be noted that the receiving module 1002, identification module 1004, detection module 1006, generation module 1008, and output module 1010 described above correspond to steps S702 to S710 in Example 3. The examples and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to those disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0236] In the above embodiment of the present application, the detection module is further used to detect the first text image using the target detection model to obtain a second recognition result.
[0237] In the above embodiments of the present application, the detection module includes: an input unit and a determination unit.
[0238] Among them, the input unit is used to input the first text image into the target detection model for prediction to obtain the prediction result of each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; the determination unit is used to determine the second recognition result based on the prediction result of each pixel.
[0239] The above prediction results may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0240] Optionally, the determination unit is also used to determine at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; and perform a regression operation based on the offset result of the at least one target pixel to obtain a second recognition result.
[0241] In the above embodiments of the present application, the device further includes: a labeling module, a generation module and a training module.
[0242] Among them, the acquisition module is also used to acquire multiple second text images; the annotation module is used to annotate each second text image according to the preset layout classification rules to obtain the annotation results of each second text image, wherein the annotation results include: the location information of the preset area of each second text image, and the type of layout corresponding to the preset area; the generation module is used to generate training samples based on multiple second text images and the annotation results of each second text image; the training module is used to train the target detection model using the training samples.
[0243] In the above embodiments of the present application, the generation module includes: a matching unit, a merging unit and a combining unit.
[0244] Among them, the matching unit is used to match each character with each target area based on the character position information of each character and the area position information of each target area, and determine the target character contained in each target area; the merging unit is used to merge the target characters contained in each target area based on the first recognition result of the target character, and generate text data corresponding to each target area; the combining unit is used to combine the text data corresponding to at least one target area to generate text data corresponding to the first text image.
[0245] In the above embodiment of the present application, the matching unit includes: a first determining subunit, a second determining subunit and a third determining subunit.
[0246] The first determining sub-unit is configured to determine, based on the character position information of each character and the region position information of each target region, a cross-area ratio of each character and each target region; the second determining sub-unit is configured to determine, for each character, a target region corresponding to a target cross-area ratio as a target region to which each character belongs, where the target cross-area ratio is greater than a preset threshold and the target cross-area ratio is a maximum cross-area ratio; and the third determining sub-unit is configured to determine, based on the target region to which each character belongs, a target character contained in each target region.
[0247] In the above embodiments of the present application, the first determining sub-unit is further configured to determine, based on the character position information of each character and the region position information of each target region, a cross-area of each character and each target region; and obtain a ratio of an area of the cross-area to an area of each character to obtain the cross-area ratio, where the area of each character is determined based on the character position information of each character.
[0248] In the above embodiments of the present application, the merging unit comprises a fourth determining sub-unit and a generating sub-unit.
[0249] The fourth determining sub-unit is configured to determine, based on the position information of the target characters, an ordering relationship between the target characters; and the generating sub-unit is configured to generate, based on the ordering relationship between the target characters and the content information corresponding to the target characters, text data corresponding to each target region.
[0250] In the above embodiments of the present application, the device further comprises a receiving module.
[0251] The output module is further configured to output the text data corresponding to each target region; the receiving module is configured to receive response data corresponding to each target region, where the response data is obtained by modifying the text data corresponding to each target region; and the combining unit is further configured to combine the text data and / or the response data corresponding to at least one target region to generate the text data corresponding to the first text image.
[0252] In the above embodiments of the present application, the combining unit comprises a first combining sub-unit and a second combining sub-unit.
[0253] The first combining sub-unit is configured to combine the type of each target region and the text data corresponding to each target region to generate structured data of each target region; and the second combining sub-unit is configured to combine the structured data of at least one target region to obtain the text data corresponding to the first text image.
[0254] In the above embodiments of the present application, the second combining sub-unit is further configured to obtain a target structured template; and combine the structured data of at least one target region according to the target structured template to generate the text data corresponding to the first text image.
[0255] In the above embodiment of the present application, the second combination subunit is further used to receive a target structured template; or, receive a structured template selected from the output multiple structured templates to obtain the target structured template.
[0256] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0257] Example 7
[0258] According to an embodiment of the present application, a text processing system is further provided, including:
[0259] processor; and
[0260] The memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining a first text image; identifying the first text image to obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and character position information of each character; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; generating text data corresponding to the first text image based on the first recognition result and the second recognition result.
[0261] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0262] Example 8
[0263] According to an embodiment of the present application, a text processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0264] Figure 11 Flowchart of the fourth text processing method according to the embodiment of the present application. Figure 11 As shown, the method may include the following steps:
[0265] Step S1102: Acquire an ancient book image.
[0266] The ancient book image in the above step can be an image of each page of the ancient book, which can be directly obtained by shooting the pages of the ancient book, or can be obtained by intercepting video frames. The video is a video shot during the reading of the entire ancient book, and contains the content of each page.
[0267] The ancient book image in the above step can include at least one target region, and the plurality of texts included in the ancient book image are typeset according to a first typesetting manner. The target region can be a region of different formats in each page of the ancient book. The types of formats in different pages are different, for example, the format of the first page of the ancient book is usually a cover, and the format of the second page is usually a table of contents.
[0268] The ancient book image includes all the texts in the page, and the texts in the page are typeset according to the typesetting manner of the ancient book. Therefore, the first typesetting manner described above can refer to the typesetting manner of the ancient book, which can be "from top to bottom, from right to left". However, the ancient book image not only includes texts in different formats, but also includes watermarks, labels added during post-processing, and other texts irrelevant to the ancient book itself.
[0269] In step S1104, the ancient book image is recognized to obtain a first recognition result of a plurality of texts in the ancient book image, wherein the first recognition result includes content information corresponding to each text and text position information of each text.
[0270] In step S1106, the ancient book image is detected to obtain a second recognition result of at least one target region in the ancient book image, wherein the second recognition result includes region position information of each target region and a type of format corresponding to each target region.
[0271] Because the typesetting manner of the ancient book is different from the typesetting manner of modern books, newspapers and periodicals, etc., the existing layout analysis algorithm cannot be directly used to detect and mark "header, footer, chart, table, title, body" and other elements. In order to achieve the purpose of detecting and classifying the format of the ancient book, a detection algorithm can be designed according to the format characteristics of the ancient book, and the format of the ancient book can be detected and classified by the detection algorithm to determine the specific position (i.e. format frame) and the specific type of the format.
[0272] In the embodiment of the present application, the specific position of each target region can be represented by four vertices of a quadrilateral as coordinates, and can be arranged in the order of top left, top right, bottom right, and bottom left, for example, the region position information in the above step: (1, 1), (100, 100), (100, 100), (1, 100). An example of the second recognition result is: [(1, 1), (100, 100), (100, 100), (1, 100), type: body].
[0273] Step S1108 : generating an ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result.
[0274] The ancient book text in the above step may refer to a serialized text obtained by sorting the content information corresponding to all characters according to the sorting relationship and layout type.
[0275] In the above embodiment of the present application, detecting the ancient book image to obtain a second recognition result of at least one target area in the ancient book image includes: detecting the ancient book image using a target detection model to obtain a second recognition result.
[0276] The target detection model in the above steps can be CenterNet, FastRCNN, MaskRCNN, FCN, etc. In the embodiment of the present application, the CenterNet model is taken as an example for illustration.
[0277] In the above embodiment of the present application, the ancient book image is detected using a target detection model, and obtaining the second recognition result includes: inputting the ancient book image into the target detection model for prediction, and obtaining the prediction result of each pixel in the ancient book image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; and determining the second recognition result based on the prediction result of each pixel.
[0278] The prediction results in the above steps may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0279] Optionally, determining the second recognition result based on the prediction result of each pixel may include: determining at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; performing a regression operation based on the offset result of the at least one target pixel to obtain the second recognition result.
[0280] In the above embodiment of the present application, before using the target detection model to detect the ancient book image and obtain the second recognition result, the method also includes: obtaining multiple ancient book images; annotating each ancient book image according to the preset layout classification rules to obtain the annotation results of each ancient book image, wherein the annotation results include: the location information of the preset area of each ancient book image, and the type of layout corresponding to the preset area; generating training samples based on the multiple ancient book images and the annotation results of each ancient book image; and using the training samples to train the target detection model.
[0281] In the above embodiments of the present application, generating the ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result comprises: matching each character with each target region based on the character position information of each character and the region position information of each target region, to determine target characters contained in each target region; merging content information corresponding to the target characters contained in each target region to generate text data corresponding to each target region; and combining the text data corresponding to at least one target region to generate the ancient book text corresponding to the ancient book image.
[0282] In the above embodiments of the present application, matching each character with each target region based on the character position information of each character and the region position information of each target region to determine target characters contained in each target region comprises: determining a cross-area ratio of each character and each target region based on the character position information of each character and the region position information of each target region; and determining, for each character, a target region corresponding to a target cross-area ratio as a target region to which the character belongs, wherein the target cross-area ratio is greater than a preset threshold value, and the target cross-area ratio is a maximum cross-area ratio; and determining target characters contained in each target region based on the target region to which each character belongs.
[0283] The preset threshold value in the above step can be a threshold value for determining the matching of each character and each target region, and can be a fixed value preset in advance, for example, 0.7, but is not limited thereto, and can be adjusted and determined according to actual needs.
[0284] In the above embodiments of the present application, determining a cross-area ratio of each character and each target region based on the character position information of each character and the region position information of each target region comprises: determining a cross region of each character and each target region based on the character position information of each character and the region position information of each target region; and obtaining a ratio of an area of the cross region to an area of each character to obtain the cross-area ratio, wherein the area of each character is determined based on the character position information of each character.
[0285] In the above embodiments of the present application, merging content information corresponding to the target characters contained in each target region to generate text data corresponding to each target region comprises: determining an ordering relationship between the target characters based on the position information of the target characters; and generating the text data corresponding to each target region based on the ordering relationship between the target characters and the content information corresponding to the target characters.
[0286] The ordering relationship in the above step can refer to the order of different characters in the reading process, for example, assuming that the ordering relationship of character a and character b is that character a is arranged in front of character b, then in the reading process, the user reads character a first and then reads character b.
[0287] In the above embodiment of the present application, after generating text data corresponding to each target area, the method also includes: outputting the text data corresponding to each target area; receiving response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; combining the text data and / or response data corresponding to at least one target area to generate an ancient book text corresponding to the ancient book image.
[0288] In the above embodiment of the present application, combining the text data corresponding to at least one target area to generate the ancient book text corresponding to the ancient book image includes: combining the type of each target area with the text data corresponding to each target area to generate structured data of each target area; combining the structured data of at least one target area to obtain the ancient book text corresponding to the ancient book image.
[0289] In the above embodiment of the present application, combining the structured data of at least one target area to obtain the ancient book text corresponding to the ancient book image includes: obtaining a target structured template; combining the structured data of at least one target area according to the target structured template to generate the ancient book text corresponding to the ancient book image.
[0290] In the above embodiment of the present application, obtaining the target structured template includes one of the following: receiving the target structured template; receiving a structured template selected from multiple output structured templates to obtain the target structured template.
[0291] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0292] Example 9
[0293] According to an embodiment of the present application, a text processing device for implementing the above text processing method is also provided. Figure 12 As shown, the device 1200 includes: an acquisition module 1202 , an identification module 1204 , a detection module 1206 and a generation module 1208 .
[0294] Among them, the acquisition module 1202 is used to acquire the ancient book image; the recognition module 1204 is used to recognize the ancient book image and obtain a first recognition result of multiple characters in the ancient book image, wherein the first recognition result includes: content information corresponding to each character, and text position information of each character; the detection module 1206 is used to detect the ancient book image and obtain a second recognition result of at least one target area in the ancient book image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; the generation module 808 is used to generate the ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result.
[0295] It should be noted that the acquisition module 1202, identification module 1204, detection module 1206, and generation module 1208 correspond to steps S1102 to S1108 in Example 8. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0296] In the above embodiment of the present application, the detection module is also used to detect the ancient book image using the target detection model to obtain a second recognition result.
[0297] In the above embodiments of the present application, the detection module includes: an input unit and a determination unit.
[0298] Among them, the input unit is used to input the ancient book image into the target detection model for prediction to obtain the prediction result of each pixel in the ancient book image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; the determination unit is used to determine the second recognition result based on the prediction result of each pixel.
[0299] The above prediction results may include: classification results and offset results. The classification results are used to characterize whether each pixel is located at the center of any target area, and the offset results are used to characterize the offset of each pixel from the vertex of any target area in the horizontal and vertical directions.
[0300] Optionally, the determination unit is also used to determine at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; and perform a regression operation based on the offset result of the at least one target pixel to obtain a second recognition result.
[0301] In the above embodiments of the present application, the device further includes: a labeling module, a generation module and a training module.
[0302] Among them, the acquisition module is also used to obtain multiple ancient book images; the annotation module is used to annotate each ancient book image according to the preset layout classification rules to obtain the annotation results of each ancient book image, wherein the annotation results include: the location information of the preset area of each ancient book image, and the type of layout corresponding to the preset area; the generation module is used to generate training samples based on multiple ancient book images and the annotation results of each ancient book image; the training module is used to use the training samples to train the target detection model.
[0303] In the above embodiments of the present application, the generation module includes: a matching unit, a merging unit and a combining unit.
[0304] Among them, the matching unit is used to match each character with each target area based on the character position information of each character and the area position information of each target area, and determine the target character contained in each target area; the merging unit is used to merge the content information corresponding to the target character contained in each target area to generate text data corresponding to each target area; the combining unit is used to combine the text data corresponding to at least one target area to generate an ancient book text corresponding to the ancient book image.
[0305] In the above embodiment of the present application, the matching unit includes: a first determining subunit, a second determining subunit and a third determining subunit.
[0306] Among them, the first determination subunit is used to determine the intersection area ratio of each character and each target area based on the character position information of each character and the area position information of each target area; the second determination subunit is used to determine, for each character, that the target area corresponding to the target intersection area ratio is the target area to which each character belongs, wherein the target intersection area ratio is greater than a preset threshold and the target intersection area ratio is the maximum intersection area ratio; the third determination subunit is used to determine the target characters contained in each target area based on the target area to which each character belongs.
[0307] In the above embodiment of the present application, the first determination subunit is also used to determine the intersection area between each character and each target area based on the character position information of each character and the area position information of each target area; obtain the ratio of the area of the intersection area to the area of each character to obtain the intersection area ratio, wherein the area of each character is determined based on the character position information of each character.
[0308] In the above embodiment of the present application, the merging unit includes: a fourth determining subunit and a generating subunit.
[0309] The fourth determining subunit is used to determine the sorting relationship between target characters based on the position information of the target characters; the generating subunit is used to generate text data corresponding to each target area based on the sorting relationship between the target characters and the content information corresponding to the target characters.
[0310] In the above embodiments of the present application, the device further includes: an output module and a receiving module.
[0311] Among them, the output module is used to output the text data corresponding to each target area; the receiving module is used to receive the response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; the combination unit is also used to combine the text data and / or response data corresponding to at least one target area to generate an ancient book text corresponding to the ancient book image.
[0312] In the above embodiments of the present application, the combination unit includes: a first combination sub-unit and a second combination sub-unit.
[0313] Among them, the first combination subunit is used to combine the type of each target area with the text data corresponding to each target area to generate structured data of each target area; the second combination subunit is used to combine the structured data of at least one target area to obtain the ancient book text corresponding to the ancient book image.
[0314] In the above embodiment of the present application, the second combining subunit is further used to obtain a target structured template; according to the target structured template, the structured data of at least one target area are combined to generate an ancient book text corresponding to the ancient book image.
[0315] In the above embodiment of the present application, the second combination subunit is further used to receive a target structured template; or, receive a structured template selected from the output multiple structured templates to obtain the target structured template.
[0316] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0317] Example 10
[0318] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0319] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0320] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the text processing method: obtaining a first text image; identifying the first text image to obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and text position information of each character; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; based on the first recognition result and the second recognition result, generating text data corresponding to the first text image.
[0321] Optionally, Figure 13 This is a structural block diagram of a computer terminal according to an embodiment of the present application. Figure 13 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1302 and a memory 1304.
[0322] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the text processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned text processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0323] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a first text image; identify the first text image to obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and text position information of each character; detect the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; based on the first recognition result and the second recognition result, generate text data corresponding to the first text image.
[0324] Optionally, the processor may further execute program code of the following steps: detecting the first text image using a target detection model to obtain a second recognition result.
[0325] Optionally, the processor may also execute the program code of the following steps: inputting the first text image into the target detection model for prediction to obtain a prediction result for each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; and determining a second recognition result based on the prediction result of each pixel.
[0326] Optionally, the processor may also execute the program code of the following steps: determining at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; performing a regression operation based on the offset result of the at least one target pixel to obtain a second recognition result.
[0327] Optionally, the processor may also execute the program code of the following steps: obtaining multiple second text images; annotating each second text image according to preset layout classification rules to obtain an annotation result for each second text image, wherein the annotation result includes: location information of a preset area of each second text image, and the type of layout corresponding to the preset area; generating training samples based on multiple second text images and the annotation result of each second text image; and training the target detection model using the training samples.
[0328] Optionally, the processor may also execute the program code of the following steps: based on the text position information of each text and the area position information of each target area, match each text with each target area to determine the target text contained in each target area; merge the content information corresponding to the target text contained in each target area to generate text data corresponding to each target area; combine the text data corresponding to at least one target area to generate text data corresponding to the first text image.
[0329] Optionally, the processor may also execute the program code of the following steps: based on the text position information of each text and the area position information of each target area, determine the intersection area ratio of each text and each target area; for each text, determine that the target area corresponding to the target intersection area ratio is the target area to which each text belongs, wherein the target intersection area ratio is greater than a preset threshold, and the target intersection area ratio is the maximum intersection area ratio; based on the target area to which each text belongs, determine the target text contained in each target area.
[0330] Optionally, the processor may also execute the program code of the following steps: determining the intersection area between each character and each target area based on the character position information of each character and the area position information of each target area; obtaining the ratio of the area of the intersection area to the area of each character to obtain the intersection area ratio, wherein the area of each character is determined based on the character position information of each character.
[0331] Optionally, the processor may also execute program code for the following steps: determining the sorting relationship between target characters based on the position information of the target characters; and generating text data corresponding to each target area based on the sorting relationship between the target characters and the content information corresponding to the target characters.
[0332] Optionally, the processor may also execute program code for the following steps: outputting text data corresponding to each target area; receiving response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; and combining text data and / or response data corresponding to at least one target area to generate text data corresponding to the first text image.
[0333] Optionally, the processor may also execute the program code of the following steps: combining the type of each target area with the text data corresponding to each target area to generate structured data for each target area; combining the structured data of at least one target area to obtain text data corresponding to the first text image.
[0334] Optionally, the processor may further execute program code of the following steps: obtaining a target structured template; combining structured data of at least one target area according to the target structured template to generate text data corresponding to the first text image.
[0335] Optionally, the processor may further execute program code of the following steps: receiving a target structured template; or receiving a structured template selected from a plurality of output structured templates to obtain a target structured template.
[0336] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: displaying a first text image; marking a first recognition result of multiple characters in the first text image in the first text image, wherein the first recognition result is obtained by recognizing the first text image, and the first recognition result includes: content information corresponding to each character, and text position information of each character; marking a second recognition result of at least one target area in the first text image in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; displaying text data corresponding to the first text image, wherein the text data is generated based on the first recognition result and the second recognition result.
[0337] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receive a first text image; identify the first text image to obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and character position information of each character; detect the first text image to obtain a second recognition result of at least one target area of the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; generate text data corresponding to the first text image based on the first recognition result and the second recognition result; and output the text data corresponding to the first text image.
[0338] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain the ancient book image; identify the ancient book image to obtain a first recognition result of multiple characters in the ancient book image; detect the ancient book image to obtain a second recognition result of at least one target area in the ancient book image, wherein the second recognition result includes: the area location information of each target area, and the type of layout corresponding to each target area; based on the first recognition result and the second recognition result, generate the ancient book text corresponding to the ancient book image.
[0339] The embodiments of the present application provide a text image processing solution. After identifying the content and location information of each character, the text image's layout can be detected and classified in combination with the unique layout characteristics of the text image. The corresponding text data is generated based on the detection and classification results, enabling automated text entry without the need for manual entry of individual characters. This achieves the technical effect of improving layout detection and classification effectiveness and increasing entry efficiency, thereby resolving the technical problem in related technologies where text processing methods can only recognize individual characters, requiring manual entry of individual characters, resulting in low entry efficiency.
[0340] It can be understood by those skilled in the art that Figure 13 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 13 It does not limit the structure of the above electronic device. For example, the computer terminal A may also include Figure 13 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 13 Different configurations shown.
[0341] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0342] Example 11
[0343] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the text processing method provided in the above embodiment.
[0344] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0345] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: acquiring a first text image; identifying the first text image to obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and character position information of each character; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; generating text data corresponding to the first text image based on the first recognition result and the second recognition result.
[0346] Optionally, the storage medium is further configured to store program codes for executing the following steps: detecting the first text image using a target detection model to obtain a second recognition result.
[0347] Optionally, the above-mentioned storage medium is also configured to store program code for executing the following steps: inputting the first text image into the target detection model for prediction to obtain a prediction result for each pixel in the first text image, wherein the prediction result is used to characterize the positional relationship between each pixel and any target area; and determining a second recognition result based on the prediction result of each pixel.
[0348] Optionally, the processor may also execute the program code of the following steps: determining at least one target pixel based on the classification result of each pixel, wherein the at least one target pixel is located at the center of at least one target area; performing a regression operation based on the offset result of the at least one target pixel to obtain a second recognition result.
[0349] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a plurality of second text images; labeling each second text image according to a preset layout classification rule to obtain a labeling result of each second text image, wherein the labeling result comprises position information of a preset region of each second text image and a type of a layout corresponding to the preset region; generating a training sample based on the plurality of second text images and the labeling result of each second text image; and training the target detection model using the training sample.
[0350] Optionally, the storage medium is further configured to store program code for performing the following steps: matching each character and each target region based on the character position information of each character and the region position information of each target region to determine target characters contained in each target region; merging content information corresponding to the target characters contained in each target region to generate text data corresponding to each target region; and combining the text data corresponding to at least one target region to generate text data corresponding to the first text image.
[0351] Optionally, the storage medium is further configured to store program code for performing the following steps: determining a cross-area ratio of each character and each target region based on the character position information of each character and the region position information of each target region; for each character, determining a target region corresponding to a target cross-area ratio as a target region to which each character belongs, wherein the target cross-area ratio is greater than a preset threshold and the target cross-area ratio is a maximum cross-area ratio; and determining target characters contained in each target region based on the target region to which each character belongs.
[0352] Optionally, the storage medium is further configured to store program code for performing the following steps: determining a cross-area of each character and each target region based on the character position information of each character and the region position information of each target region; and obtaining a ratio of an area of the cross-area to an area of each character to obtain a cross-area ratio, wherein the area of each character is determined based on the character position information of each character.
[0353] Optionally, the storage medium is further configured to store program code for performing the following steps: determining an ordering relationship between target characters based on the position information of the target characters; and generating text data corresponding to each target region based on the ordering relationship between the target characters and content information corresponding to the target characters.
[0354] Optionally, the above-mentioned storage medium is also configured to store program code for executing the following steps: outputting text data corresponding to each target area; receiving response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; combining the text data and / or response data corresponding to at least one target area to generate text data corresponding to the first text image.
[0355] Optionally, the above-mentioned storage medium is also configured to store program code for executing the following steps: combining the type of each target area with the text data corresponding to each target area to generate structured data for each target area; combining the structured data of at least one target area to obtain text data corresponding to the first text image.
[0356] Optionally, the storage medium is further configured to store program codes for executing the following steps: obtaining a target structured template; combining structured data of at least one target area according to the target structured template to generate text data corresponding to the first text image.
[0357] Optionally, the storage medium is further configured to store program codes for executing the following steps: receiving a target structured template; or receiving a structured template selected from a plurality of output structured templates to obtain a target structured template.
[0358] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: displaying a first text image; marking a first recognition result of multiple characters in the first text image in the first text image, wherein the first recognition result is obtained by recognizing the first text image, and the first recognition result includes: content information corresponding to each character, and text position information of each character; marking a second recognition result of at least one target area in the first text image in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; displaying text data corresponding to the first text image, wherein the text data is generated based on the first recognition result and the second recognition result.
[0359] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving a first text image; identifying the first text image to obtain a first recognition result of multiple characters in the first text image, wherein the first recognition result includes: content information corresponding to each character, and character position information of each character; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area position information of each target area, and the type of layout corresponding to each target area; generating text data corresponding to the first text image based on the first recognition result and the second recognition result; and outputting the text data corresponding to the first text image.
[0360] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: acquiring an ancient book image; identifying the ancient book image to obtain a first recognition result of multiple characters in the ancient book image; detecting the ancient book image to obtain a second recognition result of at least one target area in the ancient book image, wherein the second recognition result includes: area location information of each target area, and the type of layout corresponding to each target area; generating an ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result.
[0361] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0362] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0363] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0364] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0365] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0366] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0367] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A text processing method, characterized in that: include: Acquire a first text image, wherein the first text image contains a plurality of characters arranged in a first typesetting manner, and the first text image contains at least one target area, and different target areas have different typesets; Recognize the first text image to obtain a first recognition result of a plurality of characters in the first text image; Detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and a type of layout corresponding to each target area; generating text data corresponding to the first text image based on the first recognition result and the second recognition result; Generating text data corresponding to the first text image based on the first recognition result and the second recognition result includes: Performing area matching on the first recognition result and the second recognition result to determine the target text contained in each target area; The target characters contained in the at least one target area are sorted according to a reading order between the target characters to generate text data corresponding to the first text image, wherein the text data is typeset in a second typesetting manner determined by a user's reading habits.
2. The method according to claim 1, characterized in that Detecting the first text image to obtain a second recognition result of at least one target area in the first text image includes: The first text image is detected using an object detection model to obtain the second recognition result.
3. The method according to claim 2, characterized in that Detecting the first text image using the object detection model to obtain the second recognition result includes: Inputting the first text image into the target detection model for prediction, obtaining a prediction result for each pixel in the first text image, wherein the prediction result is used to represent a positional relationship between each pixel and any target area; The second recognition result is determined based on the prediction result of each pixel.
4. The method according to claim 2, characterized in that Before detecting the first text image using the object detection model to obtain the second recognition result, the method further includes: obtaining a plurality of second text images; Annotating each second text image according to a preset layout classification rule to obtain an annotation result for each second text image, wherein the annotation result includes: position information of a preset area of each second text image, and a layout type corresponding to the preset area; generating a training sample based on the plurality of second text images and the annotation result of each second text image; The target detection model is trained using the training samples.
5. The method according to claim 1, wherein The first recognition result includes: content information corresponding to each character, and character position information of each character, wherein generating text data corresponding to the first text image based on the first recognition result and the second recognition result includes: Based on the character position information of each character and the area position information of each target area, matching each character with each target area to determine the target character contained in each target area; Merging the content information corresponding to the target text contained in each target area to generate text data corresponding to each target area; The text data corresponding to the at least one target area are combined to generate text data corresponding to the first text image.
6. The method according to claim 5, characterized in that Matching each character with each target area based on the character position information of each character and the area position information of each target area, and determining the target character contained in each target area includes: determining an intersection area ratio between each character and each target area based on the character position information of each character and the area position information of each target area; For each of the characters, determining a target area corresponding to a target cross-area ratio as the target area to which each character belongs, wherein the target cross-area ratio is greater than a preset threshold and the target cross-area ratio is a maximum cross-area ratio; Based on the target area to which each character belongs, the target characters contained in each target area are determined.
7. The method according to claim 6, characterized in that Determining, based on the character position information of each character and the area position information of each target area, an intersection area ratio between each character and each target area includes: Determining an intersection area between each character and each target area based on the character position information of each character and the area position information of each target area; The ratio of the area of the intersection region to the area of each character is obtained to obtain the intersection area ratio, wherein the area of each character is determined based on the character position information of each character.
8. The method according to claim 5, characterized in that Merging the content information corresponding to the target text contained in each target area to generate text data corresponding to each target area includes: Determining a sorting relationship between the target characters based on the position information of the target characters; Based on the order relationship between the target characters and the content information corresponding to the target characters, text data corresponding to each target area is generated.
9. The method according to claim 5, characterized in that After generating the text data corresponding to each target area, the method further includes: Outputting text data corresponding to each target area; receiving response data corresponding to each target area, wherein the response data is obtained by modifying the text data corresponding to each target area; The text data and / or response data corresponding to the at least one target area are combined to generate text data corresponding to the first text image.
10. The method according to claim 5, characterized in that Combining the text data corresponding to the at least one target area to generate text data corresponding to the first text image includes: Combining the type of each target area with the text data corresponding to each target area to generate structured data for each target area; The structured data of the at least one target area are combined to obtain text data corresponding to the first text image.
11. The method according to claim 10, characterized in that Combining the structured data of the at least one target area to obtain text data corresponding to the first text image includes: Get the target structured template; The structured data of the at least one target area are combined according to the target structured template to generate text data corresponding to the first text image.
12. The method according to claim 11, characterized in that Obtaining the target structured template includes one of the following: receiving the target structured template; A structured template selected from the output multiple structured templates is received to obtain the target structured template.
13. A text processing method, characterized in that: include: Displaying a first text image, wherein the first text image includes a plurality of characters that are typeset according to a first typeset manner, and the first text image includes at least one target area, and different target areas have different typesets; marking, in the first text image, a first recognition result of a plurality of characters in the first text image, wherein the first recognition result is obtained by recognizing the first text image; marking, in the first text image, a second recognition result of at least one target area in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area location information of each target area, and a type of layout corresponding to each target area; The text data corresponding to the first text image is displayed, wherein the text data is generated by performing area matching of the first recognition result and the second recognition result to determine the target text contained in each target area, and then sorting the target text contained in at least one target area according to the reading order between the target texts. The text data is typeset in a second typesetting method determined by the user's reading habits.
14. A text processing method, characterized in that: include: Receive a first text image, wherein the first text image includes a plurality of characters arranged in a first typesetting manner, and the first text image includes at least one target area, and different target areas have different typesets; Recognize the first text image to obtain a first recognition result of a plurality of characters in the first text image; Detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and a type of layout corresponding to each target area; generating text data corresponding to the first text image based on the first recognition result and the second recognition result; outputting text data corresponding to the first text image; Generating text data corresponding to the first text image based on the first recognition result and the second recognition result includes: Performing area matching on the first recognition result and the second recognition result to determine the target text contained in each target area; The target characters contained in the at least one target area are sorted according to a reading order between the target characters to generate text data corresponding to the first text image, wherein the text data is typeset in a second typesetting manner determined by a user's reading habits.
15. A text processing method, characterized in that: include: Acquire an ancient book image, wherein a plurality of characters contained in the ancient book image are typeset according to a first typeset mode, and the ancient book image includes at least one target area, and different target areas have different typesets; Recognizing the ancient book image to obtain a first recognition result of a plurality of characters in the ancient book image; Detecting the ancient book image to obtain a second recognition result of at least one target area in the ancient book image, wherein the second recognition result includes: area location information of each target area and the type of layout corresponding to each target area; generating an ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result; Generating the ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result includes: The first recognition result and the second recognition result are matched by area to determine the target text contained in each target area; the target text contained in at least one target area is sorted according to the reading order between the target texts to generate the ancient book text corresponding to the ancient book image, wherein the ancient book text is typeset according to the second typesetting method determined by the user's reading habits.
16. The method according to claim 15, characterized in that Detecting the ancient book image to obtain a second recognition result of at least one target area in the ancient book image includes: The ancient book image is detected using a target detection model to obtain the second recognition result.
17. The method according to claim 15, characterized in that The first recognition result includes: content information corresponding to each character, and character position information of each character, wherein generating the ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result includes: Based on the character position information of each character and the area position information of each target area, matching each character with each target area to determine the target character contained in each target area; Merging the content information corresponding to the target text contained in each target area to generate text data corresponding to each target area; The text data corresponding to the at least one target area is combined to generate the ancient book text.
18. A text processing device, characterized in that: include: an acquisition module, configured to acquire a first text image, wherein the first text image includes a plurality of characters arranged in a first typesetting manner, and the first text image includes at least one target area, and different target areas have different typesets; a recognition module, configured to recognize the first text image and obtain a first recognition result of a plurality of characters in the first text image; a detection module, configured to detect the first text image and obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and a type of layout corresponding to each target area; a generating module, configured to generate text data corresponding to the first text image based on the first recognition result and the second recognition result, wherein the text data is a serialized text obtained by sorting the plurality of characters according to the sorting relationship, the second typesetting method, and the layout type in the second recognition result; The text processing device further includes a matching module for performing area matching on the first recognition result and the second recognition result to determine the target text contained in each target area; The generation module is further used to sort the target text contained in the at least one target area according to the reading order between the target text, and generate text data corresponding to the first text image, wherein the text data is typeset according to a second typesetting method determined by the user's reading habits.
19. A text processing device, characterized in that: include: A first display module is configured to display a first text image, wherein the first text image includes a plurality of characters arranged in a first typesetting manner, and the first text image includes at least one target area, and different target areas have different typesets; a first marking module, configured to mark, in the first text image, first recognition results of a plurality of characters in the first text image; a second marking module, configured to mark, in the first text image, a second recognition result of at least one target area in the first text image, wherein the second recognition result is obtained by detecting the first text image, and the second recognition result includes: area location information of each target area, and a type of layout corresponding to each target area; A second display module is used to display text data corresponding to the first text image, wherein the text data is generated by performing area matching on the first recognition result and the second recognition result, determining the target text contained in each target area, and then sorting the target text contained in at least one target area according to the reading order between the target texts. The text data is typeset in a second typesetting method determined by the user's reading habits.
20. A text processing device, characterized in that: include: a receiving module configured to receive a first text image, wherein the first text image includes a plurality of characters arranged in a first typesetting manner, and the first text image includes at least one target area, and different target areas have different typesets; a recognition module, configured to recognize the first text image and obtain a first recognition result of a plurality of characters in the first text image; a detection module, configured to detect the first text image and obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and a type of layout corresponding to each target area; a generating module, configured to generate text data corresponding to the first text image based on the first recognition result and the second recognition result, wherein the text data is a serialized text obtained by sorting the plurality of characters according to the sorting relationship, the second typesetting method, and the layout type in the second recognition result; an output module, configured to output text data corresponding to the first text image; The generation module is further used to perform area matching on the first recognition result and the second recognition result to determine the target text contained in each target area; sort the target text contained in at least one target area according to the reading order between the target text, and generate text data corresponding to the first text image, wherein the text data is typeset according to a second typesetting method determined by the user's reading habits.
21. A text processing device, characterized in that: include: an acquisition module, configured to acquire an ancient book image, wherein a plurality of characters contained in the ancient book image are typeset according to a first typeset mode; a recognition module, configured to recognize the ancient book image and obtain a first recognition result of a plurality of characters in the ancient book image; a detection module, configured to detect the ancient book image and obtain a second recognition result of at least one target area in the ancient book image, wherein the second recognition result includes: area location information of each target area and a type of layout corresponding to each target area; a generating module, configured to generate an ancient book text corresponding to the ancient book image based on the first recognition result and the second recognition result; The generation module is further used to perform area matching on the first recognition result and the second recognition result to determine the target text contained in each target area; sort the target text contained in at least one target area according to the reading order between the target texts, and generate the ancient book text corresponding to the ancient book image, wherein the ancient book text is typeset according to the second typesetting method determined by the user's reading habits.
22. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the text processing method according to any one of claims 1 to 17.
23. A computer terminal, characterized in that: The invention comprises a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the text processing method according to any one of claims 1 to 17 is executed when the program is run.
24. A text processing system, characterized in that: include: processor; as well as The memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining a first text image, wherein a plurality of characters contained in the first text image are typeset according to a first typeset mode, and the first text image includes at least one target area, and the layouts of different target areas are different; recognizing the first text image to obtain a first recognition result of the plurality of characters in the first text image; detecting the first text image to obtain a second recognition result of at least one target area in the first text image, wherein the second recognition result includes: area location information of each target area and a type of layout corresponding to each target area; generating text data corresponding to the first text image based on the first recognition result and the second recognition result; the generating text data corresponding to the first text image based on the first recognition result and the second recognition result includes: Performing area matching on the first recognition result and the second recognition result to determine the target text contained in each target area; sorting the target text contained in at least one target area according to the reading order of the target text to generate text data corresponding to the first text image, wherein, The text data is typeset according to a second typeset method determined by the user's reading habits.
Citation Information
Patent Citations
Structured processing method and device of image, storage medium and electronic equipment
CN111144210A
Text classification and recognition method and device based on target detection
CN112036395A