Data conversion method and device based on OCR, storage medium and electronic device
The OCR model extracts text and image recognition technology in PDF files and performs information fusion processing, which solves the problem that PDF files cannot be edited and realizes high-quality document conversion and editing functions.
Patent Information
- Application Number
- CN202311552700.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to convert PDF files into editable document files, resulting in users being unable to edit the received PDF files, which brings inconvenience to users.
The pre-set OCR model extracts the pending text from the PDF file, and extracts the layout information of non-text content through image recognition technology, performs information fusion processing, and generates editable document files.
It realizes the function of converting PDF files into editable document files, highly restores the layout format in the original file, and users can directly edit the generated documents without additional processing.
Smart Images

Figure CN120030991A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an OCR-based data conversion method, device, storage medium and electronic device. Background Art
[0002] Portable Document Format (PDF) is a common file format, widely used in daily work and life. PDF is a layout document, and the pages are relatively independent, so that the document layout can be accurately described and the document layout can be displayed.
[0003] In the related art, in order to prevent the document content from being tampered with during transmission, in most cases, document files such as WORD need to be converted into PDF files before transmission. However, since PDF files do not record the logical structure of the document, the end user cannot edit the PDF file after receiving it, which brings inconvenience to the user.
[0004] Based on this, there is an urgent need for a data conversion method that can convert a PDF file into an editable document file to facilitate users to edit it. Summary of the invention
[0005] The purpose of this application is to provide an OCR-based data conversion method, device, storage medium and electronic device, which are used to convert PDF files into editable document files and highly restore the typesetting format in the original file. Users can directly edit the generated editable document without performing additional processing.
[0006] The present application provides a data conversion method based on OCR, comprising:
[0007] The text to be processed is extracted from the first file by a preset OCR model, and the layout information of the non-text content is extracted from the first file by image recognition technology; the text to be processed and the layout information are fused to obtain an editable second file; wherein the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the layout information includes: location information and layout information of the non-text content.
[0008] Optionally, the extracting of the layout information of the non-text content from the first file by using the image recognition technology includes: identifying a target illustration and a first text adjacent to the target illustration from the first file by using the image recognition technology; wherein the target illustration is: any one of the at least one non-text content identified from the first file by using the image recognition technology.
[0009] Optionally, after identifying the target illustration and the first text adjacent to the target illustration from the first file through image recognition technology, the method further includes: processing the first text according to a preset processing method to obtain a second text, and processing the text to be processed according to the preset processing method to obtain a third text; determining the first position information of the second text in the third text based on the string query result, and determining the second position information of the target illustration according to the first position information and the number of bytes of the preceding text and the succeeding text in the second text; wherein the preset processing method is used to remove space information and line break information in the first text; the first position information and the second position information are both expressed based on the number of bytes; and the second position information is used to indicate the relative position relationship between the target illustration and the second text.
[0010] Optionally, the first text is processed according to a preset processing method to obtain a second text, and the text to be processed is processed according to the preset processing method to obtain a third text, including: searching for a space identifier and a line break identifier in the target text to be processed, and obtaining space information corresponding to the space identifier and line break information corresponding to the line break identifier; deleting the space identifier based on the space information, and deleting the line break identifier based on the line break information; wherein the target text to be processed is the first text, or any one of the texts to be processed.
[0011] Optionally, after identifying the target illustration and the first text adjacent to the target illustration from the first file through image recognition technology, the method also includes: determining the relative position relationship between the target illustration and the first text through image recognition technology; screening out a target layout that matches the relative position relationship from a layout library, and obtaining layout information corresponding to the target layout.
[0012] Optionally, the text to be processed and the typesetting information are subjected to information fusion processing to obtain an editable second file, including: identifying and typeset the input content of the text to be processed to obtain a semantically coherent target text with a title, and inputting the target text into a target file; inserting non-text content into the target file containing the target text according to the position information of the non-text content indicated by the typesetting information; adjusting the layout of the non-text content inserted into the target file according to the layout information indicated by the typesetting information, and generating the second file.
[0013] Optionally, inserting the non-text content into the target file containing the target text according to the position information of the non-text content indicated by the typesetting information includes: determining row and column information based on the position information of the non-text content indicated by the typesetting information, and inserting the non-text content into the target file according to the row number indicated by the row and column information and the column number indicated by the row and column information; and adjusting the layout of the non-text content inserted into the target file according to the layout information indicated by the typesetting information includes: adjusting the placement position and text wrapping method of the non-text content according to the layout information indicated by the typesetting information.
[0014] The present application also provides a data conversion device based on OCR, comprising:
[0015] An information acquisition module is used to extract the text to be processed from the first file through a preset OCR model, and to extract the layout information of the non-text content from the first file through image recognition technology; an information processing module is used to perform information fusion processing on the text to be processed and the layout information to obtain an editable second file; wherein, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the layout information includes: location information and layout information of the non-text content.
[0016] Optionally, the information acquisition module is specifically used to identify a target illustration and a first text adjacent to the target illustration from the first file through image recognition technology; wherein the target illustration is: any one of at least one non-text content identified from the first file through image recognition technology.
[0017] Optionally, the information acquisition module is specifically used to process the first text according to a preset processing method to obtain a second text, and to process the text to be processed according to the preset processing method to obtain a third text; the information acquisition module is also specifically used to determine the first position information of the second text in the third text based on the string query result, and determine the second position information of the target illustration based on the first position information and the number of bytes of the preceding text and the succeeding text in the second text; wherein the preset processing method is used to remove space information and line break information in the first text; the first position information and the second position information are both expressed based on the number of bytes; the second position information is used to indicate the relative position relationship between the target illustration and the second text.
[0018] Optionally, the information acquisition module is specifically used to search for space identifiers and line break identifiers in the target text to be processed, and obtain space information corresponding to the space identifiers and line break information corresponding to the line break identifiers; the information acquisition module is also specifically used to delete the space identifiers based on the space information, and delete the line break identifiers based on the line break information; wherein the target text to be processed is the first text, or any item of the texts to be processed.
[0019] Optionally, the information acquisition module is specifically used to determine the relative position relationship between the target illustration and the first text through image recognition technology; the information acquisition module is also specifically used to filter out a target layout that matches the relative position relationship from a layout library, and obtain layout information corresponding to the target layout.
[0020] Optionally, the information processing module is specifically used to identify and typeset the text input content to be processed, obtain a semantically coherent target text with a title, and input the target text into a target file; the information processing module is also specifically used to insert non-text content into a target file containing the target text according to the position information of the non-text content indicated by the typesetting information; the information processing module is also specifically used to adjust the layout of the non-text content inserted into the target file according to the layout information indicated by the typesetting information, and generate the second file.
[0021] Optionally, the information processing module is specifically used to determine row and column information based on the position information of the non-text content indicated by the typesetting information, and insert the non-text content into the target file according to the row number indicated by the row and column information and the column number indicated by the row and column information; the information processing module is also specifically used to adjust the placement position and text wrapping method of the non-text content according to the layout information indicated by the typesetting information.
[0022] The present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the steps of any of the above-mentioned OCR-based data conversion methods through the computer program.
[0023] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein the program, when executed at runtime, implements the steps of any of the above-mentioned OCR-based data conversion methods.
[0024] The present application also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned OCR-based data conversion methods are implemented.
[0025] The OCR-based data conversion method, device, storage medium and electronic device provided in the present application first extract the text to be processed from the first file by a preset OCR model, and extract the typesetting information of the non-text content from the first file by image recognition technology; then, the text to be processed and the typesetting information are subjected to information fusion processing to obtain an editable second file. Among them, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of the non-text content. In this way, the PDF file can be directly converted into an editable document file, and the typesetting format in the original file can be highly restored. The user can directly edit the generated editable document without any additional processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0027] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 is a schematic diagram of a hardware environment of an interaction method of a smart device according to an embodiment of the present application;
[0029] Figure 2 It is a flowchart of the OCR-based data conversion method provided by this application;
[0030] Figure 3 It is a schematic diagram of illustration positioning of the OCR-based data conversion method provided by this application;
[0031] Figure 4 is a schematic diagram of the structure of the OCR-based data conversion device provided by the present application;
[0032] Figure 5 It is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] According to one aspect of an embodiment of the present application, a data conversion method based on OCR is provided. The data conversion method based on OCR is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned data conversion method based on OCR can be applied to Figure 1 In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.
[0036] The network may include but is not limited to at least one of the following: wired network, wireless network. The wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be but is not limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.
[0037] The following is an explanation of the professional terms involved in the embodiments of this application:
[0038] Optical Character Recognition (OCR) refers to the process in which an electronic device (such as a scanner or digital camera) examines characters printed on paper, determines their shapes by detecting dark and light patterns, and then uses character recognition methods to translate the shapes into computer text; that is, for printed characters, an optical method is used to convert the text in a paper document into a black and white dot matrix image file, and recognition software is used to convert the text in the image into text format for further editing and processing by word processing software.
[0039] The Generative Pre-Trained Transformer (GPT) model is a deep learning model for text generation based on the Internet and trained with available data. After obtaining the text recognized by OCR, you can write a prompt for sentence splicing, merge them and input them into GPT, which will output the spliced text. Then write a prompt for title extraction to extract the title from the text. In the next step, perform the same operation and write a prompt for paragraph splitting, input it into GPT, and output the segmented text and title. Finally, save the result in md file format and distinguish the titles in the md file.
[0040] The purpose of designing PDF is to support cross-platform, multimedia integrated information publishing and release, especially to provide support for network information release. To achieve this goal, PDF has many advantages that other electronic document formats cannot match. The PDF file format can encapsulate text, fonts, formats, colors, and graphics and images independent of equipment and resolution in one file. This format file can also contain electronic information such as hypertext links, sounds and dynamic images, support special files, and have high integration and security reliability.
[0041] For ordinary users, e-books made with PDF have the texture and reading effect of paper books, can "realistically" show the original appearance of the original book, and the display size can be adjusted arbitrarily, providing readers with a personalized reading method. Since PDF files do not rely on the language and fonts of the operating system and display devices, they are very convenient to read. These advantages enable readers to quickly adapt to electronic reading and online reading, which is undoubtedly conducive to the popularization of computers and the Internet in daily life.
[0042] After receiving the PDF file, if the end user wants to edit it, the PDF file can be converted into a WORD file through some format conversion software or websites in the related art. However, since the PDF file does not record the logical structure of the document, the WORD file obtained after the format conversion in the related art has the technical problem of chaotic typesetting, especially when the PDF file content contains pictures, formulas and tables, the converted WORD file needs to be processed before it can be edited normally.
[0043] In response to the above-mentioned technical problems existing in the related technology, an embodiment of the present application provides a data conversion method based on OCR, which can convert a PDF file into an editable document file and highly restore the typesetting format in the original file. Users can directly edit the generated editable document without any additional processing.
[0044] The following is a detailed description of the OCR-based data conversion method provided by the embodiment of the present application through specific embodiments and application scenarios in conjunction with the accompanying drawings.
[0045] like Figure 2 As shown, an embodiment of the present application provides an OCR-based data conversion method, which may include the following steps 201 and 202:
[0046] Step 201: extract the text to be processed from the first file by using a preset OCR model, and extract the layout information of the non-text content from the first file by using image recognition technology.
[0047] Among them, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of non-text content.
[0048] Exemplarily, the first file may be a PDF file, or an image file in a jpg, bmp, png, etc. format. Since PDF files and image files in jpg, bmp, png, etc. formats can be freely converted, the first file is mainly described as a PDF file in the embodiments of the present application.
[0049] Exemplarily, the first file may be an original document file that is printed as a PDF or saved as a PDF. The original document file may be a document file in a format such as word, ppt, excel, or txt.
[0050] It should be noted that the original files described in the embodiments of the present application are all the above-mentioned original document files. In the related art, the user converts the original file into a PDF file and sends it to the target user. After receiving the PDF file, the target user can usually only read it but cannot edit it. After processing the PDF file using the OCR-based data conversion method provided in the embodiments of the present application, an editable file with a layout and content highly similar to the original file is obtained, and the target user can edit the content in the editable file.
[0051] Specifically, the step of extracting the layout information of the non-text content from the first file by using the image recognition technology in the above step 201 may include the following steps 201a:
[0052] Step 201a: Identify a target illustration and a first text adjacent to the target illustration from the first file by using image recognition technology.
[0053] The target illustration is: any one of at least one non-text content identified from the first file by image recognition technology.
[0054] Exemplarily, at least one non-text content can be identified from the first file by image recognition technology, and the non-text content can be any of the following: table content, image content, and formula content. That is, in the embodiment of the present application, the table content, image content, and formula content in the first file are all treated as illustrations for processing, and even after the first file is converted into an editable file, the table content, image content, and formula content cannot be edited.
[0055] It should be noted that the OCR-based data conversion method provided in the embodiment of the present application can be combined with the recognition and restoration methods for table content and formula content in the relevant technology, so as to achieve the technical effect of editing the above-mentioned table content, image content and formula content in the final editable file.
[0056] Exemplarily, the first text is the text adjacent to the target illustration among all the texts contained in the first file. Figure 3 As shown, in the PDF file, text A and text B are adjacent to the image. Therefore, in order to accurately locate the position of the image in the PDF file, a search needs to be performed based on text A and text B (ie, the first text mentioned above).
[0057] Specifically, after the above step 201a, the position information of the non-text content may be determined by the following steps, that is, the above step 201 may further include the following steps 201b1 and 201b2:
[0058] Step 201b1, processing the first text according to a preset processing method to obtain a second text, and processing the to-be-processed text according to the preset processing method to obtain a third text.
[0059] The preset processing method is used to remove space information and line break information in the first text.
[0060] Exemplarily, the first text may be obtained by recognizing only the region including the target illustration and the first text through a preset OCR model.
[0061] Specifically, the above step 201b1 may further include the following steps 201b11 and 201b12:
[0062] Step 201b11, searching for a space identifier and a line break identifier in the target text to be processed, and obtaining space information corresponding to the space identifier and line break information corresponding to the line break identifier.
[0063] Step 201b12: Delete the space identifier based on the space information, and delete the line break identifier based on the line break information.
[0064] The target text to be processed is the first text, or any one of the texts to be processed.
[0065] Exemplarily, the specific implementation of the above preset processing method includes: after finding the space symbol or line break symbol in the text, directly deleting it through the application program interface.
[0066] Step 201b2, determine the first position information of the second text in the third text based on the string query result, and determine the second position information of the target illustration according to the first position information and the number of bytes of the preceding text and the succeeding text in the second text.
[0067] The first position information and the second position information are both expressed based on the number of bytes; the second position information is used to indicate the relative position relationship between the target illustration and the second text.
[0068] Exemplarily, the length of the preceding text and the succeeding text may be user-defined, or the data may be sampled as a whole line.
[0069] For example, after obtaining the relative position information of the first text in the second text, the position of the target illustration in the second text can be further calculated based on the relative position information of the target illustration and the first text. Figure 3 As shown, assuming that the starting position of the text consisting of text A (i.e. the above-mentioned previous text) and text B (i.e. the above-mentioned following text) is the position of the 9527th byte in the entire article (i.e. the above-mentioned second text), and the image is located after text A and before text B, then, based on the number of bytes of text A and the number of bytes of text B, the position of the image in the entire article can be calculated (expressed in the number of bytes).
[0070] It should be noted that the text to be processed obtained by recognizing the above-mentioned first file through the preset OCR model and the above-mentioned first text recognized by the image recognition technology may contain special symbols such as spaces and line breaks in the text content. Therefore, in order to accurately calculate the position information of the target illustration, it is necessary to remove the above-mentioned special symbols in the text content to avoid interference with the position information represented based on the number of bytes.
[0071] Specifically, after the above step 201a, the layout information of the non-text content may be determined by the following steps, that is, the above step 201 may further include the following steps 201c1 and 201c2:
[0072] Step 201c1: Determine the relative position relationship between the target illustration and the first text by using image recognition technology.
[0073] Step 201c2: Filter out a target layout that matches the relative position relationship from the layout library, and obtain layout information corresponding to the target layout.
[0074] For example, in a WORD file, the layout between images and texts may include embedded, all-around, tight, through, up and down, and other layout methods. By combining the characteristics of each layout method, image recognition technology can be used to identify the layout type of the target illustration in the first file, and then obtain the corresponding layout information.
[0075] Step 202: perform information fusion processing on the text to be processed and the typesetting information to obtain an editable second file.
[0076] Exemplarily, after obtaining the position information and layout information (ie, the typesetting information) of the text to be processed and the non-text content, the final file can be generated step by step based on the obtained information.
[0077] Exemplarily, the generated second file may also retain relevant font information such as the font, font size, color, etc. of the text content in the first file.
[0078] Specifically, the above step 202 may include the following steps 202a1 to 202a3:
[0079] Step 202a1, recognize and typeset the input text to be processed to obtain a semantically coherent target text with a title, and input the target text into a target file.
[0080] Exemplarily, the above language model may be a GPT language model or other generative language models. In the embodiments of the present application, the language model is described as a GPT language model.
[0081] Step 202a2: insert the non-text content into the target file containing the target text according to the position information of the non-text content indicated by the typesetting information.
[0082] Specifically, the above step 202a2 may further include the following step 202a21:
[0083] Step 202a21, based on the position information of the non-text content indicated by the typesetting information, determine the row and column information, and insert the non-text content into the target file according to the row number indicated by the row and column information and the column number indicated by the row and column information.
[0084] Exemplarily, based on the position information of the non-text content indicated by the above layout information, the row number and column number where the non-text content needs to be inserted can be determined, and then the non-text content can be inserted into the position indicated by the row number and column number.
[0085] Step 202a3: adjust the layout of the non-text content inserted into the target file according to the layout information indicated by the typesetting information, and generate the second file.
[0086] Exemplarily, the GPT language model is first used to process the text to be processed to obtain a target text that is semantically coherent and has a title, so as to increase the readability of the generated editable file. Afterwards, the text content output by the GPT language model is input into the target file. The target file is an intermediate file of the second file.
[0087] Exemplarily, after the text content output by the GPT language model is input into the target file, the next step is to insert the non-text content into the corresponding position in the form of a picture according to the position information of the non-text content indicated by the above-mentioned typesetting information.
[0088] For example, finally, after all the content is inserted into the target file, the layout of all the non-text content inserted into the target file can be adjusted according to the layout information indicated by the typesetting information, thereby obtaining the final second file.
[0089] Specifically, the above step 202a3 may further include the following step 202a31:
[0090] Step 202a31: adjust the placement position and text wrapping mode of the non-text content according to the layout information indicated by the typesetting information.
[0091] Exemplarily, the above-mentioned placement positions may include: left-aligned, right-aligned, centered, etc.; the above-mentioned text wrapping methods may include: embedded, all-around, tight wrapping, through-wrapping, top-bottom wrapping, etc.
[0092] In this way, the PDF file can be directly converted into an editable document file, and the typesetting format in the original file can be highly restored. Users can directly edit the generated editable document without any additional processing.
[0093] The OCR-based data conversion method provided in the embodiment of the present application first extracts the text to be processed from the first file by a preset OCR model, and extracts the typesetting information of the non-text content from the first file by image recognition technology; then, the text to be processed and the typesetting information are subjected to information fusion processing to obtain an editable second file. Among them, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of the non-text content. In this way, the PDF file can be directly converted into an editable document file, and the typesetting format in the original file can be highly restored. The user can directly edit the generated editable document without any additional processing.
[0094] It should be noted that the OCR-based data conversion method provided in the embodiment of the present application can be executed by an OCR-based data conversion device, or a control module in the OCR-based data conversion device for executing the OCR-based data conversion method. In the embodiment of the present application, an OCR-based data conversion device executing the OCR-based data conversion method is taken as an example to illustrate the OCR-based data conversion device provided in the embodiment of the present application.
[0095] It should be noted that in the embodiments of the present application, the data conversion methods based on OCR shown in the drawings of the above-mentioned methods are all illustrated by combining an accompanying drawing in the embodiments of the present application as an example. In specific implementation, the data conversion methods based on OCR shown in the accompanying drawings of the above-mentioned methods can also be implemented in combination with any other accompanying drawings that can be combined as shown in the above-mentioned embodiments, which will not be repeated here.
[0096] The OCR-based data conversion device provided by the present application is described below, and the OCR-based data conversion method described below and described above can be referenced to each other.
[0097] Figure 4 A schematic diagram of the structure of an OCR-based data conversion device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, specifically including:
[0098] The information acquisition module 401 is used to extract the text to be processed from the first file through a preset OCR model, and to extract the layout information of the non-text content from the first file through image recognition technology; the information processing module 402 is used to perform information fusion processing on the text to be processed and the layout information to obtain an editable second file; wherein, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the layout information includes: location information and layout information of the non-text content.
[0099] Optionally, the information acquisition module 401 is specifically used to identify a target illustration and a first text adjacent to the target illustration from the first file through image recognition technology; wherein the target illustration is: any one of at least one non-text content identified from the first file through image recognition technology.
[0100] Optionally, the information acquisition module 401 is specifically used to process the first text according to a preset processing method to obtain a second text, and to process the text to be processed according to the preset processing method to obtain a third text; the information acquisition module 401 is also specifically used to determine the first position information of the second text in the third text based on the string query result, and determine the second position information of the target illustration according to the first position information and the number of bytes of the preceding text and the succeeding text in the second text; wherein the preset processing method is used to remove space information and line break information in the first text; the first position information and the second position information are both expressed based on the number of bytes; the second position information is used to indicate the relative position relationship between the target illustration and the second text.
[0101] Optionally, the information acquisition module is specifically used to search for space identifiers and line break identifiers in the target text to be processed, and obtain space information corresponding to the space identifiers and line break information corresponding to the line break identifiers; the information acquisition module is also specifically used to delete the space identifiers based on the space information, and delete the line break identifiers based on the line break information; wherein the target text to be processed is the first text, or any item of the texts to be processed.
[0102] Optionally, the information acquisition module 401 is specifically used to determine the relative position relationship between the target illustration and the first text through image recognition technology; the information acquisition module 401 is also specifically used to filter out a target layout that matches the relative position relationship from a layout library, and obtain layout information corresponding to the target layout.
[0103] Optionally, the information processing module 402 is specifically used to input the text to be processed into a text recognition typesetting model to obtain a semantically coherent target text with a title, and input the target text into a target file; the information processing module 402 is also specifically used to insert non-text content into the target file containing the target text according to the position information of the non-text content indicated by the typesetting information; the information processing module 402 is also specifically used to adjust the layout of the non-text content inserted into the target file according to the layout information indicated by the typesetting information, and generate the second file.
[0104] Optionally, the information processing module is specifically used to determine row and column information based on the position information of the non-text content indicated by the typesetting information, and insert the non-text content into the target file according to the row number indicated by the row and column information and the column number indicated by the row and column information; the information processing module is also specifically used to adjust the placement position and text wrapping method of the non-text content according to the layout information indicated by the typesetting information.
[0105] The OCR-based data conversion device provided in the present application first extracts the text to be processed from the first file by means of a preset OCR model, and extracts the typesetting information of the non-text content from the first file by means of image recognition technology; then, the text to be processed and the typesetting information are subjected to information fusion processing to obtain an editable second file. Among them, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of the non-text content. In this way, the PDF file can be directly converted into an editable document file, and the typesetting format in the original file can be highly restored. The user can directly edit the generated editable document without any additional processing.
[0106] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the OCR-based data conversion method, which includes: extracting the text to be processed from the first file by a preset OCR model, and extracting the typesetting information of the non-text content from the first file by an image recognition technology; performing information fusion processing on the text to be processed and the typesetting information to obtain an editable second file; wherein the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of the non-text content.
[0107] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0108] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the OCR-based data conversion method provided by the above-mentioned methods, and the method includes: extracting the text to be processed from the first file through a preset OCR model, and extracting the typesetting information of the non-text content from the first file through image recognition technology; performing information fusion processing on the text to be processed and the typesetting information to obtain an editable second file; wherein the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of the non-text content.
[0109] On the other hand, the present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, the OCR-based data conversion method provided by the above-mentioned methods is executed, and the method includes: extracting the text to be processed from the first file through a preset OCR model, and extracting the typesetting information of the non-text content from the first file through image recognition technology; performing information fusion processing on the text to be processed and the typesetting information to obtain an editable second file; wherein the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of the non-text content.
[0110] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data conversion method based on OCR, It is characterized in that include: Extracting the text to be processed from the first file by using a preset OCR model, and extracting the layout information of the non-text content from the first file by using image recognition technology; Performing information fusion processing on the text to be processed and the typesetting information to obtain an editable second file; Among them, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of non-text content.
2. The OCR-based data conversion method according to claim 1, It is characterized in that The extracting the layout information of the non-text content from the first file by using the image recognition technology includes: Identify a target illustration and a first text adjacent to the target illustration from the first file by using an image recognition technology; The target illustration is: any one of at least one non-text content identified from the first file by image recognition technology.
3. The OCR-based data conversion method according to claim 2, It is characterized in that After identifying the target illustration and the first text adjacent to the target illustration from the first file by using the image recognition technology, the method further includes: Processing the first text according to a preset processing method to obtain a second text, and processing the to-be-processed text according to the preset processing method to obtain a third text; Determine first position information of the second text in the third text based on the string query result, and determine second position information of the target illustration according to the first position information and the number of bytes of the preceding text and the succeeding text in the second text; Among them, the preset processing method is used to remove space information and line break information in the text; the first position information and the second position information are both expressed based on the number of bytes; the second position information is used to indicate the relative position relationship between the target illustration and the second text.
4. The OCR-based data conversion method according to claim 3, It is characterized in that The step of processing the first text according to a preset processing method to obtain a second text, and processing the to-be-processed text according to the preset processing method to obtain a third text, includes: Searching for a space identifier and a line break identifier in the target text to be processed, and obtaining space information corresponding to the space identifier and line break information corresponding to the line break identifier; Deleting the space identifier based on the space information, and deleting the line break identifier based on the line break information; The target text to be processed is the first text, or any one of the texts to be processed.
5. The OCR-based data conversion method according to claim 2, It is characterized in that After identifying the target illustration and the first text adjacent to the target illustration from the first file by using the image recognition technology, the method further includes: Determine the relative position relationship between the target illustration and the first text by using image recognition technology; A target layout matching the relative position relationship is screened out from the layout library, and layout information corresponding to the target layout is obtained.
6. The OCR-based data conversion method according to claim 1, It is characterized in that The step of fusing the text to be processed and the typesetting information to obtain an editable second file includes: Recognize and typeset the input text to be processed to obtain a semantically coherent target text with a title, and input the target text into a target file; inserting the non-text content into a target file containing the target text according to the position information of the non-text content indicated by the typesetting information; The non-text content inserted into the target file is adjusted in layout according to the layout information indicated by the typesetting information, and the second file is generated.
7. The OCR-based data conversion method according to claim 6, It is characterized in that The step of inserting the non-text content into a target file containing the target text according to the position information of the non-text content indicated by the typesetting information comprises: Based on the position information of the non-text content indicated by the typesetting information, determine row and column information, and insert the non-text content into the target file according to the row number indicated by the row and column information and the column number indicated by the row and column information; The step of adjusting the layout of the non-text content inserted into the target file according to the layout information indicated by the typesetting information includes: The placement position and text wrapping mode of the non-text content are adjusted according to the layout information indicated by the typesetting information.
8. A data conversion device based on OCR, It is characterized in that The device comprises: An information acquisition module, used to extract the text to be processed from the first file by using a preset OCR model, and to extract the typesetting information of the non-text content from the first file by using an image recognition technology; An information processing module, used for performing information fusion processing on the text to be processed and the typesetting information to obtain an editable second file; Among them, the first file is a non-editable file; the first file includes: text content and non-text content; the non-text content includes at least one of the following: table content, image content, formula content; the typesetting information includes: location information and layout information of non-text content.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes a stored program, wherein the OCR-based data conversion method according to any one of claims 1 to 7 is executed when the program is executed.
10. An electronic device comprising a memory and a processor, It is characterized in that The memory stores a computer program, and the processor is configured to execute the OCR-based data conversion method according to any one of claims 1 to 7 through the computer program.