Recognition Method, Device, Electronic Device and Storage Medium for Text Images
Through the image restoration model, the text images of smear traces are extracted and restored, which solves the problem that smear traces affect keyword recognition and achieves higher text recognition accuracy.
Patent Information
- Application Number
- CN202011415408.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-04
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-12-04
AI Technical Summary
In the prior art, the image restoration effect of the smear traces is poor, resulting in low keyword recognition accuracy.
The image restoration model is used to restore the text smear images of the smear traces through the keyword area, and the convolution neural network model is used to extract and deconvolution the smear traces to obtain the restored image and use optical character recognition technology to identify the text.
Without affecting the overall text restoration performance, the restoration effect of keyword areas is significantly improved and the accuracy of text recognition is improved.
Smart Images

Figure CN113537186B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and cloud technology. Specifically, the present application relates to a method, device, electronic device, and storage medium for recognizing text images. Background Art
[0002] Text image recognition refers to the technology of using a computer to capture text in an image, segment, and recognize the text content, and can be applied to many fields, such as reading, translation, retrieval of literature materials, sorting of letters and parcels, editing and proofreading of manuscripts, summarization and analysis of a large number of statistical reports and cards, processing of bank checks, statistical summarization of commercial invoices, recognition of product codes, management of product warehouses, etc.
[0003] When recognizing a text image, the obtained image may contain smeared images. If one wants to recognize the text in the smeared image, it is first necessary to perform restoration processing on the image to restore the image before smearing. Although there are already various image restoration schemes in the prior art, the accuracy of the text recognition result of the restored image obtained based on the prior art still needs to be improved. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, electronic device, and storage medium for recognizing text images. Based on this solution, the accuracy of text recognition can be effectively improved.
[0005] To achieve the above object, the specific technical solutions provided by the embodiments of the present application are as follows:
[0006] On the one hand, an embodiment of the present application provides a method for recognizing a text image, and the method includes:
[0007] Obtain a text smeared image in which the smeared trace passes through the keyword area; the keyword area is the image area where the keywords included in the text content of the text smeared image are located; the keywords are included in the keyword table;
[0008] Input the text smeared image into an image restoration model to restore the text content covered by the smeared trace in the text smeared image, and obtain a restored image corresponding to the text smeared image;
[0009] Perform text recognition on the text content in the restored image to obtain a text recognition result of the text smeared image.
[0010] On the other hand, an embodiment of the present invention also provides a device for recognizing a text image, and the device includes:
[0011] An image acquisition module, configured to acquire a text smear image in which a smear mark passes through a keyword area; the keyword area is an image area where keywords included in the text content of the text smear image are located; the keywords are included in a keyword table;
[0012] An image restoration module, configured to input the text smear image into the image restoration model, and restore the text content covered by the smear mark in the text smear image to obtain a restored image corresponding to the text smear image;
[0013] A character recognition module, configured to perform character recognition on the text content in the restored image to obtain a character recognition result of the text smear image.
[0014] An embodiment of the present invention further provides an electronic device, which includes one or more processors; a memory; one or more computer programs, wherein one or more computer programs are stored in the memory and are configured to be executed by one or more processors, and one or more computer programs are configured to execute the method as shown in the first aspect of the present application.
[0015] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store a computer program. When the computer program runs on a processor, the processor can execute the method as shown in the first aspect of the present application.
[0016] An embodiment of the present invention further provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the above-mentioned text image recognition method.
[0017] The beneficial effects brought by the technical solution provided by the present application are:
[0018] The present application provides a method, device, electronic device and storage medium for text image recognition. By using an image restoration model, a text smear image in which a smear mark passes through a keyword area is restored. Under the condition of not affecting the overall text restoration performance, the restoration effect of the keyword area in the image containing the smear mark is better, so as to facilitate character recognition for the restored image and improve the accuracy of character recognition. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application.
[0020] Figure 1 Schematic flowchart of a method for recognizing a text image provided by an embodiment of the present application;
[0021] Figure 2a Schematic diagram of an original text image provided by an embodiment of the present application;
[0022] Figure 2b Schematic diagram of the original text image after text region detection provided by an embodiment of the present application;
[0023] Figure 2c Schematic diagram of the text recognition result of the original text image provided by an embodiment of the present application;
[0024] Figure 2d Schematic diagram of the positions of each target keyword in the original text image provided by an embodiment of the present application;
[0025] Figure 3 Schematic diagram of the position of the first marked point in the original text image provided by an embodiment of the present application;
[0026] Figure 4 Schematic diagram of the positions of the first marked point and the second marked point in the original text image provided by an embodiment of the present application;
[0027] Figure 5 Schematic diagram of the connection of the smeared points in the original text image provided by an embodiment of the present application;
[0028] Figure 6 Schematic diagram of the original text image with smeared traces provided by an embodiment of the present application;
[0029] Figure 7 Schematic flowchart of the process for obtaining training samples provided by an embodiment of the present application;
[0030] Figure 8 Schematic diagram of the image to be restored provided by an embodiment of the present application;
[0031] Figure 9 Schematic diagram of the restored image provided by an embodiment of the present application;
[0032] Figure 10 Schematic diagram of the structure of the text image recognition device provided by an embodiment of the present application;
[0033] Figure 11 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0034] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.
[0035] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0036] The embodiment of the present application provides a method for recognizing a text image to solve the problem in the prior art that the restoration effect of keywords in an image with smear marks is not good, thus affecting the recognition of keywords. By using an image restoration model, the text smear image with smear marks passing through the keyword area is restored. Under the condition of not affecting the overall text restoration performance, the restoration effect of the keyword area in the image containing smear marks is better, so as to facilitate text recognition for the restored image and improve the accuracy of text recognition.
[0037] The execution subject of the technical solution of the present application is a computer device, including but not limited to a server, a personal computer, a laptop computer, a tablet computer, a smart phone, etc. The computer device includes a user device and a network device. Among them, the user device includes but not limited to a computer, a smart phone, a PAD, etc.; the network device includes but not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers in cloud computing. Among them, cloud computing is a kind of distributed computing and consists of a super virtual computer composed of a group of loosely coupled computer sets. Among them, the computer device can run alone to implement the present application, or can be connected to the network and implement the present application through interaction with other computer devices in the network. Among them, the network where the computer device is located includes but not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, etc.
[0038] The solution provided by the embodiment of the present application relates to fields such as cloud technology, big data, and artificial intelligence in computer technology.
[0039] In the embodiments of the present application, the data processing involved can be implemented through cloud technology, and the data calculation involved can be implemented through cloud computing in cloud technology.
[0040] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0041] As a basic capability provider of cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose from. The cloud computing resource pool mainly includes: computing devices (virtual machines, including operating systems), storage devices, and network devices.
[0042] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. The SaaS layer can also be directly deployed on the IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0043] Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources in a on-demand and easily scalable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a on-demand and easily scalable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance.
[0044] With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will drive a revolutionary change in the entire Internet model and enterprise management model conceptually.
[0045] The model training involved in the embodiments of this application can be achieved through machine learning in artificial intelligence technology.
[0046] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0047] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, natural language processing technology, and machine learning / deep learning.
[0048] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace the human eye to identify and measure targets, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0049] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph, etc. technologies.
[0050] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0051] The training data required for model training involved in the embodiments of this application can be big data obtained from the Internet.
[0052] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process a large amount of data that tolerates over time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.
[0053] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.
[0054] The embodiment of this application provides a method for recognizing a text image. The execution subject of this method can be any electronic device. For example, this method can be executed by a server, such as Figure 1 As shown, this method may include:
[0055] Step S101, obtaining a text smear image of the smear trace passing through the keyword area; the keyword area is the image area where the keywords included in the text content of the text smear image are located; the keywords are included in the keyword table;
[0056] Among them, the text smear image is a text image with a smear trace. The smear trace passes through the keyword area. The text content of the text smear image includes at least one keyword. The keyword can be a single keyword, or it can be a word composed of two or more keywords that are consecutive in the text content. The keyword area is the image area where at least one keyword is located, and the keywords are included in the keyword table. The keyword table is a set of keywords obtained by a computer or manually extracting keywords from documents and arranging them in a certain order. Multiple keywords can be determined in advance according to the actual application needs, and these keywords are used to establish the keyword table.
[0057] Since the smear trace partially obscures the text in the text image, it will affect the recognition of the text. Therefore, it is necessary to first perform restoration processing on the image with the smear trace and then perform text recognition.
[0058] Step S102, inputting the text smear image into an image restoration model to restore the text content covered by the smear trace in the text smear image, and obtaining a restored image corresponding to the text smear image;
[0059] Among them, the image restoration model can be a neural network model. The specific network structure of the image restoration model is not limited in the embodiments of the present application. Optionally, the image restoration model can be a Convolutional Neural Networks (CNN). Optionally, the specific structure of the image restoration model can include a cascaded encoder and decoder. Among them, the encoder can include at least one convolutional layer, and the decoder can include at least one deconvolutional layer. The text image to be recognized is subjected to feature extraction through the convolutional layer to obtain the feature map of the text image, and the extracted feature map is subjected to deconvolution processing through the deconvolutional layer to obtain the restored image.
[0060] Step S103: Perform text recognition on the text content in the restored image to obtain the text recognition result of the text smeared image.
[0061] The restored image obtained through the image restoration model no longer includes smearing marks or no longer includes all smearing marks, and the text in the restored image can be recognized. Optionally, the recognition of the text in the restored image can be implemented through Optical Character Recognition (OCR).
[0062] The text image recognition method provided by the embodiments of the present application uses an image restoration model to restore a text smeared image with smearing marks passing through the keyword area. Under the condition of not affecting the overall text restoration performance, the restoration effect of the keyword area in the image containing smearing marks is better, so as to facilitate text recognition for the restored image and improve the accuracy of text recognition.
[0063] In a possible implementation manner, the image restoration model is trained in the following way:
[0064] Obtain each training sample. Each training sample includes an original text image without smearing marks and a to-be-restored image with smearing marks corresponding to the original text image. The smearing marks in the to-be-restored image pass through the target keyword area, and the target keyword area is the image area where the keywords included in the text content of the to-be-restored image are located;
[0065] Iteratively train the initial image restoration model based on each training sample to obtain the image restoration model.
[0066] Among them, the smear marks in the image to be restored pass through the target keyword region, the target keyword is a word determined according to at least one keyword in the keyword table, and the target keyword region is the image region where the keyword included in the text content of the image to be restored is located. The text content of the image to be restored includes at least one keyword. The target keyword can be a single keyword, or can also be a word composed of two or more keywords with continuous positions in the text content.
[0067] Use the original text image without smear marks and the image to be restored with smear marks corresponding to the original text image as a training sample pair. Among them, the original text image without smear marks is used as supervision to iteratively train the initial image restoration model. During the training process, the model restores the image to be restored with smear marks corresponding to the original text image to obtain a restored image. According to the restored image output by the model and the original text image without smear marks, calculate the value of the loss function. Among them, the value of the loss function characterizes the difference between the restored image output by the model and the original text image without smear marks. Optionally, the value of the loss function can be obtained based on the pixel value difference between the restored image output by the model and the original text image without smear marks. The model obtained when the loss function converges is used as the image restoration model.
[0068] In the embodiment of the present application, the original text image without smear marks and the image to be restored with smear marks passing through the target keyword region are used as training samples to train the image restoration model. The image restoration model obtained in this way has a better restoration effect on the keyword region of the smeared image with smear marks passing through the keyword region without affecting the overall text restoration performance, thereby facilitating text recognition for the restored image and improving the accuracy of text recognition.
[0069] In a possible implementation manner, obtaining each training sample includes:
[0070] Obtain each original text image;
[0071] For each original text image, determine each target keyword region in the original text image;
[0072] For each original text image, perform a smearing process on each target keyword region in the original text image to obtain the image to be restored corresponding to the original text image.
[0073] In practical applications, for the text in an image with smudging marks, usually more attention is paid to certain keywords among them. Therefore, in order to improve the text recognition effect of the image, when restoring the image, it is necessary to improve the restoration ability of the image restoration model for the smudging marks in the keyword region. In order to improve the restoration ability for the keyword region, in the embodiments of the present application, when obtaining the training samples of the image restoration model, first determine each target keyword region in the original text image, and then perform smudging processing on each target keyword region. The image restoration model trained with the training samples obtained in this way can effectively improve the restoration effect of the smudging marks in the keyword region of the image to be restored, that is, it can better remove the smudging marks in the keyword region of the image. Therefore, when performing text recognition based on the image after removing the smudging marks, the text recognition effect can be effectively improved.
[0074] Optionally, when training an image restoration model applied to different fields, the keyword table corresponding to this field can be used. Since the keywords in different fields are different, different keyword tables can be determined according to the differences in application fields. For this application field, construct a keyword table with the keywords commonly used in the application field. When training an image restoration model for this application field, use the keyword table corresponding to this field to determine the target keyword regions of the original text image in the training samples, and smear the target keyword regions. The training samples obtained in this way are more targeted, and the image restoration model trained with such training samples is more conducive to the restoration of text images in this application field, and the restoration effect on the text in this application field is better.
[0075] In addition, in order to further improve the restoration effect of the image restoration model on the images in this application field, when determining the target keyword regions in the original text image, each target keyword region in the image to be restored in the training samples can include all the keywords in the keyword table corresponding to this field. This can make it such that when using the trained image restoration model for image restoration, the keyword regions in the smeared text image to be restored have all been smeared in the training samples, so that the coverage of the training samples can be more comprehensive. That is to say, the more comprehensive the training samples are covered, the more information the image restoration model trained with these training samples learns, which is more conducive to the restoration of all keyword regions in the image to be restored.
[0076] In some alternative embodiments, when training an image restoration model corresponding to the tourism field, a keyword list for the tourism field can be selected. Keywords determined according to this keyword list may include "train ticket", "air ticket", "hotel", "homestay", etc. The image regions corresponding to keywords such as "train ticket", "air ticket", "hotel", and "homestay" in the original text image of the training sample are used as target keyword regions. The above target keyword regions are smeared to obtain a smeared image corresponding to the original text image, thereby obtaining each training sample of the image restoration model corresponding to the tourism field. The image restoration model trained using these training samples has a better image restoration effect on the smeared images of keywords in the tourism field. Additionally, to make the training samples cover comprehensively, keywords determined by all keywords in the tourism field can be smeared in the original text image of the training sample. In this way, when using the trained image restoration model to perform image restoration in the tourism field, the keyword regions in the smeared text image to be restored have all been smeared in the training samples, so that the coverage of the training samples can be more comprehensive. The more information about the keywords in the tourism field the image restoration model learned from these training samples, the more conducive it is for the image restoration model to restore the target keyword regions in the image to be restored in the tourism field.
[0077] Optionally, if the trained image restoration model is to be applied to multiple fields, keywords in each of the multiple fields can be selected, and a keyword list common to the multiple fields can be established using the keywords in the multiple fields. When training the image restoration model for the multiple application fields, the keyword list common to the multiple fields is used to determine the target keyword regions of the original text images in the training samples, and the target keyword regions are smeared. In this way, the image restoration model trained with the obtained training samples can restore smeared text images in multiple fields. This may increase the number of training samples. However, when the image restoration model trained with such training samples is used, it has strong domain scalability and can be applicable to the restoration of text smeared images in multiple fields.
[0078] As an example of obtaining each training sample, N original text images containing text (generally, N >= 20000) can be collected, and a keyword list (which can also be called a keyword table) that needs to be focused on during actual application is prepared. Based on this keyword list, the regions where the keywords belonging to the keyword list appear in each original text image can be determined, thereby determining each target keyword region in the original text image. By focusing on smearing each target keyword region in the original text image, a smeared image is obtained. The smeared image and the corresponding original text image without smearing marks form a training sample pair to train the image restoration model.
[0079] In the embodiments of the present application, the original text image is smeared based on the target keyword regions to obtain the image to be restored. The original text image and the image to be restored obtained in this way are used as training samples of the image restoration model to train the image restoration model. The image restoration model obtained in this way has a good restoration effect on the keyword regions of the smeared images with smear marks passing through the keyword regions without affecting the overall text restoration performance, so as to facilitate text recognition for the restored image and improve the accuracy of text recognition.
[0080] In a possible implementation manner, determining each target keyword region in the original text image includes:
[0081] Performing text recognition on the original text image to obtain a text recognition result;
[0082] Based on the keyword table and the text recognition result, determine each target keyword region in the original text image.
[0083] In practical applications, after obtaining the original text image, the text in the original text image is recognized to obtain the recognized text in the original text image. Search for the same keywords as those in the keyword table among the recognized text, determine the regions where each recognized keyword is located, and determine each target keyword region in the text recognition result according to the regions where each keyword is located. The target keyword region is the image region where at least one keyword is located. Among them, the keyword table can be a keyword table for different fields. Because the keywords in different fields are different, different keyword tables can be determined according to the different application fields. For this application field, establish a keyword table with the keywords commonly used in the application field. The keyword tables for different fields are sets of keywords obtained by a computer or manually extracting keywords in different fields from literature and arranging them in a certain order.
[0084] In an example, for a given original text image without smear marks, such as Figure 2a shown, first use a text detection algorithm to obtain a text region detection frame, such as Figure 2b the text box corresponding to each line of text shown, and then use a text recognition algorithm to recognize the text content in each text box, such as Figure 2c shown. According to the recognized text result, in combination with the keyword table, locate the positions of the target keywords included in the image. Among them, assuming that the target keywords included in the text image shown in Figure 2a are determined to be "Heilongjiang" and "cycling" according to the keyword table, the regions corresponding to each target keyword can be framed with a position box as the target keyword region, such as Figure 2d shown.
[0085] Among them, the specific algorithms adopted by the text detection algorithm and the text recognition algorithm can be configured according to actual requirements, and the embodiments of the present application do not make limitations. For example, the scene text detection (An Efficient and Accurate Scene Text Detector, East) algorithm can be adopted for the text detection algorithm; the convolutional recurrent neural network (CRNN) can be used to implement the text recognition algorithm.
[0086] In the embodiments of the present disclosure, the target keyword regions are determined from the recognized text according to the keyword table, and the keywords that actually need to be focused on can be determined in the image, which is convenient for subsequent smearing processing of the keyword regions.
[0087] In a possible implementation manner, smearing processing is performed on each target keyword region in the original text image to obtain a to-be-restored image corresponding to the original text image, including:
[0088] For each target keyword region in the original text image, the target keyword region is marked in the original text image to obtain at least two first marking points;
[0089] Based on the smearing points in the original text image, smearing processing is performed on the original text image to obtain a to-be-restored image corresponding to the original text image, and the smearing points include all the first marking points corresponding to the target keyword regions.
[0090] For each original text image, after determining the target keyword regions in the original text image, the to-be-restored image can be obtained by smearing. Among them, the specific smearing method is not limited in the embodiments of the present application, as long as the final smearing trace passes through each target keyword region. As an optional method, the smearing trace can be a curve. The target keyword regions can be marked in the original text image. For example, for each target keyword region, at least two first marking points can be marked, and then based on all the first marking points corresponding to all the target keyword regions, curve fitting can be performed to obtain the to-be-restored image with the smearing trace. Since the curve is fitted based on the marking points of each target keyword region, the smearing trace in the to-be-restored image obtained by this method passes through the keyword regions.
[0091] Optionally, the specific marking method of the first marking point can be to randomly determine the marking position in the target keyword region for marking.
[0092] In an example, such as Figure 3As shown, two first marking points are marked in the position boxes corresponding to each target keyword in the original text image, and the two first marking points in each position box are used as the end points of curve fitting, so as to ensure that the curve passes through the target keyword area. Taking Figure 3 the keyword "Heilongjiang" in
[0093] as an example, two first marking points are randomly generated in the position box corresponding to the keyword "Heilongjiang" in the original text image. Then, based on these two first marking points, curve fitting can be performed. Since the curve is obtained by fitting the first marking points of the target keyword "Heilongjiang", the smeared traces in the image to be restored obtained in this way include the smeared traces of the area where the target keyword "Heilongjiang" is located.
[0094] In a possible implementation manner, the method further includes:
[0095] Determine at least one second marking point in the area outside each target keyword area in the original text image, and the smeared points include the second marking points.
[0096] In practical applications, in order to ensure that other texts outside the target keyword area can also be covered by the smeared traces, markings can also be made in the area outside the target keyword area to obtain at least one second marking point. The first marking points and the second marking points are jointly used as the smeared points, and the original text image is smeared according to these smeared points.
[0097] Optionally, the specific marking method of the second marking point can be to randomly determine the marking position in the area outside the target keyword area for marking.
[0098] Among them, the number of second marking points can be determined according to the actual application situation. For example, it can be 5-20.
[0099] In an example, as Figure 4 shown, after marking the first marking points at the target keywords "Heilongjiang" and "riding" in the original text image, multiple second marking points are marked in the area outside "Heilongjiang" and "riding". For example, in Figure 4Mark 12 second marked points at the positions corresponding to the words "line", "ride", "peace", etc. in . After that, curve fitting can be performed based on the first marked points and the second marked points. Since the curve is fitted based on the regions where the words "Heilongjiang", "riding", "line", "ride", "peace", etc. are located, the smudge marks in the image to be restored obtained in this way pass through the regions where the words "Heilongjiang", "riding", "line", "ride", "peace", etc. are located.
[0100] In a possible implementation, based on the smudge points in the original text image, smear processing is performed on the original text image, including:
[0101] Connect the smudge points in the original image in a preset direction to obtain a smudge mark;
[0102] Based on the smudge mark, smear processing is performed on the original text image.
[0103] In practical applications, the positions corresponding to each smudge point in the original text image can be connected in a preset direction to obtain a smudge mark, and smear processing is performed on the original text image according to the smudge mark. Among them, the preset direction can be pre-configured according to specific needs.
[0104] In an example, as Figure 5 shown, take the first marked points corresponding to "Heilongjiang" and "riding", and the 12 second marked points corresponding to the words "line", "ride", "peace", etc. as smudge points, connect them in a preset direction, and perform curve fitting to obtain a smudge mark, and smear the original text image based on the smudge mark. Since the curve of the smudge mark is fitted based on the marked points corresponding to the words "Heilongjiang", "riding", "line", "ride", "peace", etc., the smudge marks in the image to be restored obtained in this way pass through the regions where the words "Heilongjiang", "riding", "line", "ride", "peace", etc. are located.
[0105] Take each first marked point and second marked point in the original text image as a smudge point, connect the positions corresponding to each smudge point to obtain a smudge mark, and use this smudge mark to smear the original text image.
[0106] In a possible implementation, the preset direction includes the horizontal direction in the original text image or the vertical direction in the original text image.
[0107] In practical applications, the preset direction of the smear trace can be pre-configured according to specific needs. Optionally, according to the horizontal direction in the original text image, that is, the smear points are arranged in the order of the abscissa from left to right or from right to left to obtain the smear trace. Optionally, according to the vertical direction in the original text image, that is, the smear points are arranged in the order of the ordinate from top to bottom or from bottom to top to obtain the smear trace.
[0108] In one example, as Figure 6 shown, in the original text image, the marked points corresponding to the words "Heilongjiang", "riding", "line", "ride", "peace", etc. are used as smear points, and they are connected in the order from left to right horizontally, and two-by-two fitting is performed to obtain a curved smear trace. The curve of the smear trace is obtained by fitting the smear points corresponding to the words "Heilongjiang", "riding", "line", "ride", "peace", etc. Therefore, the smear trace in the image to be restored passes through the areas where the words "Heilongjiang", "riding", "line", "ride", "peace", etc. are located.
[0109] In one possible implementation, the method further includes:
[0110] Obtaining the attribute information of the smear trace;
[0111] Based on the smear trace, performing a smear process on the original text image, including:
[0112] Performing a smear process on the original text image according to the preset direction and the attribute information.
[0113] In practical applications, for the smear trace, the attribute information can also be pre-configured. The attribute information can be information related to the display effect of the smear trace. Perform a smear process on the original text image according to the preset direction and the attribute information of the smear trace.
[0114] In one possible implementation, the attribute information includes at least one of the curve color or the curve width of the smear trace.
[0115] In practical applications, the attribute information of the smear trace can be at least one of the curve color or the curve width of the smear trace. That is to say, the curve of the smear trace can be of different colors and different widths. Using the smear image obtained based on such a curve of the smear trace as a training sample can make the training sample more diverse, and the image restoration ability of the image restoration model trained in this way is stronger.
[0116] Next, through a specific application scenario, the process of obtaining the training sample of the technical solution of the present application will be described in detail. This embodiment is only one embodiment of the technical solution of the present application and does not represent all implementation manners of the technical solution of the present application.
[0117] AsFigure 7 As shown in Figure 7 , perform step S201 to input the original text image and the keyword table;
[0118] Among them, the original text image is a text image without smudging marks; the keyword table is a set of keywords obtained by a computer or manually extracting keywords from the literature and arranging them in a certain order. Multiple keywords can be determined in advance according to the actual application needs, and these keywords are used to establish the keyword table.
[0119] Perform step S202 to obtain the text positions and contents by using text detection and recognition algorithms;
[0120] After obtaining the original text image, detect and recognize the text in the original text image to obtain the text in the original text image. Specifically, a text detection algorithm can be used to obtain the detection frame corresponding to the text position, and then a text recognition algorithm can be used to recognize the text content in each text box.
[0121] Perform step S203 to locate the positions where the target keywords appear in the original text image;
[0122] According to the recognized text and in combination with the keyword table, locate the positions of the target keywords included in the original text image, and the positions of each target keyword area can be framed with position boxes. Specifically, the target keyword is a word determined according to at least one keyword in the keyword table, and the target keyword area is the image area where the keyword is located in the text content of the image to be restored. The text content of the image to be restored includes at least one keyword. The target keyword can be one keyword, or it can be a word composed of two or more keywords that are consecutive in the text content.
[0123] Among them, the keyword table can be a keyword table for different fields. Because the keywords in different fields are different, different keyword tables can be determined according to the different application fields. For this application field, the keywords commonly used in the application field are used to establish the keyword table. The keyword tables for different fields are sets of keywords obtained by a computer or manually extracting keywords from different fields of literature and arranging them in a certain order. In addition, keywords from each field in multiple fields can also be selected, and a keyword table common to multiple fields can be established using the keywords from multiple fields.
[0124] Perform step S204 to randomly generate two points within each target keyword area;
[0125] Mark the area where the target keyword is located in the original text image, and randomly generate two points as smearing points in the position box of each target keyword. Determining the smearing points according to the position where the target keyword is located can ensure that the target keyword area can be smeared during subsequent smearing, so that the obtained smeared image can meet the requirements for the training samples of the image restoration model.
[0126] Execute step S205, and randomly generate several points in the area outside the target keyword area;
[0127] In order to ensure that other text areas outside the target keyword area can also be covered by the smearing marks, mark the area outside the target keyword area and randomly generate several points as smearing points. The marked points in the target keyword area and the marked points in the area outside the target keyword area are used as smearing points together, and the original text image can be smeared according to these smearing points subsequently.
[0128] Execute step S206, and sort all the points in a preset manner;
[0129] Specifically, sort all the smearing points in the horizontal direction or the vertical direction. Specifically, in accordance with the horizontal direction in the original text image, that is, sort the smearing points in the order of the abscissa from left to right or from right to left. In accordance with the vertical direction in the original text image, that is, sort the smearing points in the order of the ordinate from top to bottom or from bottom to top.
[0130] Execute step S207, perform curve fitting on the sorted points in pairs to obtain a smearing trace;
[0131] According to the positions of the sorted smearing points, perform curve fitting in pairs to obtain a curve as the smearing trace of the original text image.
[0132] Execute step S208, set the attribute parameters of the smearing trace, including color and width, and draw the smearing trace in the original text image;
[0133] Set the attribute parameters for the smearing trace. The attribute parameters can be parameters related to the display effect such as the curve color parameter and curve width parameter of the smearing trace, and draw the smearing trace curve in the original text image according to the attribute parameters. Among them, the curves of the smearing trace can be of different colors and different widths. Using the smearing image obtained based on such smearing trace curves as the training sample can make the training sample more diverse, and the image restoration ability of the trained image restoration model is stronger.
[0134] Execute step S209, and save the original text image after adding the smearing trace.
[0135] After adding smudging marks to the original text image, save it as a training sample for the image restoration model.
[0136] Next, through a specific embodiment, the training process and application process of the image restoration model of the technical solution of the present application will be described in detail. This embodiment is only one embodiment of the technical solution of the present application and does not represent all implementation manners of the technical solution of the present application.
[0137] First, obtain each training sample of the image restoration model.
[0138] Specifically, obtain each original text image. For each original text image, perform text recognition. Based on the text recognition result and the keyword table, determine each target keyword region in the original text image. For each target keyword region, mark the target keyword region in the original text image to obtain at least two first marking points; determine at least one second marking point in the region outside each target keyword region in the original text image. Connect each first marking point and second marking point in the original image in a preset direction to obtain a smudging mark. Obtain the curve color and curve width of the smudging mark. Perform smudging processing on the original text image according to the preset direction, the curve color and curve width of the smudging mark to obtain the image to be restored corresponding to the original text image. Use each original text image and the image to be restored corresponding to each original text image as the training sample of the image restoration model.
[0139] Among them, the keyword table can be established for keywords in different fields. When training an image restoration model applied to different fields, the keyword table corresponding to that field can be used. Because the keywords in different fields are different, different keyword tables can be determined according to the differences in the application fields. For this application field, construct a keyword table with the commonly used keywords in the application field. When training an image restoration model for this application field, use the keyword table corresponding to this field to determine the target keyword regions of the original text images in the training samples and smudge the target keyword regions. In this way, the obtained training samples are more targeted, and the image restoration model trained with such training samples is more conducive to the restoration of text images in this application field, and the restoration effect of the text in this application field is better.
[0140] In addition, when determining the target keyword regions in the original text image, each target keyword region in the to-be-restored image of the training samples can include all the keywords in the keyword list corresponding to this field. This can ensure that when using the trained image restoration model to restore an image, the keyword regions in the smeared text image to be restored have all been smeared in the training samples, thus making the coverage of the training samples more comprehensive. That is to say, the more comprehensive the training samples are, the more information the image restoration model trained with these training samples can learn, which is more conducive to restoring all keyword regions in the image to be restored.
[0141] Secondly, use each training sample to train the initial image restoration model.
[0142] Specifically, in each training sample, the original text image without smearing marks is used as supervision to iteratively train the initial image restoration model. During the training process, the model restores the to-be-restored image with smearing marks corresponding to the original text image to obtain a restored image. Based on the pixel value difference between the restored image output by the model and the original text image without smearing marks, the value of the loss function is obtained. The model when the loss function converges is used as the image restoration model.
[0143] Thirdly, use the image restoration model to restore the text smeared image.
[0144] Specifically, obtain a text smeared image with smearing marks passing through the keyword region, as Figure 8 shown; the keyword region is the image region where the keywords included in the text content of the text smeared image are located; the keywords are included in the keyword list, which is the same as the keyword list used when determining the target keyword regions of the training samples, and this keyword list can also include other keywords. In this embodiment, the target keyword regions determined according to the keyword list are the regions where "Quanzhou" and "Silk Road" are located; this target keyword region can be the target keyword region smeared in the original text image of the training samples during the model training of the image restoration model. Input the text smeared image into the image restoration model to restore the text content covered by the smearing marks in the text smeared image to obtain the restored image corresponding to the text smeared image. For example, the restored image obtained after inputting the smeared image as Figure 8 shown into the image restoration model is as Figure 9 shown, and the restored image no longer includes smearing marks;
[0145] Finally, recognize the text in the restored image as Figure 9 shown.
[0146] Specifically, the text in the restored image can be recognized by OCR to obtain the text recognition result.
[0147] The text image recognition method provided by the embodiments of the present application uses an image restoration model to restore a text smeared image with smear marks passing through a keyword area. Under the condition of not affecting the overall text restoration performance, the restoration effect of the keyword area in the image containing smear marks is better, so as to facilitate text recognition for the restored image and improve the accuracy of text recognition.
[0148] And Figure 1 Based on the same principle as the method shown in Figure 10 An embodiment of the present disclosure also provides a text image recognition device 30, as
[0149] An image acquisition module 31, configured to acquire a text smeared image with smear marks passing through a keyword area; the keyword area is an image area where keywords included in the text content of the text smeared image are located; the keywords are included in a keyword table;
[0150] An image restoration module 32, configured to input the text smeared image into the image restoration model to restore the text content covered by the smear marks in the text smeared image, and obtain a restored image corresponding to the text smeared image;
[0151] A text recognition module 33, configured to perform text recognition on the text content in the restored image to obtain a text recognition result of the text smeared image.
[0152] In a possible implementation manner, the image restoration model is trained in the following manner:
[0153] Obtain each training sample, each training sample includes an original text image without smear marks and a to-be-restored image with smear marks corresponding to the original text image. The smear marks in the to-be-restored image pass through a target keyword area, and the target keyword area is an image area where keywords included in the text content of the to-be-restored image are located;
[0154] Iteratively train an initial image restoration model based on each training sample to obtain an image restoration model.
[0155] In a possible implementation manner, when the image restoration module 32 obtains each training sample, it is configured to:
[0156] Obtain each original text image;
[0157] For each original text image, determine each target keyword area included in the original text image;
[0158] For each original text image, perform a smearing process on each target keyword area in the original text image to obtain a to-be-restored image corresponding to the original text image.
[0159] In a possible implementation, when the image restoration module 32 determines each target keyword included in the original text image, it is used for:
[0160] Perform optical character recognition on the original text image to obtain an optical character recognition result;
[0161] Based on the keyword table and the optical character recognition result, determine each target keyword area in the original text image.
[0162] In a possible implementation, when the image restoration module 32 performs smearing processing on each target keyword included in the original text image to obtain a to-be-restored image corresponding to the original text image, it is used for:
[0163] For each target keyword area in the original text image, mark the target keyword area in the original text image to obtain at least two first marking points;
[0164] Based on the smearing points in the original text image, perform smearing processing on the original text image to obtain a to-be-restored image corresponding to the original text image, where the smearing points include the first marking points corresponding to all target keyword areas.
[0165] In a possible implementation, the apparatus 30 further includes a marking point determination module, which is used for:
[0166] Determine at least one second marking point in the area outside each target keyword area in the original text image, and the smearing points include the second marking points.
[0167] In a possible implementation, when the image restoration module 32 performs smearing processing on the original text image based on the smearing points in the original text image, it is used for:
[0168] Connect the smearing points in the original image in a preset direction to obtain a smearing trace;
[0169] Based on the smearing trace, perform smearing processing on the original text image.
[0170] In a possible implementation, the preset direction includes the horizontal direction in the original text image or the vertical direction in the original text image.
[0171] In a possible implementation, the apparatus 30 further includes an attribute information acquisition module, which is used for:
[0172] Acquire the attribute information of the smearing trace;
[0173] When the image restoration module 32 performs smearing processing on the original text image based on the smearing trace, it is used for:
[0174] Smear the original text image according to the preset direction and attribute information.
[0175] In a possible implementation, the attribute information includes at least one of the curve color or curve width of the smear mark.
[0176] The text image recognition device according to the embodiments of the present disclosure can execute the text image recognition method provided by the embodiments of the present disclosure, and its implementation principle is similar. The actions performed by each module in the text image recognition device according to the embodiments of the present disclosure correspond to the steps in the text image recognition method according to the embodiments of the present disclosure. For the detailed function description of each module of the text image recognition device, reference can be specifically made to the description in the corresponding text image recognition method shown above, and details are not described herein again. Figure 1 The text image recognition device provided by the embodiments of the present application uses an image restoration model to restore the text smear image with smear marks passing through the keyword area. Under the condition of not affecting the overall text restoration performance, the restoration effect of the keyword area in the image containing smear marks is better, so as to facilitate text recognition for the restored image and improve the accuracy of text recognition.
[0177] Among them, the text image recognition device may be a computer program (including program code) running in a computer device. For example, the text image recognition device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiments of the present application.
[0178] In some embodiments, the text image recognition device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the text image recognition device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the text image recognition method provided by the embodiments of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0179] In other embodiments, the text image recognition device provided by the embodiments of the present invention can be implemented in software.
[0180] In other embodiments, the text image recognition device provided by the embodiments of the present invention can be implemented in software. Figure 10An identification device for a text image stored in a memory, which may be software in the form of a program, a plug-in, etc., and includes a series of modules, including an image acquisition module 31, an image restoration module 32, and a character recognition module 33, for implementing the text image recognition method provided by the embodiments of the present invention.
[0181] The above embodiments introduce the text image recognition device from the perspective of virtual modules. The following introduces an electronic device from the perspective of physical modules, as specifically shown below:
[0182] An embodiment of the present application provides an electronic device, such as Figure 11 shown, Figure 11 The electronic device 8000 shown includes: a processor 8001 and a memory 8003. Among them, the processor 8001 and the memory 8003 are connected, such as connected through a bus 8002. Optionally, the electronic device 8000 may further include a transceiver 8004. It should be noted that in practical applications, the transceiver 8004 is not limited to one, and the structure of the electronic device 8000 does not constitute a limitation to the embodiments of the present application.
[0183] The processor 8001 may be a CPU, a general-purpose processor, a GPU, a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 8001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0184] The bus 8002 may include a path for transmitting information between the above components. The bus 8002 may be a PCI bus or an EISA bus, etc. The bus 8002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0185] The memory 8003 may be a ROM or other type of static storage device that can store static information and instructions, a RAM, or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM, a CD-ROM, or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0186] The memory 8003 is used to store the application program code for executing the solution of this application, and is controlled by the processor 8001 for execution. The processor 8001 is used to execute the application program code stored in the memory 8003 to implement the content shown in any of the foregoing method embodiments.
[0187] An embodiment of this application provides an electronic device. The electronic device in the embodiment of this application includes: one or more processors; a memory; one or more computer programs, where one or more computer programs are stored in the memory and are configured to be executed by one or more processors. When the one or more programs are executed by the processor, a text smear image in which the smear trace passes through the keyword area is obtained; the keyword area is an image area where the keywords included in the text content of the text smear image are located; the keywords are included in the keyword table; the text smear image is input into an image restoration model to restore the text content covered by the smear trace in the text smear image, and a restored image corresponding to the text smear image is obtained; the text content in the restored image is subjected to character recognition to obtain the character recognition result of the text smear image.
[0188] An embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a processor, the processor can execute the corresponding content in the foregoing method embodiment.
[0189] According to one aspect of this application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the above text image recognition method.
[0190] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly limited in order, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0191] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for recognizing text images, characterized in that, The method includes: Obtaining a text smear image in which the smear mark passes through a keyword area; the keyword area is an image area where the keywords included in the text content of the text smear image are located; the keywords are included in a keyword table; Inputting the text smear image into an image restoration model to restore the text content covered by the smear mark in the text smear image, and obtaining a restored image corresponding to the text smear image; Performing character recognition on the text content in the restored image to obtain a character recognition result of the text smear image; Wherein, the image restoration model is trained in the following manner: Obtaining each original text image; For each of the original text images, determining each target keyword area in the original text image; For each of the original text images, performing a smearing process on each of the target keyword areas in the original text image to obtain a to-be-restored image corresponding to the original text image, where the target keyword area is an image area where the keywords included in the text content of the to-be-restored image are located; Using each of the original text images and the corresponding to-be-restored images as each training sample, and performing iterative training on an initial image restoration model based on the training samples to obtain the image restoration model.
2. The method according to claim 1, wherein The determining each target keyword area in the original text image includes: Performing character recognition on the original text image to obtain a character recognition result; Based on the keyword table and the character recognition result, determining each of the target keyword areas in the original text image.
3. The method according to claim 1, wherein The performing a smearing process on each of the target keyword areas in the original text image to obtain a to-be-restored image corresponding to the original text image includes: For each of the target keyword areas in the original text image, marking the target keyword area in the original text image to obtain at least two first marking points; Based on the smearing points in the original text image, performing a smearing process on the original text image to obtain a to-be-restored image corresponding to the original text image, where the smearing points include the first marking points corresponding to all target keyword areas.
4. The method according to claim 3, wherein The method further includes: Determining at least one second marking point in an area outside each of the target keyword areas in the original text image, and the smearing points include the second marking point.
5. The method according to claim 3 or 4, characterized in that, The based on the smearing points in the original text image, performing a smearing process on the original text image includes: Connecting the smearing points in the original text image in a preset direction to obtain a smear mark; Based on the smear mark, performing a smearing process on the original text image.
6. The method according to claim 5, characterized in that, The preset direction includes a horizontal direction in the original text image or a vertical direction in the original text image.
7. The method according to claim 5, characterized in that The method further includes: Obtaining attribute information of the smear mark; The based on the smear mark, performing a smearing process on the original text image includes: Performing a smearing process on the original text image according to the preset direction and the attribute information.
8. The method according to claim 7, wherein The attribute information includes at least one of the curve color or curve width of the smear mark.
9. An apparatus for recognizing a text image, characterized in that The device includes: An image acquisition module for acquiring a text smear image in which a smear mark passes through a keyword area; the keyword area is an image area where keywords included in the text content of the text smear image are located; the keywords are included in a keyword table; An image restoration module for inputting the text smear image into an image restoration model to restore the text content covered by the smear mark in the text smear image, and obtaining a restored image corresponding to the text smear image; A character recognition module for performing character recognition on the text content in the restored image to obtain a character recognition result of the text smear image; Wherein, the image restoration model is obtained by the image restoration module through the following method: Obtain each original text image; For each of the original text images, determine each target keyword area in the original text image; For each of the original text images, perform a smear process on each of the target keyword areas in the original text image to obtain a to-be-restored image corresponding to the original text image, and the target keyword area is an image area where the keyword included in the text content of the to-be-restored image is located; Use each of the original text images and the corresponding to-be-restored images as each training sample, and perform iterative training on an initial image restoration model based on each of the training samples to obtain the image restoration model.
10. The device according to claim 9, characterized in that, When the image restoration module is used to determine each target keyword area in the original text image, it is specifically used for: Performing character recognition on the original text image to obtain a character recognition result; Based on the keyword table and the character recognition result, determine each of the target keyword areas in the original text image.
11. The device according to claim 9, characterized in that When the image restoration module is used to perform a smear process on each of the target keyword areas in the original text image to obtain a to-be-restored image corresponding to the original text image, it is specifically used for: For each of the target keyword areas in the original text image, mark the target keyword area in the original text image to obtain at least two first marked points; Based on the smear points in the original text image, perform a smear process on the original text image to obtain a to-be-restored image corresponding to the original text image, and the smear points include the first marked points corresponding to all target keyword areas.
12. The device according to claim 11, wherein It further includes: A marked point determination module for determining at least one second marked point in an area outside each of the target keyword areas in the original text image, and the smear points include the second marked points.
13. The device according to claim 11 or 12, characterized in that, When the image restoration module is used to perform a smear process on the original text image based on the smear points in the original text image, it is specifically used for: Connect the smear points in the original text image in a preset direction to obtain a smear mark; Based on the smear mark, perform a smear process on the original text image.
14. The device according to claim 13, wherein The preset direction includes a horizontal direction in the original text image or a vertical direction in the original text image.
15. The device according to claim 13, wherein It further includes: An attribute information acquisition module for: Acquiring the attribute information of the smear mark; When the image restoration module is used to perform smearing processing on the original text image based on the smear marks, it is specifically configured to: Perform smearing processing on the original text image according to the preset direction and the attribute information.
16. The device according to claim 15, characterized in that, The attribute information includes at least one of the curve color or the curve width of the smear marks.
17. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory; One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 8.
19. A computer program product, characterized in that, The computer program product includes computer instructions, and the processor executes the computer instructions to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Smearing image restoration method and device, storage medium and terminal equipment
CN111402156A